Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “machine learning for science”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Learning protein fitness models from evolutionary and assay-labeled data

Machine learning-based models of protein fitness typically learn from either unlabeled, evolutionarily related sequences or variant sequences with experimentally measured labels. For regimes where only limited experimental data are available, recent work has suggested methods for combining both sources of information. Toward that goal, we propose a simple combination approach that is competitive with, and on average outperforms more sophisticated methods. Our approach uses ridge regression on site-specific amino acid features combined with one probability density feature from modeling the evolutionary data. Within this approach, we find that a variational autoencoder-based probability density model showed the best overall performance, although any evolutionary density model can be used. Moreover, our analysis highlights the importance of systematic evaluations and sufficient baselines.

59 BASIC BIOLOGICAL SCIENCES↗

DIPS-Plus: The enhanced database of interacting protein structures for interface prediction

Abstract In this work, we expand on a dataset recently introduced for protein interface prediction (PIP), the Database of Interacting Protein Structures (DIPS), to present DIPS-Plus, an enhanced, feature-rich dataset of 42,112 complexes for machine learning of protein interfaces. While the original DIPS dataset contains only the Cartesian coordinates for atoms contained in the protein complex along with their types, DIPS-Plus contains multiple residue-level features including surface proximities, half-sphere amino acid compositions, and new profile hidden Markov model (HMM)-based sequence features for each amino acid, providing researchers a curated feature bank for training protein interface prediction methods. We demonstrate through rigorous benchmarks that training an existing state-of-the-art (SOTA) model for PIP on DIPS-Plus yields new SOTA results, surpassing the performance of some of the latest models trained on residue-level and atom-level encodings of protein complexes to date.

59 BASIC BIOLOGICAL SCIENCES↗

Understanding, discovery, and synthesis of 2D materials enabled by machine learning

Machine learning (ML) is becoming an effective tool for studying 2D materials. Taking as input computed or experimental materials data, ML algorithms predict the structural, electronic, mechanical, and chemical properties of 2D materials that have yet to be discovered. Such predictions expand investigations on how to synthesize 2D materials and use them in various applications, as well as greatly reduce the time and cost to discover and understand 2D materials. This tutorial review focuses on the understanding, discovery, and synthesis of 2D materials enabled by or benefiting from various ML techniques. Here, we introduce the most recent efforts to adopt ML in various fields of study regarding 2D materials and provide an outlook for future research opportunities. The adoption of ML is anticipated to accelerate and transform the study of 2D materials and their heterostructures.

2D Materials↗

Reducing Southern Ocean Shortwave Radiation Errors in the ERA5 Reanalysis with Machine Learning and 25 Years of Surface Observations

Earth system models struggle to simulate clouds and their radiative effects over the Southern Ocean, partly due to a lack of measurements and targeted cloud microphysics knowledge. We have evaluated biases of downwelling shortwave radiation in the ERA5 climate reanalysis using 25 years (1995–2019) of summertime surface measurements, collected on the Research and Supply Vessel (RSV) Aurora Australis, the Research Vessel (R/V) Investigator, and at Macquarie Island. During October–March daylight hours, the ERA5 simulation of SW down exhibited large errors (mean bias = 54 W m -2 , mean absolute error = 82 W m -2 , root-mean-square error = 132 W m -2 , and R 2 = 0.71). To determine whether we could improve these statistics, we bypassed ERA5’s radiative transfer model for SW down with machine learning–based models using a number of ERA5’s gridscale meteorological variables as predictors. These models were trained and tested with the surface measurements of SW down using a 10-fold shuffle split. An extreme gradient boosting (XGBoost) and a random forest–based model setup had the best performance relative to ERA5, both with a near complete reduction of the mean bias error, a decrease in the mean absolute error and root-mean-square error by 25% ± 3%, and an increase in the R 2 value of 5% ± 1% over the 10 splits. Large improvements occurred at higher latitudes and cyclone cold sectors, where ERA5 performed most poorly. We further interpret our methods using Shapley additive explanations. Our results indicate that data-driven techniques could have an important role in simulating surface radiation fluxes and in improving reanalysis products.

54 ENVIRONMENTAL SCIENCES↗

Membrane lipids drive formation of KRAS4b-RAF1 RBDCRD nanoclusters on the membrane

The oncogene RAS, extensively studied for decades, presents persistent gaps in understanding, hindering the development of effective therapeutic strategies due to a lack of precise details on how RAS initiates MAPK signaling with RAF effector proteins at the plasma membrane. Recent advances in X-ray crystallography, cryo-EM, and super-resolution fluorescence microscopy offer structural and spatial insights, yet the molecular mechanisms involving protein-protein and protein-lipid interactions in RAS-mediated signaling require further characterization. This study utilizes single-molecule experimental techniques, nuclear magnetic resonance spectroscopy, and the computational Machine-Learned Modeling Infrastructure (MuMMI) to examine KRAS4b and RAF1 on a biologically relevant lipid bilayer. MuMMI captures long-timescale events while preserving detailed atomic descriptions, providing testable models for experimental validation. Both in vitro and computational studies reveal that RBDCRD binding alters KRAS lateral diffusion on the lipid bilayer, increasing cluster size and decreasing diffusion. RAS and membrane binding cause hydrophobic residues in the CRD region to penetrate the bilayer, stabilizing complexes through β-strand elongation. These cooperative interactions among lipids, KRAS4b, and RAF1 are proposed as essential for forming nanoclusters, potentially a critical step in MAP kinase signal activation.

59 BASIC BIOLOGICAL SCIENCES↗

Machine learning and shallow groundwater chemistry to identify geothermal prospects in the Great Basin, USA

This study discovers various geothermal prospects in the Great Basin, USA based on shallow groundwater chemical (geochemical) data. The geochemical data are expected to include hidden (latent) information that is a proxy for geothermal prospectivity. We processed the sparse geochemical data in the Great Basin at 14,341 locations including 18 attributes. Next, a non-negative matrix factorization with customized k-means clustering is applied to the geochemical data matrix that automatically finds three hidden geothermal signatures representing modestly, moderately, and highly confident geothermal prospects. The algorithm also evaluated the probability of occurrence of these types of resources through the studied region. There is a consistency between regional geothermal prospectivity as estimated by our ML methodology and the traditional play fairway analysis conducted over a portion of the study area. We also identify the dominant data attributes associated with each signature. Finally, our ML analyses allow us to reconstruct attributes from sparse into continuous over the study domain. The predicted continuous attributes can be used for future detailed geothermal explorations in the Great Basin.

15 GEOTHERMAL ENERGY↗

Machine-learning-based automatic small-angle measurement between planar surfaces in interferometer images: A 2D multilayer Laue lenses case

Here, we report a new machine-learning-based approach to automatically measure the small angle between multiple planar surfaces characterized by white light interferometers. By applying an unsupervised clustering algorithm, DBSCAN (Density-Based Spatial Clustering of Applications with Noise), the multiple surfaces in an interferometer image are automatically identified as distinct surfaces. The angles between every two surfaces are then calculated through the surface fitting. This method can be applied to multiple surfaces regardless of their shapes and locations and significantly simplifies the angle measurement procedure. Using the developed method, we have demonstrated a quick and precise angle measurement for the alignment of 2D Multilayer Laue Lenses (MLLs) for the development of high-resolution x-ray microscopy. This automatic, accurate, and robust small-angle measurement method is compatible with widely used white light interferometers and can be further applied to other metrology applications of interferometer results.

36 MATERIALS SCIENCE↗

The drivers and predictability of wildfire re-burns in the western United States (US)

Evidence is mounting that the effectiveness of using prescribed burns as a management tactic may be diminishing due to the higher incidence of wildfire re-burns. The development of predictive models of re-burns is thus essential to better understand their primary drivers so that forest management practices can be updated to account for these events. First, we assess the potential for human activity as a driver of re-burns by evaluating re-burn trends both within and outside of the wildland–urban interface (WUI) of the western US. Next, we investigate the predictability of re-burns through the application of both random forest and the explanatory machine learning non-negative matrix factorization using k-means clustering (NMFk) algorithms to predict re-burn occurrence over California based on a number of climate factors. Our findings indicate that while most states showed increasing trends within the WUI when trends were conducted over longer moving windows (e.g. 20 years), California was the only state where the rate of increase was consistently higher in the WUI, indicating a stronger potential for human activity as a driver in that location. Furthermore, we find model performance was found to be robust over most of California (Testing F1 scores = 0.688), although results were highly variable based on EPA level III Ecoregion (F1 scores = 0.0–0.778). Insights provided from this study will lead to a better understanding of climate and human activity drivers of re-burns and how these vary at broad spatial scales so that improvements in forest management practices can be tuned according to the level of change that is expected for a given region.

54 ENVIRONMENTAL SCIENCES↗

Invariant surface elastic properties in FCC metals and their correlation to bulk properties revealed by machine learning methods

In this work, we present a combination of machine-learned models that predicts the surface elastic properties of general free surfaces in face-centered cubic (FCC) metals. These models are built by combining a semi-analytical method based on atomistic simulations to calculate surface properties with the artificial neural network (ANN) method or the boosted regression tree (BRT) method. The latter is also used to link bulk properties and surface orientation to surface properties. The surface elastic properties are represented by their invariants considering plane elasticity within a polar method. The resulting models are shown to accurately predict the surface elastic properties of seven pure FCC metals (Cu, Ni, Ag, Au, Al, Pd, Pt). The BRT model reveals the correlations between bulk and corresponding surface properties in terms of invariants, which can be used to guide the design of complex nano-sized particles, wires and films. Finally, by expressing the surface excess energy density as a function of surface elastic invariants, fast predictions of surface energy as a function of in-plane deformations can be made from these model constructs.

36 MATERIALS SCIENCE↗

FY23 Progress Report: SRNL Analysis of ICCWR LCM and WAMS data for Corrosion and Cracking

Algorithms for Machine Learning (ML) and data analysis for the 3013 Surveillance Program have been developed in an ongoing collaborative effort by the Savannah River National Laboratory (SRNL) and the University of South Carolina (USC). The objective of the algorithms is to automate the identification of corrosion and crack formation in the Inner Container Closure Weld Region (ICCWR) of the canister system used to store Pu-bearing material. Data for corrosion and cracking is collected from large binary files generated by a Laser Confocal Microscope (LCM), the Wide Area 3D Measurement System (WAMS), and in a recent proposal, by a Scanning Electron Microscope (SEM). The ML software uses the physical attributes in the data files (e.g., one or all of: height, color, and grayscale values as functions of position in a plane projection) to detect the presence of surface corrosion and cracking after being trained on similar data with the features to be detected labeled.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Machine learning in nuclear materials research

Nuclear materials are often demanded to function for extended time in extreme environments, including high radiation fluxes with associated transmutations, high temperature and temperature gradients, mechanical stresses, and corrosive coolants. They also have a wide range of microstructural and chemical makeups, resulting in multifaceted and often out-of-equilibrium interactions. Machine learning (ML) is increasingly being used to tackle these complex time-dependent interactions and aid researchers in developing models and making predictions, sometimes with better accuracy than traditional modeling that focuses on one or two parameters at a time. Conventional practices of acquiring new experimental data in nuclear materials research are often slow and expensive, limiting the opportunity for data-centric ML, but new methods are changing that paradigm. Here we review high-throughput computational and experimental data approaches, especially robotic experimentation and active learning that is based on Gaussian process and Bayesian optimization. We show ML examples in structural materials (e.g., reactor pressure vessel (RPV) alloys and radiation detecting scintillating materials) and highlight new techniques of high-throughput sample preparation and characterizations, and automated radiation/environmental exposures and real-time online diagnostics. Herein, this review suggests that ML models of material constitutive relations in plasticity, damage, and even electronic and optical responses to radiation are likely to become powerful tools as they develop. Finally, we speculate on how the recent trends of using natural language processing (NLP) to aid the collection and analysis of literature data, interpretable artificial intelligence (AI), and the use of streamlined scripting, database, workflow management, and cloud computing platforms that will soon make the utilization of ML techniques as commonplace as the spreadsheet curve-fitting practices of today.

36 MATERIALS SCIENCE↗

Generative AI models for learning flow maps of stochastic dynamical systems in bounded domains

Simulating stochastic differential equations (SDEs) in bounded domains, presents significant computational challenges due to particle exit phenomena, which requires accurate modeling of interior stochastic dynamics and boundary interactions. Despite the success of machine learning-based methods in learning SDEs, existing learning methods are not applicable to SDEs in bounded domains because they cannot accurately capture the particle exit dynamics. We present a unified hybrid data-driven approach that combines a conditional diffusion model with an exit prediction neural network to capture both interior stochastic dynamics and boundary exit phenomena. Our ML model consists of two major components: a neural network that learns exit probabilities using binary cross-entropy loss with rigorous convergence guarantees, and a training-free diffusion model that generates state transitions for non-exiting particles using closed-form score functions. The two components are integrated through a probabilistic sampling algorithm that determines particle exit at each time step and generates appropriate state transitions. Here, the performance of the proposed approach is demonstrated via three test cases: a one-dimensional simplified problem for theoretical verification, a two-dimensional advection-diffusion problem in a bounded domain, and a three-dimensional problem of interest to magnetically confined fusion plasmas.

Bounded domains↗

Improving streamflow predictions across CONUS by integrating advanced machine learning models and diverse data

Accurate streamflow prediction is crucial to understand climate impacts on water resources and develop effective adaption strategies. A global long short-term memory (LSTM) model, using data from multiple basins, can enhance streamflow prediction, yet acquiring detailed basin attributes remains a challenge. To overcome this, we introduce the Geo-vision transformer (ViT)-LSTM model, a novel approach that enriches LSTM predictions by integrating basin attributes derived from remote sensing with a ViT architecture. Applied to 531 basins across the Contiguous United States, our method demonstrated superior prediction accuracy in both temporal and spatiotemporal extrapolation scenarios. Geo-ViT-LSTM marks a significant advancement in land surface modeling, providing a more comprehensive and effective tool for better understanding the environment responses to climate change.

Tayal, Kshitij↗

Toward machine learning interatomic potentials for modeling uranium mononitride

Uranium mononitride (UN) is a promising accident-tolerant fuel because of its high fissile density and high thermal conductivity. In this study, we developed the first machine learning interatomic potentials for reliable atomic-scale modeling of UN at finite temperatures. We constructed a training set using density functional theory (DFT) calculations that was enriched through an active learning procedure, and two neural network potentials were generated. Both potentials successfully reproduce key thermophysical properties of interest, such as temperature-dependent lattice parameter, specific heat capacity, and bulk modulus. We also evaluated the energy of stoichiometric defect reactions and defect migration barriers and found close agreement with DFT predictions, demonstrating that our potentials can be used for modeling defects in UN. Additional tests provide evidence that our potentials are reliable for simulating diffusion, noble gas impurities, and radiation damage.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Extension of Large Fire Emissions From Summer to Autumn and Its Drivers in the Western US

Abstract Burned areas in the western US have increased ten‐fold since 1980s, which are attributable to multiple factors, including increasing heat, changing precipitation patterns, and extended drought. To better understand how these factors contribute to large fire emissions (gridded monthly fire emissions >95th percentile of all the fire emissions in the western US; 0.009 Gg/month), we build a machine learning model to predict fire emissions (PM 2.5 ) over the western US at 0.25° resolution, interpreted using explainable artificial intelligence (XAI). From the predictor contributions derived from XAI, we conduct k‐means clustering analysis to identify four clusters of predictor variables representing different drivers of large fire emissions. The four clusters feature the contributions of fuel load (Cluster 1) and different levels of dryness (Cluster 2–4), controlled by fuel moisture, drought condition, and fire‐favorable large‐scale meteorological patterns featuring high temperature, high pressure, and low relative humidity. In the past two decades, large fire emissions peak in summer. However, large fire emissions increased significantly in September and October in 2010–2020 relative to 2000–2009, extending the peak large fire emissions from summer to autumn. The larger enhancements of large fire emissions during autumn compared to summer are contributed by decreased fuel moisture, along with more frequent concurrent fire‐favorable large‐scale meteorological patterns and drought. These results highlight fuel drying as a common driver supported by multiple drivers, such as warmer temperature and more frequent synoptic patterns favorable for fires, in increasing the autumn risk of large fire emissions across the western US.

54 ENVIRONMENTAL SCIENCES↗

Gene network centrality analysis identifies key regulators coordinating day-night metabolic transitions in Synechococcus elongatus PCC 7942 despite limited accuracy in predicting direct regulator-gene interactions

Synechococcus elongatus PCC 7942 is a model organism for studying circadian regulation and bioproduction, where precise temporal control of metabolism significantly impacts photosynthetic efficiency and CO 2 -to-bioproduct conversion. Despite extensive research on core clock components, our understanding of the broader regulatory network orchestrating genome-wide metabolic transitions remains incomplete. We address this gap by applying machine learning tools and network analysis to investigate the transcriptional architecture governing circadian-controlled gene expression. While our approach showed moderate accuracy in predicting individual transcription factor-gene interactions - a common challenge with real expression data - network-level topological analysis successfully revealed the organizational principles of circadian regulation. Our analysis identified distinct regulatory modules coordinating day-night metabolic transitions, with photosynthesis and carbon/nitrogen metabolism controlled by day-phase regulators, while nighttime modules orchestrate glycogen mobilization and redox metabolism. Through network centrality analysis, we identified potentially significant but previously understudied transcriptional regulators: HimA as a putative DNA architecture regulator, and TetR and SrrB as potential coordinators of nighttime metabolism, working alongside established global regulators RpaA and RpaB. This work demonstrates how network-level analysis can extract biologically meaningful insights despite limitations in predicting direct regulatory interactions. The regulatory principles uncovered here advance our understanding of how cyanobacteria coordinate complex metabolic transitions and may inform metabolic engineering strategies for enhanced photosynthetic bioproduction from CO 2 .

59 BASIC BIOLOGICAL SCIENCES↗

Identifying Key Drivers of Wildfires in the Contiguous US Using Machine Learning and Game Theory Interpretation

Abstract Understanding the complex interrelationships between wildfire and its environmental and anthropogenic controls is crucial for wildfire modeling and management. Although machine learning (ML) models have yielded significant improvements in wildfire predictions, their limited interpretability has been an obstacle for their use in advancing understanding of wildfires. This study builds an ML model incorporating predictors of local meteorology, land‐surface characteristics, and socioeconomic variables to predict monthly burned area at grid cells of 0.25° × 0.25° resolution over the contiguous United States. Besides these predictors, we construct and include predictors representing the large‐scale circulation patterns conducive to wildfires, which largely improves the temporal correlations in several regions by 14%–44%. The Shapley additive explanation is introduced to quantify the contributions of the predictors to burned area. Results show a key role of longitude and latitude in delineating fire regimes with different temporal patterns of burned area. The model captures the physical relationship between burned area and vapor pressure deficit, relative humidity (RH), and energy release component (ERC), in agreement with the prior findings. Aggregating the contribution of predictor variables of all the grids by region, analyses show that ERC is the major contributor accounting for 14%–27% to large burned areas in the western US. In contrast, there is no leading factor contributing to large burned areas in the eastern US, although large‐scale circulation patterns featuring less active upper‐level ridge‐trough and low RH two months earlier in winter contribute relatively more to large burned areas in spring in the southeastern US.

54 ENVIRONMENTAL SCIENCES↗

Machine learning for ultrasonic nondestructive examination of welding defects: A systematic review

Recent years have seen a substantial increase in the application of machine learning (ML) for automated analysis of nondestructive examination (NDE) data. One of the applications of interest is the use of ML for the analysis of data from in-service inspection of welds in nuclear power and other industries. These types of inspections are performed in accordance with criteria described in the ASME Boiler and Pressure Vessel Code and require the use of reliable NDE techniques. The rapid growth in ML methods and the diversity of possible approaches indicate a need to assess the current capabilities of ML and automated data analysis for NDE and identify any gaps or shortcomings in current ML technologies as applied to the automated analysis of NDE data. In particular, there is a need to determine the impact of ML on the NDE reliability. This paper discusses the findings from a literature survey on the current state of ML for the automated analysis of data from ultrasonic NDE of weld flaws. It discusses an overview of ultrasonic NDE as used for weld inspections in nuclear power and other industries. Herein, data sets and ML models used in the literature are summarized, along with a generally applicable workflow for ML. Findings on the capabilities, limitations and potential gaps in feature selection, data selection, and ML model optimization are discussed. The paper identified several needs for quantifying and validating the performance of ML methods for ultrasonic NDE, including the need for common data sets.

36 MATERIALS SCIENCE↗