Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “machine learning for science”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

FracML: A Machine Learning Based Tool to Quantify Reservoir Scale Fracture Network for CO2 Storage

Poster on “FRACML: A Machine Learning Based Tool to Quantify Reservoir Scale Fracture Network for CO2 Storage” for the CCUS 2025 conference held in Houston, Texas March 3-5, 2025. The accurate characterization of subsurface fracture networks is essential for the secure operation of carbon capture, utilization, and storage (CCUS) projects. A thorough understanding of the spatial distribution of subsurface faults and fractures is crucial for predicting CO2 plume evolution and minimizing risks such as potential leakage into overlying formations or induced seismicity. In this context, robust fracture network quantification plays a pivotal role in reservoir management, providing the data necessary to fine-tune operational parameters, and ensure the environmental and economic viability of CCUS projects. As part of the U.S. Department of Energy’s SMART (Science-informed Machine Learning for Accelerating Real-time Decisions in Subsurface Applications) initiative, we focused on the development and application of a machine learning-based tool (FRACML) designed to quantify and map fracture networks using real-world (non-synthetic) data from an active CO2 injection site. Our objective is to demonstrate the utility of this tool in improving operational efficiency and safety across CCUS sites.

artifical intelligence / machine learning (AI/ML)↗

Protein sequence design with a learned potential

The task of protein sequence design is central to nearly all rational protein engineering problems, and enormous effort has gone into the development of energy functions to guide design. Here, we investigate the capability of a deep neural network model to automate design of sequences onto protein backbones, having learned directly from crystal structure data and without any human-specified priors. The model generalizes to native topologies not seen during training, producing experimentally stable designs. We evaluate the generalizability of our method to a de novo TIM-barrel scaffold. The model produces novel sequences, and high-resolution crystal structures of two designs show excellent agreement with in silico models. Our findings demonstrate the tractability of an entirely learned method for protein sequence design.

59 BASIC BIOLOGICAL SCIENCES↗

Enhanced Surgical Decision-Making Tools in Breast Cancer: Predicting 2-Year Postoperative Physical, Sexual, and Psychosocial Well-Being following Mastectomy and Breast Reconstruction (INSPiRED 004)

Background: We sought to predict clinically meaningful changes in physical, sexual, and psychosocial well-being for women undergoing cancer-related mastectomy and breast reconstruction 2 years after surgery using machine learning (ML) algorithms trained on clinical and patient-reported outcomes data. Patients and Methods: We used data from women undergoing mastectomy and reconstruction at 11 study sites in North America to develop three distinct ML models. We used data of ten sites to predict clinically meaningful improvement or worsening by comparing pre-surgical scores with 2 year follow-up data measured by validated Breast-Q domains. We employed ten-fold cross-validation to train and test the algorithms, and then externally validated them using the 11th site’s data. We considered area-under-the-receiver-operating-characteristics-curve (AUC) as the primary metric to evaluate performance. Results: Overall, between 1454 and 1538 patients completed 2 year follow-up with data for physical, sexual, and psychosocial well-being. In the hold-out validation set, our ML algorithms were able to predict clinically significant changes in physical well-being (chest and upper body) (worsened: AUC range 0.69–0.70; improved: AUC range 0.81–0.82), sexual well-being (worsened: AUC range 0.76–0.77; improved: AUC range 0.74–0.76), and psychosocial well-being (worsened: AUC range 0.64–0.66; improved: AUC range 0.66–0.66). Baseline patient-reported outcome (PRO) variables showed the largest influence on model predictions. Conclusions: Machine learning can predict long-term individual PROs of patients undergoing postmastectomy breast reconstruction with acceptable accuracy. This may better help patients and clinicians make informed decisions regarding expected long-term effect of treatment, facilitate patient-centered care, and ultimately improve postoperative health-related quality of life.

60 APPLIED LIFE SCIENCES↗

Modeling Spatial Distribution of Snow Water Equivalent by Combining Meteorological and Satellite Data with Lidar Maps

Abstract An accurate characterization of the water content of snowpack, or snow water equivalent (SWE), is necessary to quantify water availability and constrain hydrologic and land surface models. Recently, airborne observations (e.g., lidar) have emerged as a promising method to accurately quantify SWE at high resolutions (scales of ∼100 m and finer). However, the frequency of these observations is very low, typically once or twice per season in the Rocky Mountains of Colorado. Here, we present a machine learning framework that is based on random forests to model temporally sparse lidar-derived SWE, enabling estimation of SWE at unmapped time points. We approximated the physical processes governing snow accumulation and melt as well as snow characteristics by obtaining 15 different variables from gridded estimates of precipitation, temperature, surface reflectance, elevation, and canopy. Results showed that, in the Rocky Mountains of Colorado, our framework is capable of modeling SWE with a higher accuracy when compared with estimates generated by the Snow Data Assimilation System (SNODAS). The mean value of the coefficient of determination R 2 using our approach was 0.57, and the root-mean-square error (RMSE) was 13 cm, which was a significant improvement over SNODAS (mean R 2 = 0.13; RMSE = 20 cm). We explored the relative importance of the input variables and observed that, at the spatial resolution of 800 m, meteorological variables are more important drivers of predictive accuracy than surface variables that characterize the properties of snow on the ground. This research provides a framework to expand the applicability of lidar-derived SWE to unmapped time points. Significance Statement Snowpack is the main source of freshwater for close to 2 billion people globally and needs to be estimated accurately. Mountainous snowpack is highly variable and is challenging to quantify. Recently, lidar technology has been employed to observe snow in great detail, but it is costly and can only be used sparingly. To counter that, we use machine learning to estimate snowpack when lidar data are not available. We approximate the processes that govern snowpack by incorporating meteorological and satellite data. We found that variables associated with precipitation and temperature have more predictive power than variables that characterize snowpack properties. Our work helps to improve snowpack estimation, which is critical for sustainable management of water resources.

54 ENVIRONMENTAL SCIENCES↗

The effect of 10 at.% Al addition on the hydrogen storage properties of the Ti 0.33 V 0.33 Nb 0.33 multi-principal element alloy

We report here a thorough study on the effect of 10 at.% Al addition into the ternary equimolar Ti 0.33 V 0.33 Nb 0.33 alloy on the hydrogen storage properties. Despite a decrease of the storage capacity by 20%, several other properties are enhanced by the presence of Al. The hydride formation is destabilized in the quaternary alloy as compared to the pristine ternary composition, as also confirmed by machine learning approach. The hydrogen desorption occurs at lower temperature in the Al-containing alloy relative to the initial material. Moreover, the Al presence improves the stability during hydrogen absorption/desorption cycling without significant loss of the capacity and phase segregation. Finally, this study proves that Al addition into multi-principal element alloys is a promising strategy for the design of novel materials for hydrogen storage.

36 MATERIALS SCIENCE↗

Automated model-predictive design of synthetic promoters to control transcriptional profiles in bacteria

Abstract Transcription rates are regulated by the interactions between RNA polymerase, sigma factor, and promoter DNA sequences in bacteria. However, it remains unclear how non-canonical sequence motifs collectively control transcription rates. Here, we combine massively parallel assays, biophysics, and machine learning to develop a 346-parameter model that predicts site-specific transcription initiation rates for any σ 70 promoter sequence, validated across 22132 bacterial promoters with diverse sequences. We apply the model to predict genetic context effects, design σ 70 promoters with desired transcription rates, and identify undesired promoters inside engineered genetic systems. The model provides a biophysical basis for understanding gene regulation in natural genetic systems and precise transcriptional control for engineering synthetic genetic systems.

42 ENGINEERING↗

LandScan Global 30 Arcsecond Annual Global Gridded Population Datasets from 2000 to 2022

Abstract Oak Ridge National Laboratory (ORNL) annually develops the LandScan Global (LSG) dataset, a 30 arcsecond global gridded population dataset representing global ambient human population distribution. This multivariable dasymetric model disaggregates census counts within administrative boundaries using ancillary data. Each country’s distribution reflects cultural and socioeconomic patterns; manual validations yield a unique global dataset for assessing populations at risk. For over two decades, LSG has been a standard for estimating populations at risk, aiding U.S. federal government, academia and humanitarian organizations. During disasters such as the 2004 Indian Ocean tsunami and the 2010 Haiti earthquake and geopolitical crises such as the Syrian civil war and the 2022 Russian invasion of Ukraine, LSG supported scientific and operational communities in emergency response and recovery. In 2022, LSG datasets from 2000 onward were made publicly available through ORNL’s LandScan Portal. This data descriptor details our methodology and the application of geospatial science and machine learning to geographic and demographic data, highlighting uses in urban resiliency, emergency management, disaster response, and human health and security.

Science & Technology - Other Topics↗

Machine learning prediction of the mechanical properties of refractory multicomponent alloys based on a dataset of phase and first principles simulation

In this work, a dataset including structural and mechanical properties of refractory multicomponent alloys was developed by fusing computations of phase diagram (CALPHAD) and density functional theory (DFT). The refractory multicomponent alloys, also named refractory complex concentrated alloys (CCAs) which contain 2–5 types of refractory elements were constructed based on Special Quasi-random Structure (SQS). The phase of alloys was predicted using CALPHAD and the mechanical property of alloys with stable and single body-centered cubic (BCC) at high temperature (over 1,500°C) was investigated using DFT-based simulation. As a result, a dataset with 393 refractory alloys and 12 features, including volume, melting temperature, density, energy, elastic constants, mechanical moduli, and hardness, were produced. To test the capability of the dataset on supporting machine learning (ML) study to investigate the property of CCAs, CALPHAD, and DFT calculations were compared with principal components analysis (PCA) technique and rule of mixture (ROM), respectively. It is demonstrated that the CALPHAD and DFT results are more in line with experimental observations for the alloy phase, structural and mechanical properties. Furthermore, the data were utilized to train a verity of ML models to predict the performance of certain CCAs with advanced mechanical properties, highlighting the usefulness of the dataset for ML technique on CCA property prediction.

36 MATERIALS SCIENCE↗

Accelerating the discovery of novel magnetic materials using machine learning–guided adaptive feedback

Magnetic materials are essential for energy generation and information devices, and they play an important role in advanced technologies and green energy economies. Currently, the most widely used magnets contain rare earth (RE) elements. An outstanding challenge of notable scientific interest is the discovery and synthesis of novel magnetic materials without RE elements that meet the performance and cost goals for advanced electromagnetic devices. Here, we report our discovery and synthesis of an RE-free magnetic compound, Fe 3 CoB 2 , through an efficient feedback framework by integrating machine learning (ML), an adaptive genetic algorithm, first-principles calculations, and experimental synthesis. Magnetic measurements show that Fe 3 CoB 2 exhibits a high magnetic anisotropy ( K 1 = 1.2 MJ/m 3 ) and saturation magnetic polarization ( J s = 1.39 T), which is suitable for RE-free permanent-magnet applications. Our ML-guided approach presents a promising paradigm for efficient materials design and discovery and can also be applied to the search for other functional materials.

36 MATERIALS SCIENCE↗

MLSPICE: Machine Learning based SPICE Modeling Platform for Power Magnetics

Electrical power converters are critical to a wide range of applications ranging from renewable integration to transportation electrification, and can be a key factor determining the size, weight, and efficiency of energy conversion systems. Magnetic components are typically the largest and least efficient components in power electronics. While there have been major strides in the modeling and analysis of power semiconductor devices and circuit simulations, the necessary advances in the design of power magnetics have lagged. In this project, we have transformed the modeling and design of power magnetics with machine learning enabled methods and catalyze simultaneous disruptive improvements for ML-based power electronics design tools. A fully automated open-source machine learning based magnetics modeling platform – the MagNet project - with innovations in full stack have been developed to greatly accelerate the design process and provide new insights to magnetic material and geometry design. The ARPA-E funded MagNet platform contains three major building blocks: 1) a ML-Integrated Data Acquisition System (MIDAS): a highly automated data acquisition testbed which is capable of measuring a large number of magnetic cores with a wide range of electrical circuit excitations; 2) a ML-integrated Core Loss Model (MICLM): a machine-learning trained modeling method for modeling the core loss and saturation effects of magnetic materials for arbitrary excitation waveforms; 3) ML-guided Magnetics SPICE Simulation Tool (PMSPICE): a fully integrated CAD tool which can simulate the magnetics in SPICE. It can help the designers to quickly model the linear and non-linear characteristics of magnetic components and evaluate their behavior in SPICE simulations. The developed MagNet system has fully demonstrated the proposed performance target and has been open sourced to the entire power electronics community to advance the modeling and design of power magnetics from many different angles.

36 MATERIALS SCIENCE↗

Machine Learning based Correlation of the Mechanical Properties of Sub-sized and Standard-sized Specimens

Mechanical testing with sub-sized specimens is essential in the nuclear industry, offering the ability to conduct tests in confined spaces with lower irradiation and expediting material qualification. However, smaller specimens exhibit different material behavior across scales, a phenomenon known as the "specimen size effect". In this study, we compiled over 1,000 tensile testing records, covering 54 parameters such as material type, composition, manufacturing details, irradiation conditions, specimen dimensions, and tensile properties through a comprehensive literature review. We focus on correlating sub-sized and standard specimens’ tensile mechanical properties on SS316 alloy, which has the most extensive dataset available. We explore ML-based models and uncertainty quantification for tensile properties, analyze key factors influencing these properties, and compare the effectiveness of ML models with existing analytical methods in addressing the specimen size effect.

tensile properties↗

Controlling reversible phase transitions in rare-earth nickelates for novel memory devices

Resistive switching in correlated complex oxides is lucrative for emerging applications in neuromorphic computing, and densely scaled non-volatile memory. Electrical conductance of such complex oxides can be controllable switched across multiple orders of magnitude by either (a) electroforming a conduction channel (e.g., in tungsten oxide), or (b) inducing Mott-Hubbard transition (e.g., in rare-earth nickelates)– both via controlled migration of defects (such as oxygen vacancies) under applied bias. Nevertheless, the promise of such defect-driven electronic transitions are far from realized due to a lack of fundamental understanding of the atomic-scale processes that underlie migration and spatiotemporal evolution of oxygen vacancies over nano-to-mesoscopic length/timescales under applied electric field. In this project, we employ a synergistic integration of density functional theory (DFT) calculations, ab initio/classical molecular dynamics (AIMD/CMD) simulations, machine learning (ML), precision synthesis, and multi-modal X-ray imaging experiments to address this knowledge gap. Such an integrated approach offers to elucidate the correlations between subtle structural distortion and oxidation states; treat localized charge carriers; describe defect/ion transport in the presence of electric field; and, in turn, greatly advance the current understanding of microstructural evolution in complex oxides under applied bias. The fundamental knowledge gained from this work will enable precise control over hierarchical defect structures and unravel new routes to manipulate resistance states in complex oxides. This, in turn, will accelerate design of novel devices with desired set of neural functionalities, and high-speed densely-scaled resistive random access memory technologies.

36 MATERIALS SCIENCE↗

Development and Validation of a Non-Invasive, Chairside Oral Cavity Cancer Risk Assessment Prototype Using Machine Learning Approach

Oral cavity cancer (OCC) is associated with high morbidity and mortality rates when diagnosed at late stages. Early detection of increased risk provides an opportunity for implementing prevention strategies surrounding modifiable risk factors and screening to promote early detection and intervention. Historical evidence identified a gap in the training of primary care providers (PCPs) surrounding the examination of the oral cavity. The absence of clinically applicable analytical tools to identify patients with high-risk OCC phenotypes at point-of-care (POC) causes missed opportunities for implementing patient-specific interventional strategies. This study developed an OCC risk assessment tool prototype by applying machine learning (ML) approaches to a rich retrospectively collected data set abstracted from a clinical enterprise data warehouse. We compared the performance of six ML classifiers by applying the 10-fold cross-validation approach. Accuracy, recall, precision, specificity, area under the receiver operating characteristic curve, and recall–precision curves for the derived voting algorithm were: 78%, 64%, 88%, 92%, 0.83, and 0.81, respectively. The performance of two classifiers, multilayer perceptron and AdaBoost, closely mirrored the voting algorithm. Integration of the OCC risk assessment tool developed by clinical informatics application into an electronic health record as a clinical decision support tool can assist PCPs in targeting at-risk patients for personalized interventional care.

60 APPLIED LIFE SCIENCES↗

Expanded analysis of machine learning models for nuclear transient identification using TPOT

Industries around the world are becoming more and more data driven. The nuclear field is no exception with several different applications being proposed. One popular area of research is the use of machine learning in transient detection. This paper seeks to build upon a previous study which made use of the AutoML package TPOT to train traditional machine learning models to classify transient events occurring with a reactor. Synthetic data was once again collected using a GPWR reactor simulator. Data on 12 different events was collected using 15 different initial conditions. Here, a dataset consisting of over 100,000 data points was compiled and used to train 7 different machine learning models using a pre-defined TPOT dictionary with 12 different preprocessing techniques. Three of the trained models were able to produce validation results in the 90s with the expanded dataset. Once the models were trained, it was possible to look into where during the simulation, misclassifications occurred. Using these three models, analysis was done to determine if TPOT could be used to train models that were effective if important features were missing. The results from this were positive with the newly trained models scoring close to the original models. Finally, to conclude this study, the three high performing models were retrained using different random states to see if there was any major variation when different states were used.

42 ENGINEERING↗

Learning protein fitness models from evolutionary and assay-labeled data

Machine learning-based models of protein fitness typically learn from either unlabeled, evolutionarily related sequences or variant sequences with experimentally measured labels. For regimes where only limited experimental data are available, recent work has suggested methods for combining both sources of information. Toward that goal, we propose a simple combination approach that is competitive with, and on average outperforms more sophisticated methods. Our approach uses ridge regression on site-specific amino acid features combined with one probability density feature from modeling the evolutionary data. Within this approach, we find that a variational autoencoder-based probability density model showed the best overall performance, although any evolutionary density model can be used. Moreover, our analysis highlights the importance of systematic evaluations and sufficient baselines.

59 BASIC BIOLOGICAL SCIENCES↗

DIPS-Plus: The enhanced database of interacting protein structures for interface prediction

Abstract In this work, we expand on a dataset recently introduced for protein interface prediction (PIP), the Database of Interacting Protein Structures (DIPS), to present DIPS-Plus, an enhanced, feature-rich dataset of 42,112 complexes for machine learning of protein interfaces. While the original DIPS dataset contains only the Cartesian coordinates for atoms contained in the protein complex along with their types, DIPS-Plus contains multiple residue-level features including surface proximities, half-sphere amino acid compositions, and new profile hidden Markov model (HMM)-based sequence features for each amino acid, providing researchers a curated feature bank for training protein interface prediction methods. We demonstrate through rigorous benchmarks that training an existing state-of-the-art (SOTA) model for PIP on DIPS-Plus yields new SOTA results, surpassing the performance of some of the latest models trained on residue-level and atom-level encodings of protein complexes to date.

59 BASIC BIOLOGICAL SCIENCES↗

Understanding, discovery, and synthesis of 2D materials enabled by machine learning

Machine learning (ML) is becoming an effective tool for studying 2D materials. Taking as input computed or experimental materials data, ML algorithms predict the structural, electronic, mechanical, and chemical properties of 2D materials that have yet to be discovered. Such predictions expand investigations on how to synthesize 2D materials and use them in various applications, as well as greatly reduce the time and cost to discover and understand 2D materials. This tutorial review focuses on the understanding, discovery, and synthesis of 2D materials enabled by or benefiting from various ML techniques. Here, we introduce the most recent efforts to adopt ML in various fields of study regarding 2D materials and provide an outlook for future research opportunities. The adoption of ML is anticipated to accelerate and transform the study of 2D materials and their heterostructures.

2D Materials↗

Reducing Southern Ocean Shortwave Radiation Errors in the ERA5 Reanalysis with Machine Learning and 25 Years of Surface Observations

Earth system models struggle to simulate clouds and their radiative effects over the Southern Ocean, partly due to a lack of measurements and targeted cloud microphysics knowledge. We have evaluated biases of downwelling shortwave radiation in the ERA5 climate reanalysis using 25 years (1995–2019) of summertime surface measurements, collected on the Research and Supply Vessel (RSV) Aurora Australis, the Research Vessel (R/V) Investigator, and at Macquarie Island. During October–March daylight hours, the ERA5 simulation of SW down exhibited large errors (mean bias = 54 W m -2 , mean absolute error = 82 W m -2 , root-mean-square error = 132 W m -2 , and R 2 = 0.71). To determine whether we could improve these statistics, we bypassed ERA5’s radiative transfer model for SW down with machine learning–based models using a number of ERA5’s gridscale meteorological variables as predictors. These models were trained and tested with the surface measurements of SW down using a 10-fold shuffle split. An extreme gradient boosting (XGBoost) and a random forest–based model setup had the best performance relative to ERA5, both with a near complete reduction of the mean bias error, a decrease in the mean absolute error and root-mean-square error by 25% ± 3%, and an increase in the R 2 value of 5% ± 1% over the 10 splits. Large improvements occurred at higher latitudes and cyclone cold sectors, where ERA5 performed most poorly. We further interpret our methods using Shapley additive explanations. Our results indicate that data-driven techniques could have an important role in simulating surface radiation fluxes and in improving reanalysis products.

54 ENVIRONMENTAL SCIENCES↗