Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “machine learning for science”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20

Large language models generate functional protein sequences across diverse families

Deep-learning language models have shown promise in various biotechnological applications, including protein design and engineering. Here, in this paper, we describe ProGen, a language model that can generate protein sequences with a predictable function across large protein families, akin to generating grammatically and semantically correct natural language sentences on diverse topics. The model was trained on 280 million protein sequences from >19,000 families and is augmented with control tags specifying protein properties. ProGen can be further fine-tuned to curated sequences and tags to improve controllable generation performance of proteins from families with sufficient homologous samples. Artificial proteins fine-tuned to five distinct lysozyme families showed similar catalytic efficiencies as natural lysozymes, with sequence identity to natural proteins as low as 31.4%. ProGen is readily adapted to diverse protein families, as we demonstrate with chorismate mutase and malate dehydrogenase.

59 BASIC BIOLOGICAL SCIENCES↗

Discovery of widespread transcription initiation at microsatellites predictable by sequence-based deep neural network

Using the Cap Analysis of Gene Expression (CAGE) technology, the FANTOM5 consortium provided one of the most comprehensive maps of transcription start sites (TSSs) in several species. Strikingly, ~72% of them could not be assigned to a specific gene and initiate at unconventional regions, outside promoters or enhancers. Here, we probe these unassigned TSSs and show that, in all species studied, a significant fraction of CAGE peaks initiate at microsatellites, also called short tandem repeats (STRs). To confirm this transcription, we develop Cap Trap RNA-seq, a technology which combines cap trapping and long read MinION sequencing. We train sequence-based deep learning models able to predict CAGE signal at STRs with high accuracy. These models unveil the importance of STR surrounding sequences not only to distinguish STR classes, but also to predict the level of transcription initiation. Importantly, genetic variants linked to human diseases are preferentially found at STRs with high transcription initiation level, supporting the biological and clinical relevance of transcription initiation at STRs. Together, our results extend the repertoire of non-coding transcription associated with DNA tandem repeats and complexify STR polymorphism.

59 BASIC BIOLOGICAL SCIENCES↗

Atikokan Digital Twin: Machine learning in a biomass energy system

The Atikokan Generating Station, operated by Ontario Power Generation, has a 200 MW, biomass-fired tower boiler that operates on a dispatch schedule with a five-minute cycle. The boiler is generally operated in the range of 40–100 MW using two of five burner levels. In order to optimize boiler performance, we propose the implementation of a unique digital twin. Our digital twin abstraction couples Bayesian inference from science-based models and from observations (machine learning) with decision theory to predict operating-variable set points that optimize the physical asset (the boiler) in the presence of uncertainty (artificial intelligence). We focus this paper on the continuous Bayesian machine learning part of the Atikokan Digital Twin; we discuss decision theory in a companion paper. We identify and learn about 12 operational, model, and measured-output parameters and their uncertainties from high-fidelity, science-based simulations of the Atikokan boiler and from the observed measurements at the power plant. Since the goal of the Atikokan Digital Twin is to implement it online in real time, we require fast function evaluations for the quantities of interest extracted from the simulations in the Bayesian analysis. We use Gaussian process regression/interpolation to create accurate, robust surrogate models. We define the Bayesian priors and likelihood function and solve for the posterior distributions of the 12 parameters. Here we then propagate these distributions (i.e., parameters with uncertainty) into the predicted distributions of 790 quantities of interest to learn about the relative importance of various sources of error including experimental, model, and operating-parameter errors.

09 BIOMASS FUELS↗

Materials Genomics Search for Possible Helium‐Absorbing Nano‐Phases in Fusion Structural Materials

Abstract Civilian fusion demands structural materials that can withstand the harsh environments imposed inside fusion plasma reactors. The structural materials often transmute under 14.1 MeV fast neutrons, producing helium (He), which embrittles the grain boundary (GB) network. Here, it is shown that neutron‐friendly and mechanically strong nano‐phases with atomic‐scale free volume can have low He‐embedding energy and >10 at.% He‐absorbing capacity, and can be especially advantageous for soaking up He on top of resisting radiation damage and creep, provided they have thermodynamic compatibility with the matrix phase, satisfactory equilibrium wetting angle, as well as a high enough melting point. The preliminary experimental demonstration proves that is a good ab initio predictor of He shielding potency in nano‐heterophase materials, and thus, is used as a key feature for computational screening. In this context, a list of viable compounds expected to be good He‐absorbing nano‐phases is presented, taking into account , the neutron absorption and activation cross‐sections, the elastic moduli, melting temperature, the thermodynamic compatibility, and the equilbrium wetting angle of the nano‐phases with the Fe matrix as an example.

36 MATERIALS SCIENCE↗

High-Throughput Analysis of Materials for Chemical Looping Processes

Chemical looping is a promising approach for improving the energy efficiency of many industrial chemical processes. However, a major limitation of modern chemical looping technologies is the lack of suitable active materials to mediate the involved subreactions. Identification of suitable materials has been historically limited by the scarcity of high-temperature (>600 °C) thermochemical data to evaluate candidate materials. An accuratethermodynamic approach is demonstrated here to rapidly identify active materials which is applicable to a wide variety of chemical looping chemistries. Application of this analysis to chemical looping combustion correctly classifies 17/17 experimentally studied redox materials by their viability and identifies over 1300 promising yet previously unstudied active materials. This approach is further demonstrated by analyzing redox pairs for mediating a novel chemical looping process for producing pure SO2 from raw sulfur and air which could provide a more efficient and lower emission route to sulfuric acid. 12 promising redox materials for this process are identified, two of which are supported by previous experimental studies of their individual oxidation and reduction reactions. This approach provides the necessary foundation for connecting process design with high-throughput material discovery to accelerate the innovation and development of a wide range of chemical looping technologies.

chemical looping↗

In-situ sensor monitoring of multi-class gas porosity formation in laser powder bed fusion using convolutional neural network

In-situ monitoring of defect formation remains a significant challenge in the laser powder bed fusion (LPBF) process. Recent advances have enabled real-time defect detection with machine learning and in-situ sensing technologies; however, most studies focus on binary classification of keyhole pores, limiting nuanced multi-class pore differentiation and formation mechanisms. This work introduces a multi-class pore detection framework (no pore, small pores < 15 µm, and large pores > 15 µm) by leveraging photodiode sensor data alongside high-fidelity synchrotron X-ray imaging. The 15 µm threshold is selected to distinguish between two fundamentally different defect mechanisms, following the physical size-mechanism boundary established by prior high-resolution synchrotron X-ray characterization of Al6061 LPBF. Distinguishing these classes is critical because large keyhole pores are structurally detrimental, whereas small gas pores are often benign, requiring different process control strategies. Thermal emission monitoring data collected simultaneously with high-speed X-ray imaging at the Stanford Synchrotron Radiation Lightsource (SSRL), are correlated with subsurface melt pool dynamics to establish ground truth. Continuous Wavelet Transform (CWT) with optimized parameters converts the photodiode time-series signals into time–frequency images, facilitating feature extraction. Convolutional Neural Networks (CNN) are then applied for real-time multi-class pore classification in an average inference time of 1 ms per signal window. It achieves 79% accuracy and an Area Under the Receiver Operating Characteristic curve (AUC ROC) score of 0.89 with five-fold cross-validation. The results demonstrate that coupling CWT-based feature engineering with CNN architecture enables reliable multi-class pore detection in Al6061 builds using affordable in-situ sensors. This approach advances scalable and affordable quality assurance in additive manufacturing by moving beyond binary defect detection toward more nuanced classification of porosity mechanisms with in-situ sensors and machine learning.

Laser powder bed fusion, Multi-class pores, In-sit↗

Atomistic simulations and machine learning of solute grain boundary segregation in Mg alloys at finite temperatures

Understanding solute segregation thermodynamics is the first step in investigating grain boundary (GB) properties, such as strong yttrium (Y) effects on grain growth and texture evolution in micro-scale polycrystalline magnesium (Mg) alloys. To estimate the average GB segregation behavior in low-solute-concentration Mg alloys (e.g., 2 at.% Y), a state-of-the-art spectral approach is applied based on a per-site segregation energy spectrum for Y solute atoms at zero K obtained from molecular statistics (MS) simulations of ~10 4 GB sites in Mg symmetric tilt GBs (STGBs). Although selected MS simulation results are consistent with verification by density functional theory (DFT) calculations, estimates of average segregation tendency based on the zero-K energy spectrum deviate from experimental observations. To resolve this problem, thermodynamic integration (TI) methods based on molecular dynamics (MD) simulations are used to determine the per-site segregation free energies of Y at representative GB sites, which show contributions beyond harmonic approximations can be important for certain GB sites at high temperatures. A surrogate model of per-site segregation free energy is constructed from a small set of TI data points using stacking cross-validation regressors and physics-informed descriptors. This model is applied to predict the Y segregation free energy spectra for thousands of GB sites in Mg STGBs with uncertainty quantification. Finally, the average segregation tendency predicted by the spectral approach based on the free energy spectra agrees well (within the uncertainty range) with experimental observations for micro-scale polycrystalline Mg alloys at typical thermomechanical processing temperatures (500 ~ 800 K), where thermodynamic equilibrium states are likely to be achieved due to fast diffusion.

Atomistic simulations↗

Machine learning insights into microstructural origins of transport and mechanical properties in porous microstructures

Multifunctional porous materials are increasingly needed across various fields, but their complex microstructures create significant challenges due to the intricate microstructure-property relationships. This complexity, combined with limitations of traditional analysis methods, hinders efforts to understand and optimize microstructure–property relationships. Here, to address this, we integrate physics-based mesoscale modeling with interpretable machine learning (ML) to uncover how microstructural features govern effective diffusivity and elastic modulus. At constant porosity, we show diffusivity varies by over 150 × and modulus by ∼50 ×, highlighting the power of microstructure engineering. Statistical analysis reveals bimodal behavior in diffusivity and unimodal in modulus. ML identifies connectivity as the dominant factor, while modulus is also sensitive to domain size and feature interactions. Controlled simulations further highlight domain shape as a critical feature for modulus. This framework enables efficient exploration of microstructure-property correlations, offering new insights to guide the design of advanced porous materials.

Bicontinuous microstructure↗

Real-time tracking and analysis of gas bubble dynamics in laser powder bed fusion using in-situ X-ray characterization and machine learning

Porosity defects remain a significant challenge in the laser powder bed fusion (LPBF) process, adversely affecting the mechanical properties and reliability of additively manufactured components. Here, this study investigates the real-time formation and trajectory of gas bubbles during LPBF of Al6061 alloy using advanced in-situ X-ray characterization and machine learning. The unsupervised Gaussian mixture model and particle tracking algorithm developed are able to precisely track and quantify the properties of gas bubbles and keyhole pores. Our analysis identified five distinct types of gas bubble formation and movement patterns, emphasizing the diverse origins and behaviors of these defects. It enables precise quantification of trajectories, velocities, and morphological changes of gas bubbles, offering a granular view of the subsurface dynamics within the melt pool. Additionally, we explored keyhole-induced pore dynamics, revealing the critical role of keyhole oscillation and collapse for the formation of both large and small gas pores. It defines four different regions of gas bubble movement within the melt pool, providing a clearer understanding of how local fluid dynamics affect pore behavior. The results underscore the importance of integrating in-situ experimental observation and automated machine learning to develop a more robust predictive model for defect formation in LPBF.

In-situ X-ray imaging↗

Reinforced double-threaded slide-ring networks for accelerated hydrogel discovery and 3D printing

Traditionally, slide-ring gels are stretchable but soft as a result of an elasticity-stretchability trade-off. Herein, we introduce a new approach to breaking this trade-off and creating reinforced slide-ring networks with mobile crosslinkers. Our approach involves the construction of a polyethylene glycol double-threaded γ-cyclodextrin-based pro-slide-ring crosslinker that serves as a modular component for 3D printing and copolymerization. The resulting crystalline-domain-reinforced slide-ring hydrogels, or CrysDoS-gels, exhibit both high elasticity and high stretchability. The modular synthesis allows for high-throughput synthesis of CrysDoS-gels, generating a large amount of data for structure-property analysis. Here, by employing data science techniques, such as machine learning and linear regression, not only were we able to identify which chemical components influence the mechanical properties of CrysDoS-gels, but this analysis also aided in the discovery of better-performing CrysDoS-gels. Finally, we demonstrate the potential application of the newly discovered CrysDoS-gels as sensing devices by 3D printing them as stress sensors with high sensitivity and a broad detection range.

3D-printing↗

Evaluation of UAV-derived multimodal remote sensing data for biomass prediction and drought tolerance assessment in bioenergy sorghum

Screening for drought tolerance is critical to ensure high biomass production of bioenergy sorghum in arid or semi-arid environments. The bottleneck in drought tolerance selection is the challenge of accurately predicting biomass for a large number of genotypes. Although biomass prediction by low-altitude remote sensing has been widely investigated on various crops, the performance of the predictions are not consistent, especially when applied in a breeding context with hundreds of genotypes. In some cases, biomass prediction of a large group of genotypes benefited from multimodal remote sensing data; while in other cases, the benefits were not obvious. In this study, we evaluated the performance of single and multimodal data (thermal, RGB, and multispectral) derived from an unmanned aerial vehicle (UAV) for biomass prediction for drought tolerance assessments within a context of bioenergy sorghum breeding. The biomass of 360 sorghum genotypes grown under well-watered and water-stressed regimes was predicted with a series of UAV-derived canopy features, including canopy structure, spectral reflectance, and thermal radiation features. Biomass predictions using canopy features derived from the multimodal data showed comparable performance with the best results obtained with the single modal data with coefficients of determination (R 2 ) ranging from 0.40 to 0.53 under water-stressed environment and 0.11 to 0.35 under well-watered environment. The significance in biomass prediction was highest with multispectral followed by RGB and lowest with the thermal sensor. Finally, two well-recognized yield-based drought tolerance indices were calculated from ground truth biomass data and UAV predicted biomass, respectively. Results showed that the geometric mean productivity index outperformed the yield stability index in terms of the potential for reliable predictions by the remotely sensed data. Collectively, this study demonstrated a promising strategy for the use of different UAV-based imaging sensors to quantify yield-based drought tolerance.

09 BIOMASS FUELS↗

Accelerating the discovery of battery electrode materials through data mining and deep learning models

The availability of crystalline materials databases allows for building accurate machine learning (ML) models that can accelerate the exploration of materials chemical space for energy storage applications. In this work, we screen all inorganic materials included in the Materials Project and AFLOW databases as potential metal-ion battery electrodes. We develop an efficient protocol to mine and screen raw data in current databases and provide a new database of electrode materials by considering pairs of charged and discharged electrodes. This effort leads to a new database with over 190,000 instances, in contrast to the original battery database which contains about 5000. The expanded battery data set is then used to build regression-based deep neural network models for predicting average voltages and percentage volume changes upon charging and discharging, which present improvements of at least 28% for target properties with respect to previous models, and are now able to predict anode electrodes (low voltage region) as well as electrodes that will not work in electrochemical cells (negative voltages), overcoming the challenges identified in previous ML models for battery electrodes. Additionally, a further screening of the expanded database itself allowed us to identify 35 novel electrode candidates with excellent battery performance metrics.

25 ENERGY STORAGE↗

Austenitic parent grain reconstruction in martensitic steel using deep learning

In this work we develop a deep convolutional architecture to estimate the prior austenite structure from observed martensite electron backscatter diffraction micrographs. A novel data augmentation strategy randomizes the global reference coordinate system which makes it possible to train our model from only four micrographs. The model is much faster than algorithmic approaches and generalizes well when applied to micrographs of a different material. Empirical evidence suggests the efficacy of the model depends on the scale of the microstructure and receptive field of the vision model. Furthermore, this work demonstrates that modern computer vision approaches are well suited for capturing complex spatial-orientation patterns present in orientation imaging micrographs.

36 MATERIALS SCIENCE↗

Development of machine learning framework for interface force closures based on bubble tracking data

Interfacial force closures in the two-fluid model play a critical role for the predictive capabilities of void fraction distribution. However, the practices of interfacial force modeling have long been challenged by the inherent physical complexity of the two-phase flows. The rapidly expanding computational capabilities in the recent years have made high-fidelity data from the interface-captured direct numerical simulation become more available, and hence potential for data-driven interfacial force modeling has prevailed. In this work, we established a data-driven modeling framework integrated to the HZDR multiphase Eulerian-Eulerian framework for computational fluid dynamics simulations. The data-driven framework is verified in a benchmark problem, where a feedforward neural network managed to capture the non-linear mapping between bubble Reynolds number and drag coefficient and reproduce the void distribution resulting from the baseline model in the test case. The second focus is on utilizing the bubble tracking data set to form a closure for the bubble drag in the turbulent bubbly flow, in which the drag coefficient is set to be correlated with the bubble Reynolds number and the Eötvös number. Pseudo-steady state filtering in the Frenet Frame was carried out to obtain the drag coefficient from the turbulent bubbly flow data. The performance of the data-driven drag model is also examined through a case study, where improvement of model’s prediction near-wall is regarded necessary. In conclusion, discussion and further plans of investigation are provided.

42 ENGINEERING↗

Two-Dimensional Energy Histograms as Features for Machine Learning to Predict Adsorption in Diverse Nanoporous Materials

A major obstacle for machine learning (ML) in chemical science is the lack of physically informed feature representations that provide both accurate prediction and easy interpretability of the ML model. In this work, we describe adsorption systems using novel two-dimensional energy histogram (2D-EH) features, which are obtained from the probe-adsorbent energies and energy gradients at grid points located throughout the adsorbent. The 2D-EH features encode both energetic and structural information of the material and lead to highly accurate ML models (coefficient of determination R2 ~ 0.94–0.99) for predicting single-component adsorption capacity in metal–organic frameworks (MOFs). Here, we consider the adsorption of spherical molecules (Kr and Xe), linear alkanes with a wide range of aspect ratios (ethane, propane, n-butane, and n-hexane), and a branched alkane (2,2-dimethylbutane) over a wide range of temperatures and pressures. The interpretable 2D-EH features enable the ML model to learn the basic physics of adsorption in pores from the training data. We show that these MOF-data-trained ML models are transferrable to different families of amorphous nanoporous materials. We also identify several adsorption systems where capillary condensation occurs, and ML predictions are more challenging. Nevertheless, our 2D-EH features still outperform structural features including those derived from persistent homology. The novel 2D-EH features may help accelerate the discovery and design of advanced nanoporous materials using ML for gas storage and separation in the future.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Integrated analysis of X-ray diffraction patterns and pair distribution functions for machine-learned phase identification

Abstract To bolster the accuracy of existing methods for automated phase identification from X-ray diffraction (XRD) patterns, we introduce a machine learning approach that uses a dual representation whereby XRD patterns are augmented with simulated pair distribution functions (PDFs). A convolutional neural network is trained directly on XRD patterns calculated using physics-informed data augmentation, which accounts for experimental artifacts such as lattice strain and crystallographic texture. A second network is trained on PDFs generated via Fourier transform of the augmented XRD patterns. At inference, these networks classify unknown samples by aggregating their predictions in a confidence-weighted sum. We show that such an integrated approach to phase identification provides enhanced accuracy by leveraging the benefits of each model’s input representation. Whereas networks trained on XRD patterns provide a reciprocal space representation and can effectively distinguish large diffraction peaks in multi-phase samples, networks trained on PDFs provide a real space representation and perform better when peaks with low intensity become important. These findings underscore the importance of using diverse input representations for machine learning models in materials science and point to new avenues for automating multi-modal characterization.

36 MATERIALS SCIENCE↗

Machine Learned Hückel Theory: Interfacing Physics and Deep Neural Networks

The Hückel Hamiltonian is an incredibly simple tight-binding model known for its ability to capture qualitative physics phenomena arising from electron interactions in molecules and materials. Part of its simplicity arises from using only two types of empirically fit physics-motivated parameters: the first describes the orbital energies on each atom and the second describes electronic interactions and bonding between atoms. By replacing these empirical parameters with machine-learned dynamic values, we vastly increase the accuracy of the extended Hückel model. The dynamic values are generated with a deep neural network, which is trained to reproduce orbital energies and densities derived from density functional theory. The resulting model retains interpretability, while the deep neural network parameterization is smooth and accurate and reproduces insightful features of the original empirical parameterization. Altogether, this work shows the promise of utilizing machine learning to formulate simple, accurate, and dynamically parameterized physics models.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗