Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “computing methodologies → machine learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Exploring Structure-Sensitive Relations for Small Species Adsorption Using Machine Learning

Accurate prediction of adsorption energies on heterogeneous catalyst surfaces is crucial to predicting reactivity and screening materials. Adsorption linear scaling relations have been developed extensively but often lack accuracy and apply to one adsorbate and a single binding site type at a time. These facts undermine their ability to predict structure sensitivity and optimal catalyst structure. Using machine learning on nearly 300 density functional theory calculations, we demonstrate that generalized coordination number scaling relations hold well for oxygen- and high-valency carbon-binding species but fail for others. Here we reveal that the valency and the electronic coupling of a species with the surface, along with the site type and its coordination environment, are critical for small species adsorption. The model simultaneously predicts the adsorption energy and preferred site and significantly outperforms linear scalings in accuracy. It can expose the structure sensitivity of chemical reactions and enable enhanced catalyst activity via engineering particle shape and facet defects. The generality of our methodology is validated by training the model with transition metal data and transferring it to predict adsorption energies on single-atom alloys.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Machine-learning-accelerated multimodal characterization and multiobjective design optimization of natural porous materials

Natural porous materials such as nanoporous clays are used as green and low-cost adsorbents and catalysts. The key factors determining their performance in these applications are the pore morphology and surface activity, which are typically represented by properties such as specific surface area, pore volume, micropore content and pH. The latter may be modified and tuned to specific applications through material processing and/or chemical treatment. Characterization of the material, raw or processed, is typically performed experimentally, which can become costly especially in the context of tuning of the properties towards specific application requirements and needing numerous experiments. In this work, we present an application of tree-based machine learning methods trained on experimental datasets to accelerate the characterization of natural porous materials. The resulting models allow reliable prediction of the outcomes of experimental characterization of processed materials (R2 from 0.78 to 0.99) as well as identification of key factors contributing to those properties through feature importance analysis. Furthermore, the high throughput of the models enables exploration of processing parameter–property correlations and multiobjective optimization of prototype materials towards specific applications. We have applied these methodologies to pinpoint and rationalize optimal processing conditions for clays exploitable in acid catalysis. One of such identified materials was synthesized and tested revealing appreciable acid character improvement with respect to the pristine material. Specifically, it achieved 79% removal of chlorophyll-a in acid catalyzed degradation.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Evolving Multi-hazard Machine Learning Modeling for Advanced Risk-Informed Infrastructure Resilience Assessment

The socioeconomic impacts of pipeline incidents have escalated over the past three decades, revealing the limitation of traditional risk modeling methods when applied to extensive pipeline networks. This research aims to develop machine learning (ML) models that effectively identify, rank, and predict the diverse hazards and socioeconomic consequences associated with pipeline incidents. Utilizing historical data on pipeline incidents alongside weather and oceanographic data from the 1980s onward, the Houston metropolitan area serves as a testbed for the proposed methodologies. The research segments the combined datasets into three consecutive periods, demonstrating the efficacy of the updated model in predicting future events, particularly concerning precipitation rate data. Despite the challenges posed by a relatively limited dataset, local-level ML modeling offers valuable insights into the spatial and temporal dynamics of multiple hazards that contribute to pipeline incidents. These findings hold significant implications for future research, particularly in understanding and mitigating risks in various locations across the Gulf Coast and other coastal regions.

42 ENGINEERING↗

Artificial Intelligence Medical Support for Long-Duration Space Missions

We envision an artificial intelligence (AI) based system that will provide support and recommendations to the crew medical officer (CMO) and ground flight surgeon during long-duration space missions. Such a system would be pretrained on the knowledgebase of clinical knowledge on Earth, minimizing the amount of Earth data that needs to be transferred into space. Then during deployment, the system would be constantly refined through active learning from diverse streams of data from sensors in the spacecraft, data collected daily from individual astronauts, and human-in-the-loop feedback from the crew. The model could be interrogated for predictions and recommendations on personalized crew health based on the overall status of the spacecraft, medicinal stores, and status of other crew members. Adaptation techniques would be used to incorporate spaceflight data that have very different distributions from the training data due to the extreme environment. Edge computing and the most advanced neuromorphic processing would enable computation in scenarios with low power and bandwidth, while dimensionality reduction would be employed to ensure that the input data streams from spaceflight are as small as possible. In order to realize this long-term vision, several hardware and software aspects need to be developed and assembled. First, models pretrained on Earth biomedical data would need to be evaluated for predictive accuracy, and the best one selected. That model would need to be adapted to learn from diverse, sparse, and inconsistently measured data streams, as well as human-in-the-loop feedback. A data integration, standardization, and dimensionality reduction methodology would need to be developed to handle all data types and feed them into the model. Once the software and data infrastructure is developed, it would need to be integrated with small footprint compute processors and tested in high-radiation, high-vibration, unregulated temperature situations. As a short-term goal, we recommend to focus on the development of the data and model software structure. Several large language models (LLM) already exist that have been trained on Earth biomedical and clinical knowledgebases, including BioMedLLM, Med-PaLM, SPOKE LLM, and Foresight. These models need to be evaluated for accuracy and the best one chosen for a proof-of-concept structure, while maintaining awareness of the accelerating AI field and incorporating any newly improved model architectures as needed. Then, we recommend to develop a database of synthetic data types to mimic the diverse data streams that are expected in a long-duration space mission. This should include environmental and microbial data from the spacecraft, non-invasive data from wearables and point-of-care devices employed by astronauts, and more invasive molecular and physiological monitoring of clinical and biomarker data from astronauts. The data standardization methodology should be developed, and these data streams used to refine the clinical LLM. Several scenarios should be developed that could plausibly come up in a long-duration space mission, and changes or aberrations introduced to the data at specific times to mimic these scenarios. Then, question and answer tasks should be designed to interrogate the model for predictions and recommendations, with acceptable answers already identified.

Artificial Intelligence↗

K-means-driven Gaussian Process data collection for angle-resolved photoemission spectroscopy

Abstract We propose the combination of k-means clustering with Gaussian Process (GP) regression in the analysis and exploration of 4D angle-resolved photoemission spectroscopy (ARPES) data. Using cluster labels as the driving metric on which the GP is trained, this method allows us to reconstruct the experimental phase diagram from as low as 12% of the original dataset size. In addition to the phase diagram, the GP is able to reconstruct spectra in energy-momentum space from this minimal set of data points. These findings suggest that this methodology can be used to improve the efficiency of ARPES data collection strategies for unknown samples. The practical feasibility of implementing this technology at a synchrotron beamline and the overall efficiency implications of this method are discussed with a view on enabling the collection of more samples or rapid identification of regions of interest.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Automated Development of Molten Salt Machine Learning Potentials: Application to LiCl

The in silico modeling of molten salts is critical for emerging "carbon-free" energy applications but is inhibited by the cost of quantum mechanically treating the high polarizabilities of molten salts. Here, we integrate configurational sampling using classical force fields with active learning to automate and accelerate the generation of Gaussian approximation potentials (GAP) for molten salts. This methodology reduces the number of expensive ab initio evaluations required for training set generation to O(100), enabling the facile parametrization of a molten LiCl GAP model that exhibits a 19 000-fold speedup relative to AIMD. The developed molten LiCl GAP model is applied to sample extended spatiotemporal scales, permitting new physical insights into molten LiCl's coordination structure as well as experimentally validated predictions of structures, densities, self-diffusion constants, and ionic conductivities. The developed methodology significantly lowers the barrier to the in silico understanding and design of molten salts across the periodic table.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Wave Detection and Tracking Within a Rotating Detonation Engine Through Object Detection

As the operational time window of experimental rotating detonation engines (RDEs) is expanded and the technology matures toward integration within gas turbines, monitoring techniques must evolve to offer computationally efficient and highly time-resolved diagnostics. In this study, computer vision object detection methodology that seeks to reduce data processing time and calculate wave velocity within drastically reduced time intervals as compared to traditional high-frame-rate RDE images analysis techniques is proposed. The adapted you-only-look-once object detection network is trained to detect individual detonation waves within single down-axis RDE images. The wave location and rotational direction detected within a frame are tracked through a series of high-speed images to calculate the frame-to-frame wave velocity with the time-step resolution of $\mathrm{20 μs}$ across a series of frames. The analysis of the annotation box size and image linearization effects is presented, demonstrating the lowest frame-to-frame velocity total uncertainty of $\mathrm{±3.8\%}$ and the highest classification speed of 9.5 frames per second using linearized images. Linearized images “unwrap” the RDE annulus pixel region to a reduced image size. Here, this new method offers great reductions in data processing times and unsteady detonation behavior insight at intervals more comparable to the timescales of detonation wave interactions via the application of machine learning to experimental RDE data.

33 ADVANCED PROPULSION SYSTEMS↗

Machine learning reduces soft costs for residential solar photovoltaics

Further deployment of rooftop solar photovoltaics (PV) hinges on the reduction of soft (non-hardware) costs—now larger and more resistant to reductions than hardware costs. The largest portion of these soft costs is the expenses solar companies incur to acquire new customers. In this study, we demonstrate the value of a shift from significance-based methodologies to prediction-oriented models to better identify PV adopters and reduce soft costs. We employ machine learning to predict PV adopters and non-adopters, and compare its prediction performance with logistic regression, the dominant significance-based method in technology adoption studies. Our results show that machine learning substantially enhances adoption prediction performance: The true positive rate of predicting adopters increased from 66 to 87%, and the true negative rate of predicting non-adopters increased from 75 to 88%. We attribute the enhanced performance to complex variable interactions and nonlinear effects incorporated by machine learning. With more accurate predictions, machine learning is able to reduce customer acquisition costs by 15% ($0.07/Watt) and identify new market opportunities for solar companies to expand and diversify their customer bases. Our research methods and findings provide broader implications for the adoption of similar clean energy technologies and related policy challenges such as market growth and energy inequality.

14 SOLAR ENERGY↗

Modeling the Swift Bat Trigger Algorithm with Machine Learning

To draw inferences about gamma-ray burst (GRB) source populations based on Swift observations, it is essential to understand the detection efficiency of the Swift burst alert telescope (BAT). This study considers the problem of modeling the Swift / BAT triggering algorithm for long GRBs, a computationally expensive procedure, and models it using machine learning algorithms. A large sample of simulated GRBs from Lien et al. is used to train various models: random forests, boosted decision trees (with AdaBoost), support vector machines, and artificial neural networks. The best models have accuracies of greater than or equal to 97 percent (less than or equal to 3 percent error), which is a significant improvement on a cut in GRB flux, which has an accuracy of 89.6 percent (10.4 percent error). These models are then used to measure the detection efficiency of Swift as a function of redshift z, which is used to perform Bayesian parameter estimation on the GRB rate distribution. We find a local GRB rate density of n (sub 0) approaching 0.48 (sup plus 0.41) (sub minus 0.23) per cubic gigaparsecs per year with power-law indices of n (sub 1) approaching 1.7 (sup plus 0.6) (sub minus 0.5) and n (sub 2) approaching minus 5.9 (sup plus 5.7) (sub minus 0.1) for GRBs above and below a break point of z (redshift) (sub 1) approaching 6.8 (sup plus 2.8) (sub minus 3.2). This methodology is able to improve upon earlier studies by more accurately modeling Swift detection and using this for fully Bayesian model fitting.

gamma-ray burst: general – gamma-rays: general â↗

Plant Single-Cell Solutions for Energy and the Environment (Second Workshop Report)

Plants are important sources of energy and materials, and they collectively represent a critical component of Earth’s ecosystem. With increasing environmental stresses due to climate change and intensive agricultural practices, the need for resilient plants is greater than ever before. To secure plant resources for bioenergy, biomaterials, food, and ecosystem adaptation, a deeper understanding of the fundamental biology of plants at a cellular level is urgently needed. Plants contain a multitude of specialized cell types that compose tissues and organs. Pathogens often target specific cell types within plants, and the response of one cell to a particular stimulus is likely to be distinct from its neighbor because of underlying molecular and contextual differences. Understanding how these responses are distributed among cells, the main goal of single-cell approaches, will substantially enhance our ability to use targeted engineering for improving plant productivity and resilience. Furthermore, single-cell approaches are necessary to understand the interactions between plants and other ecosystem members such as fungi, bacteria, and archaea. Unlocking these gene-response mechanisms at a cellular level can improve our ability to adapt plants to environmental stresses, increasing their utility as feedstocks for biomaterials and bioenergy. Recent advances in high-throughput sequencing, mass spectrometry, microfluidics and miniaturization, artificial intelligence and machine learning, and bioinformatics have greatly improved our ability to detect and understand processes at a cellular level. In mammalian systems, single-cell transcriptomics has already led to many advances, such as newly identified cell types and cell-targeted treatment of diseases, and mass spectrometry-based single-cell proteomics has recently been demonstrated as a promising emerging technology. However, plant single-cell omics has lagged behind mammalian approaches due to the high cost of the technologies relative to available resources and to the innate biological features of plants, including the complexity of the cell wall and polyploidy. To better understand how single-cell methods could enable plant science, Lawrence Berkeley National Laboratory (Berkeley Lab) hosted a workshop on April 29, 2021, that brought together a diverse group of leaders in plant and/or single-cell biology. Attendees represented federal research programs and domestic and international academic institutions. During the workshop, three presenters described the current state of research in both experimental and computational approaches. While the focus of the workshop was on factors preventing plant biology researchers from fully adopting single-cell methodologies, workshop participants agreed that most barriers could be overcome with focused, strategic investment and coordinated efforts among institutions leading to significant scientific discoveries that would be difficult to obtain using more conventional technologies.

59 BASIC BIOLOGICAL SCIENCES↗

Toward ultra-efficient high-fidelity predictions of wind turbine wakes: Augmenting the accuracy of engineering models with machine learning

This study proposes a novel machine learning (ML) methodology for the efficient and cost-effective prediction of high-fidelity three-dimensional velocity fields in the wake of utility-scale turbines. The model consists of an autoencoder convolutional neural network with U-Net skipped connections, fine-tuned using high-fidelity data from large-eddy simulations (LES). The trained model takes the low-fidelity velocity field cost-effectively generated from the analytical engineering wake model as input and produces the high-fidelity velocity fields. The accuracy of the proposed ML model is demonstrated in a utility-scale wind farm for which datasets of wake flow fields were previously generated using LES under various wind speeds, wind directions, and yaw angles. Comparing the ML model results with those of LES, the ML model was shown to reduce the error in the prediction from 20% obtained from the Gauss Curl hybrid (GCH) model to less than 5%. In addition, the ML model captured the non-symmetric wake deflection observed for opposing yaw angles for wake steering cases, demonstrating a greater accuracy than the GCH model. The computational cost of the ML model is on par with that of the analytical wake model while generating numerical outcomes nearly as accurate as those of the high-fidelity LES.

Mechanics↗

VAIM-CFF: a variational autoencoder inverse mapper solution to Compton form factor extraction from deeply virtual exclusive reactions

We develop a new methodology for extracting Compton form factors (CFFs) from deeply virtual exclusive reactions such as the unpolarized DVCS cross section using a specialized inverse problem solver, a variational autoencoder inverse mapper (VAIM). The VAIM-CFF framework not only allows us access to a fitted solution set possibly containing multiple solutions in the extraction of all 8 CFFs from a single cross section measurement, but also accesses the lost information contained in the forward mapping from CFFs to cross section. We investigate various assumptions and their effects on the predicted CFFs such as cross section organization, number of extracted CFFs, use of uncertainty quantification technique, and inclusion of prior physics information. We then use dimensionality reduction techniques such as principal component analysis to visualize the missing physics information tracked in the latent space of the VAIM framework. Through re-framing the extraction of CFFs as an inverse problem, we gain access to fundamental properties of the problem not comprehensible in standard fitting methodologies: exploring the limits of the information encoded in deeply virtual exclusive experiments.

Accelerator Physics↗

Maximizing machine learning interatomic potential transferability for the discovery of the novel stellated octadecagon Bi18-Pt24 cage structure

Achieving true transferability remains the central challenge for Machine Learning Interatomic Potentials (ML-IAPs) in modeling complex bimetallic nanoclusters across their vast potential energy surfaces. We systematically investigate data selection strategies to optimize the Chebyshev Interaction Model for Efficient Simulation (ChIMES) potential for the Bi-Pt nanoclusters by comparing three innovative sampling methods: Principal Component Analysis (PCA)/k-means (structural diversity), t-distributedStochasticNeighborEmbedding (t-SNE)/k-means (force-space diversity), and hierarchical clustering. Quantitatively, the PCA/k-means strategy proved most effective for global accuracy, yielding the lowest force errors and achieving energy root mean square errors (RMSE) values competitive with Density Functional Theory (DFT), demonstrating excellent accuracy (19.16meV/atom). Structural validation on 34 unique DFT-optimized isomers further confirmed the potential’s high fidelity, with the best model PCA/k-means reproducing structures with an average root mean square deviation (RMSD) of 0.10 Å. However, the t-SNE methods, by maximizing diversity in the force space, demonstrated superior extrapolative power, leading to the more precise prediction of a novel stellated octadecagon Bi18⁢Pt24 cage structure, demonstrating the potential for exploring previously unseen morphologies. Our results establish a clear methodology for strategic data sampling that successfully maximizes ML-IAP transferability, providing an accurate and computationally efficient tool that accelerates the theoretical discovery of complex bimetallic architectures.

Vangheluwe, Raphaël [Université Paris-Saclay, CNRS↗

Inference offers a metric to constrain dynamical models of neutrino flavor transformation

The multimessenger astrophysics of compact objects presents a vast range of environments where neutrino flavor transformation may occur and may be important for nucleosynthesis, dynamics, and a detected neutrino signal. Development of efficient techniques for surveying flavor evolution solution spaces in these diverse environments, which augment and complement existing sophisticated computational tools, could leverage progress in this field. To this end we continue our exploration of statistical data assimilation (SDA) to identify solutions to a small-scale model of neutrino flavor transformation. SDA is a machine learning formula wherein a dynamical model is assumed to generate any measured quantities. Specifically, we use an optimization formulation of SDA wherein a cost function is extremized via the variational method. Regions of state space in which the extremization identifies the global minimum of the cost function will correspond to parameter regimes in which a model solution can exist. Our example study seeks to infer the flavor transformation histories of two monoenergetic neutrino beams coherently interacting with each other and with a matter background. We require that the solution be consistent with measured neutrino flavor fluxes at the point of detection, and with constraints placed upon the flavor content at various locations along their trajectories, such as the point of emission, and the locations of the Mikheyev-Smirnov-Wolfenstein resonances. We show how the procedure efficiently identifies solution regimes and rules out regimes where solutions are infeasible. Overall, results in this work intimate the promise of this “variational annealing” methodology to efficiently probe an array of fundamental questions that traditional numerical simulation codes render difficult to access.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Fast Characterization of Inducible Regions of Atrial Fibrillation Models With Multi-Fidelity Gaussian Process Classification

Computational models of atrial fibrillation have successfully been used to predict optimal ablation sites. A critical step to assess the effect of an ablation pattern is to pace the model from different, potentially random, locations to determine whether arrhythmias can be induced in the atria. In this work, we propose to use multi-fidelity Gaussian process classification on Riemannian manifolds to efficiently determine the regions in the atria where arrhythmias are inducible. We build a probabilistic classifier that operates directly on the atrial surface. We take advantage of lower resolution models to explore the atrial surface and combine seamlessly with high-resolution models to identify regions of inducibility. We test our methodology in 9 different cases, with different levels of fibrosis and ablation treatments, totalling 1,800 high resolution and 900 low resolution simulations of atrial fibrillation. When trained with 40 samples, our multi-fidelity classifier that combines low and high resolution models, shows a balanced accuracy that is, on average, 5.7% higher than a nearest neighbor classifier. We hope that this new technique will allow faster and more precise clinical applications of computational models for atrial fibrillation. All data and code accompanying this manuscript will be made publicly available at: https://github.com/fsahli/AtrialMFclass.

59 BASIC BIOLOGICAL SCIENCES↗

Flexible AI Models for Grid Resilience

The rapid growth in size and complexity of artificial intelligence (AI) and machine learning (ML) models has led to increased energy demands, posing a threat to the reliability of the existing power grid. This project addresses the challenge of highly intermittent and energy-intensive inference workloads by (1) developing fidelity-adaptive neural networks capable of dynamic response to grid conditions and (2) integrating these networks with power flow simulations to assess their impact on power grid reliability. We will explore both top-down and bottom-up approaches to create hierarchies of submodels that provide a controlled trade-off between power draw and prediction accuracy. The top-down method utilizes NN pruning to reduce a flagship model into progressively smaller, energy-efficient variants. The bottom-up approach employs geometrically principled weight setting strategies to construct depth-efficient models from the ground up. A real-time hardware-in-the-loop (HIL) platform will be developed to simulate a scaled AC power grid, integrating live AI workload power draw and enabling dynamic model switching in response to grid feedback. This work will provide a novel framework for evaluating the impact of flexible AI/ML workloads on grid performance and establish new methodologies for energy-aware computing in data centers. The outcomes will demonstrate that adaptive AI/ML can play a critical role in improving grid stability while advancing NREL's leadership in energy-efficient computing research.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Reduced-basis method for few-body bound-state emulation

Recent advances in both theoretical and computational methods have enabled large-scale, precision calculations of the properties of atomic nuclei. With the growing complexity of modern nuclear theory, however, also comes the need for novel methods to perform systematic studies and quantify the uncertainties of models when confronted with experimental data. Here, this study presents an application of such an approach, the reduced basis method, to substantially lower computational costs by constructing a significantly smaller Hamiltonian subspace informed by previous solutions. Our method shows comparable efficiency and accuracy to other dimensionality reduction techniques on an artificial three-body bound system while providing a richer representation of physical information in its projection and training subspace. This methodological advancement can be applied in other contexts and has the potential to greatly improve our ability to systematically explore theoretical models and thus enhance our understanding of the fundamental properties of nuclear systems.

cluster models↗

A machine learning approach to emulation and biophysical parameter estimation with the Community Land Model, version 5

Abstract. Land models are essential tools for understanding and predicting terrestrial processes and climate–carbon feedbacks in the Earth system, but uncertainties in their future projections are poorly understood. Improvements in physical process realism and the representation of human influence arguably make models more comparable to reality but also increase the degrees of freedom in model configuration, leading to increased parametric uncertainty in projections. In this work we design and implement a machine learning approach to globally calibrate a subset of the parameters of the Community Land Model, version 5 (CLM5) to observations of carbon and water fluxes. We focus on parameters controlling biophysical features such as surface energy balance, hydrology, and carbon uptake. We first use parameter sensitivity simulations and a combination of objective metrics including ranked global mean sensitivity to multiple output variables and non-overlapping spatial pattern responses between parameters to narrow the parameter space and determine a subset of important CLM5 biophysical parameters for further analysis. Using a perturbed parameter ensemble, we then train a series of artificial feed-forward neural networks to emulate CLM5 output given parameter values as input. We use annual mean globally aggregated spatial variability in carbon and water fluxes as our emulation and calibration targets. Validation and out-of-sample tests are used to assess the predictive skill of the networks, and we utilize permutation feature importance and partial dependence methods to better interpret the results. The trained networks are then used to estimate global optimal parameter values with greater computational efficiency than achieved by hand tuning efforts and increased spatial scale relative to previous studies optimizing at a single site. By developing this methodology, our framework can help quantify the contribution of parameter uncertainty to overall uncertainty in land model projections.

54 ENVIRONMENTAL SCIENCES↗