Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “computing methodologies → machine learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Hypothesis testing via AI: Generating physically interpretable models of scientific data with machine learning (Full Technical Report)

Deep learning has demonstrated an exceptional ability to solve complex tasks (an engineering success); however, it has done so at the expense of the ability to generate new knowledge (a scientific failure). We propose an alternative framework—entitled Deep Symbolic Regression (DSR)—in which artificial neural networks (NNs) rapidly generate hypotheses about physical relationships among inputs. This framework bypasses the need to interpret an NN altogether, while still leveraging the representational power of deep learning. The resulting models are tractable mathematical expressions, which are inherently and readily human interpretable and can provide insights into underlying physical phenomena. Further, we fold this methodology into the scientific process by allowing the scientist to directly integrate a priori knowledge and beliefs to accelerate learning. We demonstrate this methodology on symbolic regression—the problem of rediscovering underlying expressions describing a dataset—and achieve state-of-the-art performance across a wide variety of symbolic regression problems. Further, we generalize our DSR framework to apply to the more general class of symbolic optimization problems, in which one seeks to optimize a sequence of symbols or “tokens” under a black-box reward function. Examples of other symbolic optimization problems include neural architecture search and computational antibody design. Our generalized tool, Deep Symbolic Optimization (DSO), has been demonstrated on the task of learning symbolic control policies for reinforcement learning environments, and has been adopted as an enabling capability for computational antibody design.

97 MATHEMATICS AND COMPUTING↗

Thermal Management for FPGA Nodes in HPC Systems

The integration of FPGAs into large-scale computing systems is gaining attention. In these systems, real-time data handling for networking, tasks for scientific computing, and machine learning can be executed with customized datapaths on reconfigurable fabric within heterogeneous compute nodes. At the same time, thermal management, particularly battling the cooling cost and guaranteeing the reliability, is a continuing concern. The introduction of new heterogeneous components into HPC nodes only adds further complexities to thermal modeling and management. The thermal behavior of multi-FPGA systems deployed within large compute clusters is less explored. Here, we first show that the thermal behaviors of different FPGAs of the same generation can vary due to their physical locations in a rack and process variation, even though they are running the same tasks. We present a machine learning–based model to capture the thermal behavior of each individual FPGA in the cluster. We then propose two thermal management strategies guided by our thermal model. First, we mitigate thermal variation and hotspots across the cluster by proactive thermal-aware task placement. Under the tested system and benchmarks, we achieve up to 26.4° C and on average 13.3° C system temperature reduction with no performance penalty. Second, we utilize this thermal model to guide HLS parameter tuning at the task design stage to achieve improved thermal response after deployment.

97 MATHEMATICS AND COMPUTING↗

An investigation on machine learning predictive accuracy improvement and uncertainty reduction using VAE-based data augmentation

The confluence of ultrafast computers with large memory, rapid progress in Machine Learning (ML) algorithms, and the availability of large datasets place multiple engineering fields at the threshold of dramatic progress. However, a unique challenge in nuclear engineering is data scarcity because experimentation on nuclear systems is usually more expensive and time-consuming than most other disciplines. One potential way to resolve the data scarcity issue is deep generative learning, which uses certain ML models to learn the underlying distribution of existing data and generate synthetic samples that resemble the real data. In this way, one can significantly expand the dataset to train more accurate predictive ML models. In this study, our objective is to evaluate the effectiveness of data augmentation using variational autoencoder (VAE)-based deep generative models. We investigated whether the data augmentation leads to improved accuracy in the predictions of a deep neural network (DNN) model trained using the augmented data. Additionally, the DNN prediction uncertainties are quantified using Bayesian Neural Networks (BNN) and conformal prediction (CP) to assess the impact on predictive uncertainty reduction. To test the proposed methodology, we used TRACE simulations of steady-state void fraction data based on the NUPEC Boiling Water Reactor Full-size Fine-mesh Bundle Test (BFBT) benchmark. Here, we found that augmenting the training dataset using VAEs has improved the DNN model’s predictive accuracy, improved the prediction confidence intervals, and reduced the prediction uncertainties.

Bayesian neural network↗

Physics-Informed and Data-Driven Prediction of Residual Stress in Three-Dimensional Machining

Efficient and reliable prediction of machining-induced residual stress (RS) is a key requirement for truly integrated computational materials engineering (ICME). Currently available process modeling approaches, including empirical, analytical, and numerical methodologies lack predictive power and require substantial calibration and validation data. Moreover, most model-based approaches consider only two-dimensional (2D) (i.e., orthogonal), cutting processes. Meanwhile, industrial processes such as milling, turning, and drilling are inherently three-dimensional (3D). The present work attempts to bridge the gap between 2D and 3D through careful consideration of the process physics, including geometric, kinematic, and size-effect constraints to realize robust prediction of how RS develops in 3D machining. Using a novel in-situ experimental technique and digital image correlation (DIC) to determine equivalent Hertzian contact widths, contact pressures, and friction coefficients, the proposed methodology leverages a discretized conversion algorithm that includes multi-pass shakedown effects. This paper presents a semi-analytical model to predict machining-induced RS in 3D turning operations, which are used representatively for 3D processes more generally. Rather than follow a ‘brute force’ 3D FEM approach or conduct countless experiments to train a purely data-driven machine learning algorithm, the proposed approach builds on previous 2D modeling work. Through careful consideration of the process physics, including complex geometry/kinematic considerations of 3D turning, the authors demonstrated an experimentally calibrated approach, as well as validation based on published RS data. Model predictions and previously published measurement data of RS depth profiles for turning of Inconel 718 were compared for a range of process parameters. Correlation between the proposed 3D model and validation data was found to be within the margin of experimental error for most conditions. The proposed model appears to capture the overall behavior of 3D RS depth profiles with acceptable accuracy, particularly the key metrics of near-surface stress, peak stress magnitude and location, as well as overall stress profile depth. This report presents a physics-informed, data-driven approach for efficient calibration of a 2D model for machining-induced RS through DIC analysis of in-situ characterized subsurface displacement fields.

42 ENGINEERING↗

Exploring Structure-Sensitive Relations for Small Species Adsorption Using Machine Learning

Accurate prediction of adsorption energies on heterogeneous catalyst surfaces is crucial to predicting reactivity and screening materials. Adsorption linear scaling relations have been developed extensively but often lack accuracy and apply to one adsorbate and a single binding site type at a time. These facts undermine their ability to predict structure sensitivity and optimal catalyst structure. Using machine learning on nearly 300 density functional theory calculations, we demonstrate that generalized coordination number scaling relations hold well for oxygen- and high-valency carbon-binding species but fail for others. Here we reveal that the valency and the electronic coupling of a species with the surface, along with the site type and its coordination environment, are critical for small species adsorption. The model simultaneously predicts the adsorption energy and preferred site and significantly outperforms linear scalings in accuracy. It can expose the structure sensitivity of chemical reactions and enable enhanced catalyst activity via engineering particle shape and facet defects. The generality of our methodology is validated by training the model with transition metal data and transferring it to predict adsorption energies on single-atom alloys.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Machine-learning-accelerated multimodal characterization and multiobjective design optimization of natural porous materials

Natural porous materials such as nanoporous clays are used as green and low-cost adsorbents and catalysts. The key factors determining their performance in these applications are the pore morphology and surface activity, which are typically represented by properties such as specific surface area, pore volume, micropore content and pH. The latter may be modified and tuned to specific applications through material processing and/or chemical treatment. Characterization of the material, raw or processed, is typically performed experimentally, which can become costly especially in the context of tuning of the properties towards specific application requirements and needing numerous experiments. In this work, we present an application of tree-based machine learning methods trained on experimental datasets to accelerate the characterization of natural porous materials. The resulting models allow reliable prediction of the outcomes of experimental characterization of processed materials (R2 from 0.78 to 0.99) as well as identification of key factors contributing to those properties through feature importance analysis. Furthermore, the high throughput of the models enables exploration of processing parameter–property correlations and multiobjective optimization of prototype materials towards specific applications. We have applied these methodologies to pinpoint and rationalize optimal processing conditions for clays exploitable in acid catalysis. One of such identified materials was synthesized and tested revealing appreciable acid character improvement with respect to the pristine material. Specifically, it achieved 79% removal of chlorophyll-a in acid catalyzed degradation.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Evolving Multi-hazard Machine Learning Modeling for Advanced Risk-Informed Infrastructure Resilience Assessment

The socioeconomic impacts of pipeline incidents have escalated over the past three decades, revealing the limitation of traditional risk modeling methods when applied to extensive pipeline networks. This research aims to develop machine learning (ML) models that effectively identify, rank, and predict the diverse hazards and socioeconomic consequences associated with pipeline incidents. Utilizing historical data on pipeline incidents alongside weather and oceanographic data from the 1980s onward, the Houston metropolitan area serves as a testbed for the proposed methodologies. The research segments the combined datasets into three consecutive periods, demonstrating the efficacy of the updated model in predicting future events, particularly concerning precipitation rate data. Despite the challenges posed by a relatively limited dataset, local-level ML modeling offers valuable insights into the spatial and temporal dynamics of multiple hazards that contribute to pipeline incidents. These findings hold significant implications for future research, particularly in understanding and mitigating risks in various locations across the Gulf Coast and other coastal regions.

42 ENGINEERING↗

K-means-driven Gaussian Process data collection for angle-resolved photoemission spectroscopy

Abstract We propose the combination of k-means clustering with Gaussian Process (GP) regression in the analysis and exploration of 4D angle-resolved photoemission spectroscopy (ARPES) data. Using cluster labels as the driving metric on which the GP is trained, this method allows us to reconstruct the experimental phase diagram from as low as 12% of the original dataset size. In addition to the phase diagram, the GP is able to reconstruct spectra in energy-momentum space from this minimal set of data points. These findings suggest that this methodology can be used to improve the efficiency of ARPES data collection strategies for unknown samples. The practical feasibility of implementing this technology at a synchrotron beamline and the overall efficiency implications of this method are discussed with a view on enabling the collection of more samples or rapid identification of regions of interest.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Automated Development of Molten Salt Machine Learning Potentials: Application to LiCl

The in silico modeling of molten salts is critical for emerging "carbon-free" energy applications but is inhibited by the cost of quantum mechanically treating the high polarizabilities of molten salts. Here, we integrate configurational sampling using classical force fields with active learning to automate and accelerate the generation of Gaussian approximation potentials (GAP) for molten salts. This methodology reduces the number of expensive ab initio evaluations required for training set generation to O(100), enabling the facile parametrization of a molten LiCl GAP model that exhibits a 19 000-fold speedup relative to AIMD. The developed molten LiCl GAP model is applied to sample extended spatiotemporal scales, permitting new physical insights into molten LiCl's coordination structure as well as experimentally validated predictions of structures, densities, self-diffusion constants, and ionic conductivities. The developed methodology significantly lowers the barrier to the in silico understanding and design of molten salts across the periodic table.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Wave Detection and Tracking Within a Rotating Detonation Engine Through Object Detection

As the operational time window of experimental rotating detonation engines (RDEs) is expanded and the technology matures toward integration within gas turbines, monitoring techniques must evolve to offer computationally efficient and highly time-resolved diagnostics. In this study, computer vision object detection methodology that seeks to reduce data processing time and calculate wave velocity within drastically reduced time intervals as compared to traditional high-frame-rate RDE images analysis techniques is proposed. The adapted you-only-look-once object detection network is trained to detect individual detonation waves within single down-axis RDE images. The wave location and rotational direction detected within a frame are tracked through a series of high-speed images to calculate the frame-to-frame wave velocity with the time-step resolution of $\mathrm{20 μs}$ across a series of frames. The analysis of the annotation box size and image linearization effects is presented, demonstrating the lowest frame-to-frame velocity total uncertainty of $\mathrm{±3.8\%}$ and the highest classification speed of 9.5 frames per second using linearized images. Linearized images “unwrap” the RDE annulus pixel region to a reduced image size. Here, this new method offers great reductions in data processing times and unsteady detonation behavior insight at intervals more comparable to the timescales of detonation wave interactions via the application of machine learning to experimental RDE data.

33 ADVANCED PROPULSION SYSTEMS↗

Machine learning reduces soft costs for residential solar photovoltaics

Further deployment of rooftop solar photovoltaics (PV) hinges on the reduction of soft (non-hardware) costs—now larger and more resistant to reductions than hardware costs. The largest portion of these soft costs is the expenses solar companies incur to acquire new customers. In this study, we demonstrate the value of a shift from significance-based methodologies to prediction-oriented models to better identify PV adopters and reduce soft costs. We employ machine learning to predict PV adopters and non-adopters, and compare its prediction performance with logistic regression, the dominant significance-based method in technology adoption studies. Our results show that machine learning substantially enhances adoption prediction performance: The true positive rate of predicting adopters increased from 66 to 87%, and the true negative rate of predicting non-adopters increased from 75 to 88%. We attribute the enhanced performance to complex variable interactions and nonlinear effects incorporated by machine learning. With more accurate predictions, machine learning is able to reduce customer acquisition costs by 15% ($0.07/Watt) and identify new market opportunities for solar companies to expand and diversify their customer bases. Our research methods and findings provide broader implications for the adoption of similar clean energy technologies and related policy challenges such as market growth and energy inequality.

14 SOLAR ENERGY↗

Plant Single-Cell Solutions for Energy and the Environment (Second Workshop Report)

Plants are important sources of energy and materials, and they collectively represent a critical component of Earth’s ecosystem. With increasing environmental stresses due to climate change and intensive agricultural practices, the need for resilient plants is greater than ever before. To secure plant resources for bioenergy, biomaterials, food, and ecosystem adaptation, a deeper understanding of the fundamental biology of plants at a cellular level is urgently needed. Plants contain a multitude of specialized cell types that compose tissues and organs. Pathogens often target specific cell types within plants, and the response of one cell to a particular stimulus is likely to be distinct from its neighbor because of underlying molecular and contextual differences. Understanding how these responses are distributed among cells, the main goal of single-cell approaches, will substantially enhance our ability to use targeted engineering for improving plant productivity and resilience. Furthermore, single-cell approaches are necessary to understand the interactions between plants and other ecosystem members such as fungi, bacteria, and archaea. Unlocking these gene-response mechanisms at a cellular level can improve our ability to adapt plants to environmental stresses, increasing their utility as feedstocks for biomaterials and bioenergy. Recent advances in high-throughput sequencing, mass spectrometry, microfluidics and miniaturization, artificial intelligence and machine learning, and bioinformatics have greatly improved our ability to detect and understand processes at a cellular level. In mammalian systems, single-cell transcriptomics has already led to many advances, such as newly identified cell types and cell-targeted treatment of diseases, and mass spectrometry-based single-cell proteomics has recently been demonstrated as a promising emerging technology. However, plant single-cell omics has lagged behind mammalian approaches due to the high cost of the technologies relative to available resources and to the innate biological features of plants, including the complexity of the cell wall and polyploidy. To better understand how single-cell methods could enable plant science, Lawrence Berkeley National Laboratory (Berkeley Lab) hosted a workshop on April 29, 2021, that brought together a diverse group of leaders in plant and/or single-cell biology. Attendees represented federal research programs and domestic and international academic institutions. During the workshop, three presenters described the current state of research in both experimental and computational approaches. While the focus of the workshop was on factors preventing plant biology researchers from fully adopting single-cell methodologies, workshop participants agreed that most barriers could be overcome with focused, strategic investment and coordinated efforts among institutions leading to significant scientific discoveries that would be difficult to obtain using more conventional technologies.

59 BASIC BIOLOGICAL SCIENCES↗

Toward ultra-efficient high-fidelity predictions of wind turbine wakes: Augmenting the accuracy of engineering models with machine learning

This study proposes a novel machine learning (ML) methodology for the efficient and cost-effective prediction of high-fidelity three-dimensional velocity fields in the wake of utility-scale turbines. The model consists of an autoencoder convolutional neural network with U-Net skipped connections, fine-tuned using high-fidelity data from large-eddy simulations (LES). The trained model takes the low-fidelity velocity field cost-effectively generated from the analytical engineering wake model as input and produces the high-fidelity velocity fields. The accuracy of the proposed ML model is demonstrated in a utility-scale wind farm for which datasets of wake flow fields were previously generated using LES under various wind speeds, wind directions, and yaw angles. Comparing the ML model results with those of LES, the ML model was shown to reduce the error in the prediction from 20% obtained from the Gauss Curl hybrid (GCH) model to less than 5%. In addition, the ML model captured the non-symmetric wake deflection observed for opposing yaw angles for wake steering cases, demonstrating a greater accuracy than the GCH model. The computational cost of the ML model is on par with that of the analytical wake model while generating numerical outcomes nearly as accurate as those of the high-fidelity LES.

Mechanics↗

VAIM-CFF: a variational autoencoder inverse mapper solution to Compton form factor extraction from deeply virtual exclusive reactions

We develop a new methodology for extracting Compton form factors (CFFs) from deeply virtual exclusive reactions such as the unpolarized DVCS cross section using a specialized inverse problem solver, a variational autoencoder inverse mapper (VAIM). The VAIM-CFF framework not only allows us access to a fitted solution set possibly containing multiple solutions in the extraction of all 8 CFFs from a single cross section measurement, but also accesses the lost information contained in the forward mapping from CFFs to cross section. We investigate various assumptions and their effects on the predicted CFFs such as cross section organization, number of extracted CFFs, use of uncertainty quantification technique, and inclusion of prior physics information. We then use dimensionality reduction techniques such as principal component analysis to visualize the missing physics information tracked in the latent space of the VAIM framework. Through re-framing the extraction of CFFs as an inverse problem, we gain access to fundamental properties of the problem not comprehensible in standard fitting methodologies: exploring the limits of the information encoded in deeply virtual exclusive experiments.

Accelerator Physics↗

Machine-learning-assisted insight into spin ice Dy 2 Ti 2 O 7

Complex behavior poses challenges in extracting models from experiment. An example is spin liquid formation in frustrated magnets like Dy 2 Ti 2 O 7 . Understanding has been hindered by issues including disorder, glass formation, and interpretation of scattering data. Here, we use an automated capability to extract model Hamiltonians from data, and to identify different magnetic regimes. This involves training an autoencoder to learn a compressed representation of three-dimensional diffuse scattering, over a wide range of spin Hamiltonians. The autoencoder finds optimal matches according to scattering and heat capacity data and provides confidence intervals. Validation tests indicate that our optimal Hamiltonian accurately predicts temperature and field dependence of both magnetic structure and magnetization, as well as glass formation and irreversibility in Dy 2 Ti 2 O 7 . The autoencoder can also categorize different magnetic behaviors and eliminate background noise and artifacts in raw data. Our methodology is readily applicable to other materials and types of scattering problems.

36 MATERIALS SCIENCE↗

Maximizing machine learning interatomic potential transferability for the discovery of the novel stellated octadecagon Bi18-Pt24 cage structure

Achieving true transferability remains the central challenge for Machine Learning Interatomic Potentials (ML-IAPs) in modeling complex bimetallic nanoclusters across their vast potential energy surfaces. We systematically investigate data selection strategies to optimize the Chebyshev Interaction Model for Efficient Simulation (ChIMES) potential for the Bi-Pt nanoclusters by comparing three innovative sampling methods: Principal Component Analysis (PCA)/k-means (structural diversity), t-distributedStochasticNeighborEmbedding (t-SNE)/k-means (force-space diversity), and hierarchical clustering. Quantitatively, the PCA/k-means strategy proved most effective for global accuracy, yielding the lowest force errors and achieving energy root mean square errors (RMSE) values competitive with Density Functional Theory (DFT), demonstrating excellent accuracy (19.16meV/atom). Structural validation on 34 unique DFT-optimized isomers further confirmed the potential’s high fidelity, with the best model PCA/k-means reproducing structures with an average root mean square deviation (RMSD) of 0.10 Å. However, the t-SNE methods, by maximizing diversity in the force space, demonstrated superior extrapolative power, leading to the more precise prediction of a novel stellated octadecagon Bi18⁢Pt24 cage structure, demonstrating the potential for exploring previously unseen morphologies. Our results establish a clear methodology for strategic data sampling that successfully maximizes ML-IAP transferability, providing an accurate and computationally efficient tool that accelerates the theoretical discovery of complex bimetallic architectures.

Vangheluwe, Raphaël [Université Paris-Saclay, CNRS↗

Inference offers a metric to constrain dynamical models of neutrino flavor transformation

The multimessenger astrophysics of compact objects presents a vast range of environments where neutrino flavor transformation may occur and may be important for nucleosynthesis, dynamics, and a detected neutrino signal. Development of efficient techniques for surveying flavor evolution solution spaces in these diverse environments, which augment and complement existing sophisticated computational tools, could leverage progress in this field. To this end we continue our exploration of statistical data assimilation (SDA) to identify solutions to a small-scale model of neutrino flavor transformation. SDA is a machine learning formula wherein a dynamical model is assumed to generate any measured quantities. Specifically, we use an optimization formulation of SDA wherein a cost function is extremized via the variational method. Regions of state space in which the extremization identifies the global minimum of the cost function will correspond to parameter regimes in which a model solution can exist. Our example study seeks to infer the flavor transformation histories of two monoenergetic neutrino beams coherently interacting with each other and with a matter background. We require that the solution be consistent with measured neutrino flavor fluxes at the point of detection, and with constraints placed upon the flavor content at various locations along their trajectories, such as the point of emission, and the locations of the Mikheyev-Smirnov-Wolfenstein resonances. We show how the procedure efficiently identifies solution regimes and rules out regimes where solutions are infeasible. Overall, results in this work intimate the promise of this “variational annealing” methodology to efficiently probe an array of fundamental questions that traditional numerical simulation codes render difficult to access.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Fast Characterization of Inducible Regions of Atrial Fibrillation Models With Multi-Fidelity Gaussian Process Classification

Computational models of atrial fibrillation have successfully been used to predict optimal ablation sites. A critical step to assess the effect of an ablation pattern is to pace the model from different, potentially random, locations to determine whether arrhythmias can be induced in the atria. In this work, we propose to use multi-fidelity Gaussian process classification on Riemannian manifolds to efficiently determine the regions in the atria where arrhythmias are inducible. We build a probabilistic classifier that operates directly on the atrial surface. We take advantage of lower resolution models to explore the atrial surface and combine seamlessly with high-resolution models to identify regions of inducibility. We test our methodology in 9 different cases, with different levels of fibrosis and ablation treatments, totalling 1,800 high resolution and 900 low resolution simulations of atrial fibrillation. When trained with 40 samples, our multi-fidelity classifier that combines low and high resolution models, shows a balanced accuracy that is, on average, 5.7% higher than a nearest neighbor classifier. We hope that this new technique will allow faster and more precise clinical applications of computational models for atrial fibrillation. All data and code accompanying this manuscript will be made publicly available at: https://github.com/fsahli/AtrialMFclass.

59 BASIC BIOLOGICAL SCIENCES↗