Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Neural network embeddings”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Actinium–DOTA coordination in water from hybrid ML/MM: Structure, free energies, and water-exchange pathways

Quantitative simulation of trivalent ƒ-block chelates in water remains challenging because bonded and non-bonded force-field models make different approximations for coordination structure, exchange dynamics, and ion–ligand interactions in highly charged systems. Here, we develop a hybrid machine-learning/molecular-mechanics (ML/MM) framework for Ac 3+ –DOTA in explicit solvent by training an E(3)-equivariant neural network potential (MACELES) on mechanically embedded QM/MM data for Ac aquo and Ac–DOTA species and coupling it to NAMD 2.14 with particle-mesh Ewald electrostatics. Nanosecond ML/MM trajectories remain numerically stable and preserve chelate integrity, yielding a compact DOTA inner shell with an inner-sphere water coordination number of CN Ac,O w ≈ 1.7 arising from a dynamic equilibrium between one- and two-water states (37.5% and 59.9% of frames; three waters 2.5%). A 5 ns potential of mean force shows two low-lying basins at CN Ac,O w ≈ 1 and CN Ac,O w ≈ 2. DFT end-state free energies are consistent with the ML/MM profile, and DFT minimum-energy paths provide a qualitative electronic-structure reference for the observed basin connectivity. State-resolved kinetics reveal picosecond water-exchange pathways that couple hydration changes to transient DOTA arm fluctuations, and training-set comparisons show that temperature-matched Ac–DOTA data optimize energy/force accuracy while more diverse solvated data improve charge prediction. Overall, the present hybrid ML/MM model provides a practical description of Ac 3+ –DOTA hydration thermodynamics and short-time exchange behavior in explicit water at MD-like cost.

Actinium↗

E(n)-Equivariant cartesian tensor message passing interatomic potential

Machine learning potential (MLP) has been a popular topic in recent years for its capability to replace expensive first-principles calculations in some large systems. Meanwhile, message passing networks have gained significant attention due to their remarkable accuracy, and a wave of message passing networks based on Cartesian coordinates has emerged. However, the information of the node in these models is usually limited to scalars, and vectors. In this work, we propose High-order Tensor message Passing interatomic Potential (HotPP), an E(n) equivariant message passing neural network that extends the node embedding and message to an arbitrary order tensor. By performing some basic equivariant operations, high order tensors can be coupled very simply and thus the model can make direct predictions of high-order tensors such as dipole moments and polarizabilities without any modifications. The tests in several datasets show that HotPP not only achieves high accuracy in predicting target properties, but also successfully performs tasks such as calculating phonon spectra, infrared spectra, and Raman spectra, demonstrating its potential as a tool for future research.

97 MATHEMATICS AND COMPUTING↗

Multimodal representation learning for predicting molecule–disease relations

Motivation: Predicting molecule–disease indications and side effects is important for drug development and pharmacovigilance. Comprehensively mining molecule–molecule, molecule–disease and disease–disease semantic dependencies can potentially improve prediction performance. Methods: We introduce a Multi-Modal REpresentation Mapping Approach to Predicting molecular-disease relations (M2REMAP) by incorporating clinical semantics learned from electronic health records (EHR) of 12.6 million patients. Specifically, M2REMAP first learns a multimodal molecule representation that synthesizes chemical property and clinical semantic information by mapping molecule chemicals via a deep neural network onto the clinical semantic embedding space shared by drugs, diseases and other common clinical concepts. To infer molecule–disease relations, M2REMAP combines multimodal molecule representation and disease semantic embedding to jointly infer indications and side effects. Results: We extensively evaluate M2REMAP on molecule indications, side effects and interactions. Results show that incorporating EHR embeddings improves performance significantly, for example, attaining an improvement over the baseline models by 23.6% in PRC-AUC on indications and 23.9% on side effects. Further, M2REMAP overcomes the limitation of existing methods and effectively predicts drugs for novel diseases and emerging pathogens. Availability and implementation: The code is available at https://github.com/celehs/M2REMAP, and prediction results are provided at https://shiny.parse-health.org/drugs-diseases-dev/.

59 BASIC BIOLOGICAL SCIENCES↗

Evaluating the Use of Foundational Chemical Language Models in Multimodal Graph Fusion

Rapid and accurate prediction of the physicochemical properties of molecules given their structures remains a key challenge in cheminformatics. Machine learning approaches offer high-throughput options, but the optimality of inductive biases and data representations are up for debate. For example, BERT-based masked language models (MLMs) can be trained in a self-supervised way on hundreds of millions to billions of readily available SMILES strings. Another option is graph neural networks (GNNs), which can operate directly on molecular structures. Yet, generating accurate molecular geometry is computationally expensive, leading to a relative scarcity in data compared to SMILES strings. It is attractive to combine these two paradigms by pre-training an LM on a large corpus of SMILES strings and embedding these representation into a geometric graph neural network. Despite the promise of such an approach, and contrary to previous studies, we find mixed results with the combination of the LMs and GNNs on several molecule datasets. In particular, we found evidence for improvement on the FreeSolv and QM7 benchmarks, but degraded performance on the ESOL, LIPO and QM9 datasets compared to a GNN baseline.

Francel, Collin [University of Alabama]↗

Power System Event Identification Based on Deep Neural Network With Information Loading

Online power system event identification and classification are crucial to enhancing the reliability of transmission systems. In this study, we develop a deep neural network (DNN) based approach to identify and classify power system events by leveraging real-world measurements from hundreds of phasor measurement units (PMUs) and labels from thousands of events. Two innovative designs are embedded into the baseline model built on convolutional neural networks (CNNs) to improve the event classification accuracy. First, we propose a graph signal processing based PMU sorting algorithm to improve the learning efficiency of CNNs. Second, we deploy information loading based regularization to strike the right balance between memorization and generalization for the DNN. Numerical results based on real-world dataset from the Eastern Interconnection of the U.S power transmission grid show that the combination of PMU based sorting and the information loading based regularization techniques help the proposed DNN approach achieve highly accurate event identification and classification results.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Embedded, Real-Time, and Distributed Traveling Wave Fault Location Method Using Graph Convolutional Neural Networks

This work proposes and develops an implementation of a fault location method to provide a fast and resilient protection scheme for power distribution systems. The method analyzes the transient dynamics of traveling waves (TWs) to generate features using the discrete wavelet transform (DWT), which are then used to train several graph convolutional network (GCN) models. Faults are simulated in the IEEE 34-node system, which is divided into three protection zones (PZs). The goal is to identify the PZ in which the fault occurs. The GCN models create a distributed protection scheme, as all nodes are able to retrieve a prediction. Given that message-passing between nodes occurs both during training and in the execution of the model, the resiliency of such schemes to communication losses was analyzed and demonstrated. One of the models, which only uses voltage measurements, was implemented on a Texas Instruments F28379D development board. The execution times were monitored to assess the speed of the protection scheme. It is shown that the proposed method can be executed in approximately a millisecond, which is comparable to existing TW protection in the transmission system. For experimental purposes, a DWT-based detection method is employed. A design of a setup to playback TWs using two development boards is also addressed.

Jiménez-Aparicio, Miguel (ORCID:000000016864461X)↗

SpeckleNN: a unified embedding for real-time speckle pattern classification in X-ray single-particle imaging with limited labeled examples

With X-ray free-electron lasers (XFELs), it is possible to determine the three-dimensional structure of noncrystalline nanoscale particles using X-ray single-particle imaging (SPI) techniques at room temperature. Classifying SPI scattering patterns, or `speckles', to extract single-hits that are needed for real-time vetoing and three-dimensional reconstruction poses a challenge for high-data-rate facilities like the European XFEL and LCLS-II-HE. Here, we introduce SpeckleNN, a unified embedding model for real-time speckle pattern classification with limited labeled examples that can scale linearly with dataset size. Trained with twin neural networks, SpeckleNN maps speckle patterns to a unified embedding vector space, where similarity is measured by Euclidean distance. We highlight its few-shot classification capability on new never-seen samples and its robust performance despite having only tens of labels per classification category even in the presence of substantial missing detector areas. Without the need for excessive manual labeling or even a full detector image, our classification method offers a great solution for real-time high-throughput SPI experiments.

47 OTHER INSTRUMENTATION↗

Graph-EAM: An Interpretable and Efficient Graph Neural Network Potential Framework

The development of deep learning interatomic potentials has enabled efficient and accurate computations in quantum chemistry and materials science, circumventing computationally expensive ab initio calculations. However, the huge number of learnable parameters in deep learning models and their complex architectures hinder physical interpretability and affect the robustness of the derived potential. In this work, we propose graph-EAM, a lightweight graph neural network (GNN) inspired by the empirical embedded atom method to model the interatomic potential of single-element structures. Four material systems: platinum, niobium, silicon, and amorphous-carbon, for which quantum simulation data sets are publicly available, are examined to demonstrate that graph-EAM can achieve high energy and force prediction accuracy-comparable or better than existing state-of-the-art machine learning models-with much fewer parameters. It is also shown that the explicit inclusion of the angular information via three-body atomic density increases the prediction accuracy. In conclusion, the accuracy and efficiency of potentials obtained from graph-EAM can help accelerate the molecular dynamics simulation.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

An finite element analysis surrogate model with boundary oriented graph embedding approach for rapid design

Abstract In this work, we present a boundary oriented graph embedding (BOGE) approach for the graph neural network to assist in rapid design and digital prototyping. The cantilever beam problem has been solved as an example to validate its potential of providing physical field results and optimized designs using only 10 ms. Providing shortcuts for both boundary elements and local neighbor elements, the BOGE approach can embed unstructured mesh elements into the graph and performs an efficient regression on large-scale triangular-mesh-based finite element analysis (FEA) results, which cannot be realized by other machine-learning-based surrogate methods. It has the potential to serve as a surrogate model for other boundary value problems. Focusing on the cantilever beam problem, the BOGE approach with 3-layer DeepGCN model achieves the regression with mean square error (MSE) of 0.011 706 (2.41% mean absolute percentage error) for stress field prediction and 0.002 735 MSE (with 1.58% elements having error larger than 0.01) for topological optimization. The overall concept of the BOGE approach paves the way for a general and efficient deep-learning-based FEA simulator that will benefit both industry and Computer Aided Design (CAD) design-related areas.

42 ENGINEERING↗

A Deep Learning Pipeline for Optimizing Large-scale Phase Field Simulations

Phase field (PF) simulations are computationally expensive but remain a key analysis tool to understand the complex mechanisms of additive manufacturing (AM) processes. Each PF simulation-aided analysis requires thousands of node hours on leadership-class supercomputers. One of the main goals of these analyses is the study of microstructure evolution during the build process which begins with the onset of nucleation. Nucleation occurs under certain thermomechanical conditions which are not known a priori and many PF simulations are required to identify ranges of input thermo-mechanical parameters that can result in the onset of nucleation. Since many of the simulations do not result in nucleation, an analysis campaign often ends up wasting tremendous amounts of precious computing resources executing nucleation-absent simulations. The goal of this work is to design and train deep learning models to inform a PF simulation about the likelihood of the occurrence of nucleation in a future simulation time-step based on the state summary over a finite number of past time-steps of a running simulation. If the prediction determines that the running simulation is unlikely to reach nucleation in the allotted time, then its execution is stopped immediately ultimately resulting in vast reduction in wasted computations when accrued over all the PF simulations typically performed in a single or multiple analysis campaign(s). The paper presents the performance of a machine learning pipeline that uses a convolutional neural network (CNN) model to learn an embedding which is then used with a self-attention network to build a multi-task deep learning model to predict the likelihood of nucleation. The model also predicts the input parameters used in a simulation. Performance is compared with a baseline pipeline that uses an off-the-shelf LeNet-5 model to learn the initial embedding. Despite their smaller size, performance results indicate significant improvement in accuracy of the proposed models compared to the larger baseline models.

Kannan, Ramakrishnan {ramki}↗

Learned adaptive properties for mitigation of weight perturbations in embedded spiking networks

Recent years have seen an increased importance of neural network inference in edge-based scenarios, which impose size and power constraints requiring novel computing devices. These same edge scenarios may require operating over long periods of time, or exposure to extreme environments, resulting in a drift of neural network weights that cause degraded performance. In searching for ways to develop neural network approaches that perform robustly under these conditions, we propose a biologically-inspired mechanism for the dynamic adaptation of within-neuron parameters that is guided by a global context signal carrying information about perturbations and variability in incoming stimuli. Specifically, we demonstrate that adaptive voltage thresholds or neuronal time constants, when informed by a global context signal, can enable network-level mechanisms to recover from perturbed synaptic weights. Consistent with prior literature, the context-modulated approach is effective for recurrent, but not feedforward networks, by modulating network level dynamics. We demonstrate this approach successfully recovers performance in image classification tasks and spatiotemporal tracking tasks under idealized and Gaussian noise as well as for realistic perturbations from a memristive device when exposed to ionizing radiation. Finally, we discuss how this approach enables the design of robust and energy-efficient neuromorphic systems that perform well, even in resource-constrained scenarios with extreme environments such as edge processing.

context modulation↗

Identification of Flux Rope Orientation via Neural Networks

Geomagnetic disturbance forecasting is based on the identification of solar wind structures and accurate determination of their magnetic field orientation. For nowcasting activities, this is currently a tedious and manual process. Focusing on the main driver of geomagnetic disturbances, the twisted internal magnetic field of interplanetary coronal mass ejections (ICMEs), we explore a convolutional neural network’s (CNN) ability to predict the embedded magnetic flux rope’s orientation once it has been identified from in situ solar wind observations. Our work uses CNNs trained with magnetic field vectors from analytical flux rope data. The simulated flux ropes span many possible spacecraft trajectories and flux rope orientations. We train CNNs first with full duration flux ropes and then again with partial duration flux ropes. The former provides us with a baseline of how well CNNs can predict flux rope orientation while the latter provides insights into real-time forecasting by exploring how accuracy is affected by percentage of flux rope observed. The process of casting the physics problem as a machine learning problem is discussed as well as the impacts of different factors on prediction accuracy such as flux rope fluctuations and different neural network topologies. Finally, results from evaluating the trained network against observed ICMEs from Wind during 1995–2015 are presented.

Thomas Narock↗

Multi-head physics-informed neural networks for learning functional priors and uncertainty quantification

In numerous applications, the integration of prior knowledge and historical information is essential, particularly for tasks requiring the solution of ordinary or partial differential equations (ODEs/PDEs) in data-sparse or noisy environments. For instance, achieving accurate solutions to time-dependent PDEs with limited initial condition measurements necessitates an effective strategy for embedding prior knowledge. Hard-parameter sharing architectures in neural networks (NNs) have demonstrated success in both traditional and scientific machine learning domains, facilitating the learning of informative representations. Here, in this study, we introduce a novel, yet efficient, method to enhance physics-informed neural networks (PINNs) by incorporating a multi-head structure that enables the learning of functional priors from both empirical data and governing physical laws. This prior information can then be used to address data sparsity and high-level noise in solving ODE/PDE problems with uncertainty quantification (UQ). The approach, termed Multi-Head PINN (MH-PINN), consists of a shared body NN and multiple head NNs, each corresponding to an individual PINN instance. Our framework for functional prior learning is carried out in two stages: (1) training the MH-PINNs to develop a shared body NN alongside multiple head NNs, and (2) employing these trained head NNs to estimate a prior distribution through a normalizing flow-based density estimator. The learned functional prior can then be applied as a regularization mechanism in deterministic contexts or as an informative prior within a Bayesian inference framework, aiding in the resolution of subsequent ODE/PDE tasks. We evaluate the efficacy of MH-PINNs across five benchmark problems, including a high-dimensional parametric PDE, all characterized by data sparsity or substantial noise levels. Our findings reveal that MH-PINNs deliver accurate solutions and robust UQ, demonstrating adaptability across a range of complex and challenging scenarios.

Bayesian inference↗

Jet rotational metrics

Abstract Embedding symmetries in the architectures of deep neural networks can improve classification and network convergence in the context of jet substructure. These results hint at the existence of symmetries in jet energy depositions, such as rotational symmetry, arising from the physical features of the underlying processes. We introduce new jet observables, Jet Rotational Metrics (JRMs), which provide insights into the substructure of jets by comparing them to jets with perfect discrete rotational symmetry. We show that JRMs are formidable jet features, achieving good classification scores when used as inputs to deep neural networks. We also show that when used in combination with other jet observables, like N-subjettiness and EFPs, our features increase classification performance. The results suggest that JRMs may capture information not efficiently captured by the other observables, motivating the design of future jet observables for learning the underlying symmetries in the physical processes.

Physics↗

Hybrid interatomic potential for Sn

To design materials for extreme applications, it is important to understand and predict phase transitions and their influence on material properties under high pressures and temperatures. Atomistic modeling can be a useful tool to assess these behaviors. However, this can be difficult due to the lack of fidelity of the interatomic potentials in reproducing this high pressure and temperature extreme behavior. Here, in this work, a hybrid EAM-R—which is the combination of embedded atom method (EAM) and rapid artificial neural network potential—for Tin (Sn) is described which is capable of accurately modeling the complex sequence of phase transitions between different metallic polymorphs as a function of pressure. This hybrid approach ensures that a basic empirical potential like EAM is used as a lower energy bound. By using the final activation function, the neural network contribution to energy must be positive, assuring stability over the whole configuration space. This implementation has the capacity to reproduce density functional theory results at 6 orders of magnitude slower than a pair potential for molecular dynamics simulation, including elastic and plastic characteristics and relative energies of each phase. Using calculations of the Gibbs free energy, it is demonstrated that the potential precisely predicts the experimentally observed phase changes at temperatures and pressures across the whole phase diagram. At 10.2 GPa, the present potential predicts a first-order phase transition between body-centered tetragonal (BCT) β-Sn and another polymorph of BCT-Sn. This structure transforms into body-centered cubic near the experimentally reported value at 33 GPa. Thus, the Sn potential developed in this paper can be used to study complex deformation mechanisms under extreme conditions of high pressure and strain rates unlike existing potentials. Moreover, the framework developed in this paper can be extended for different material systems with complex phase diagrams.

36 MATERIALS SCIENCE↗

ScaWL: Scaling k-WL (Weisfeiler-Lehman) Algorithms in Memory and Performance on Shared and Distributed-Memory Systems

The k-dimensional Weisfeiler-Lehman (k-WL) algorithm—developed as an efficient heuristic for testing if two graphs are isomorphic—is a fundamental kernel for node embedding in the emerging field of graph neural networks. Unfortunately, the k-WL algorithm has exponential storage requirements, limiting the size of graphs that can be handled. This work presents a novel k-WL scheme with a storage requirement orders of magnitude lower while maintaining the same accuracy as the original k-WL algorithm. Due to the reduced storage requirement, our scheme allows for processing much bigger graphs than previously possible on a single compute node. For even bigger graphs, we provide the first distributed-memory implementation. Our k-WL scheme also has significantly reduced communication volume and offers high scalability. Our experimental results demonstrate that our approach is significantly faster and has superior scalability compared to five other implementations employing state-of-the-art techniques.

algorithims↗

Use of Soft Computing Technologies For Rocket Engine Control

The problem to be addressed in this paper is to explore how the use of Soft Computing Technologies (SCT) could be employed to further improve overall engine system reliability and performance. Specifically, this will be presented by enhancing rocket engine control and engine health management (EHM) using SCT coupled with conventional control technologies, and sound software engineering practices used in Marshall s Flight Software Group. The principle goals are to improve software management, software development time and maintenance, processor execution, fault tolerance and mitigation, and nonlinear control in power level transitions. The intent is not to discuss any shortcomings of existing engine control and EHM methodologies, but to provide alternative design choices for control, EHM, implementation, performance, and sustaining engineering. The approaches outlined in this paper will require knowledge in the fields of rocket engine propulsion, software engineering for embedded systems, and soft computing technologies (i.e., neural networks, fuzzy logic, and Bayesian belief networks), much of which is presented in this paper. The first targeted demonstration rocket engine platform is the MC-1 (formerly FASTRAC Engine) which is simulated with hardware and software in the Marshall Avionics & Software Testbed laboratory that

Trevino, Luis C.↗

Constrained Block Nonlinear Neural Dynamical Models

Neural network modules conditioned by known priors can be effectively trained and combined to represent systems with nonlinear dynamics. This work explores a novel formulation for data-efficient learning of deep control-oriented nonlinear dynamical models by embedding local model structure and constraints. The proposed method consists of neural network blocks that represent input, state, and output dynamics with constraints placed on the network weights and system variables. For handling partially observable dynamical systems, we utilize a state observer neural network to estimate the states of the system's latent dynamics. We evaluate the performance of the proposed architecture and training methods on system identification tasks for three nonlinear systems: a continuous stirred tank reactor, a two tank interacting system, and an aerodynamics body. Models optimized with a few thousand system state observations accurately represent system dynamics in open loop simulation over thousands of time steps from a single set of initial conditions. Experimental results demonstrate an order of magnitude reduction in open-loop simulation mean squared error for our constrained, block-structured neural models when compared to traditional unstructured and unconstrained neural network models.

Skomski, Elliott↗