Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Neural state space models”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Screening of Li-Based Solid Electrolytes Using Bond-Valence Methods and Graph Neural Networks

Li-based solid-state electrolyte (Li-SSE) materials enable safer, all-solid-state batteries but the computational search for candidates with favorable stability and Li-ion conductivity is challenging due to the size of the search space and the cost of evaluating transport properties with ab initio methods. We present a high-throughput screening approach for Li-SSE materials using a combination of bond-valence methods and graph neural networks. We demonstrate the screening approach with a dataset containing tens of thousands of Li-containing compounds. Furthermore, we combine the machine-learning screening procedure with an isovalent substitution scheme to generate and screen additional Li SSE candidates beyond existing databases. Finally, we discuss relative importances of geometric and bond-valence quantities in the training of graph neural networks, providing insight for future modeling of ionic conductivity in Li-SSE materials.

Materials discovery↗

Fast and accurate simulations of calorimeter showers with normalizing flows

In this study, we introduce caloflow, a fast detector simulation framework based on normalizing flows. For the first time, we demonstrate that normalizing flows can reproduce many-channel calorimeter showers with extremely high fidelity, providing a fresh alternative to computationally expensive geant4 simulations, as well as other state-of-the-art fast simulation frameworks based on generative adversarial networks (GANs) or variational autoencoders (VAEs). In addition to the usual histograms of physical features and images of calorimeter showers, we introduce a new metric for judging the quality of generative modeling: the performance of a classifier trained to differentiate real from generated images. We show that GAN-generated images can be identified by the classifier with nearly 100% accuracy, while images generated from caloflow are better able to fool the classifier. More broadly, normalizing flows offer several advantages compared to other state-of-the-art approaches (GANs and VAEs), including tractable likelihoods, stable and convergent training, and principled model selection. Normalizing flows also provide a bijective mapping between data and the latent space, which could have other applications beyond simulation, for example, to detector unfolding.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

One-shot gas detection with transformer paired neural networks in Mako collected longwave infrared hyperspectral imagery

To date, careful data treatment workflows and statistical detectors are used to perform hyperspectral image (HSI) detection of any gas contained in a spectral library, which is often expanded with physics models to incorporate different spectral characteristics. In general, surrounding evidence or known gas-release parameters are used to provide confidence in or confirm detection capability, respectively. This makes quantifying detection performance difficult as it is nearly impossible to develop an absolute ground truth for gas target pixel presence in collected HSI. Consequently, developing and comparing new detection methods, especially machine learning (ML)-based methods, is susceptible to subjectivity in derived detection map quality. Here, in this work, we demonstrate the first use of transformer-based paired neural networks (PNNs) for one-shot gas target detection for multiple gases while providing quantitative classification and detection metrics for their use on labeled data. Terabytes of training data are generated from a database of long-wave infrared HSI obtained from historical Mako sensor campaigns over Los Angeles. By incorporating labels, singular signature representations, and a model development pipeline, we can tune and select PNNs to detect multiple gas targets that are not seen in training on a quantitative basis. We additionally assess our test set detections using interpretability techniques widely employed with ML-based predictors, but less common with detection methods relying on learned latent spaces.

Hyperspectral imaging↗

Mitigating spectral bias in neural operators via high-frequency scaling for physical systems

Neural operators have emerged as powerful surrogates for modeling complex physical problems. However, they suffer from spectral bias making them oblivious to high-frequency modes, which are present in multiscale physical systems. Therefore, they tend to produce over-smoothed solutions, which is particularly problematic in modeling turbulence and for systems with intricate patterns and sharp gradients such as multi-phase flow systems. In this work, we introduce a new approach named high-frequency scaling (HFS) to mitigate spectral bias in convolutional-based neural operators. By integrating HFS with proper variants of UNet, we demonstrate a higher prediction accuracy by mitigating spectral bias in single and two-phase flow problems. Unlike Fourierbased techniques, HFS is directly applied to the latent space, thus eliminating the computational cost associated with the Fourier transform. Additionally, we investigate alternative spectral bias mitigation through a diffusion model conditioned on neural operators. While the diffusion model integrated with the standard neural operator may still suffer from significant errors, these errors are substantially reduced when the diffusion model is integrated with a HFS-enhanced neural operator.

97 MATHEMATICS AND COMPUTING↗

PolyODENet

Kinetics of a reaction network that follows mass-action rate laws can be described with a system of ordinary differential equations (ODEs) with polynomial right-hand side. However, it is challenging to derive such kinetic differential equations from transient kinetic data without knowing the reaction network, especially when the data are incomplete due to experimental limitations. We introduce a program, PolyODENet, toward this goal. Based on the machine-learning method Neural ODE, PolyODENet defines a generative model and predicts concentrations at arbitrary time. As such, it is possible to include unmeasurable intermediate species in the kinetic equations. Importantly, we have implemented various measures to apply physical constraints and chemical knowledge in the training to regularize the solution space.

Wu, Qin [Brookhaven National Lab. (BNL), Upton, NY↗

Spatio-spectral optical fission in time-varying subwavelength layers

Transparent conducting oxides are highly doped semiconductors that exhibit favourable optical features compared with metals, including reduced material losses, tuneable electronic and optical properties, and enhanced damage thresholds. Recently, the photonic community has renewed its attention towards these materials, recognizing their remarkable nonlinear optical properties in the near-infrared spectrum. The exceptionally large and ultrafast change in the refractive index, which can be optically induced in these compounds, extends beyond the boundaries of conventional perturbative analysis and makes this class of materials the closest approximation to a time-varying system. Here we report the spatio-spectral fission of an ultrafast pulse trespassing a thin film of aluminium zinc oxide with a non-stationary refractive index. By applying phase conservation to this time-varying layer, our model can account for both space and time refraction and explain, in quantitative terms, the spatial separation of both spectrum and energy. Our findings represent an example of extreme nonlinear phenomena on subwavelength propagation distances, which provides new insights into transparent conducting oxides’ transient optical properties. This can be critical for the ongoing research on photonic time crystals, on-chip generation of non-classical states of light, integrated optical neural networks, ultrafast beam steering and frequency-division multiplexing.

integrated optics↗

Identification and Control of Aircrafts using Multiple Models and Adaptive Critics

We compared two possible implementations of local linear models for control: one approach is based on a self-organizing map (SOM) to cluster the dynamics followed by a set of linear models operating at each cluster. Therefore the gating function is hard (a single local model will represent the regional dynamics). This simplifies the controller design since there is a one to one mapping between controllers and local models. The second approach uses a soft gate using a probabilistic framework based on a Gaussian Mixture Model (also called a dynamic mixture of experts). In this approach several models may be active at a given time, we can expect a smaller number of models, but the controller design is more involved, with potentially better noise rejection characteristics. Our experiments showed that the SOM provides overall best performance in high SNRs, but the performance degrades faster than with the GMM for the same noise conditions. The SOM approach required about an order of magnitude more models than the GMM, so in terms of implementation cost, the GMM is preferable. The design of the SOM is straight forward, while the design of the GMM controllers, although still reasonable, is more involved and needs more care in the selection of the parameters. Either one of these locally linear approaches outperform global nonlinear controllers based on neural networks, such as the time delay neural network (TDNN). Therefore, in essence the local model approach warrants practical implementations. In order to call the attention of the control community for this design methodology we extended successfully the multiple model approach to PID controllers (still today the most widely used control scheme in the industry), and wrote a paper on this subject. The echo state network (ESN) is a recurrent neural network with the special characteristics that only the output parameters are trained. The recurrent connections are preset according to the problem domain and are fixed. In a nutshell, the states of the reservoir of recurrent processing elements implement a projection space, where the desired response is optimally projected. This architecture trades training efficiency by a large increase in the dimension of the recurrent layer. However, the power of the recurrent neural networks can be brought to bear on practical difficult problems. Our goal was to implement an adaptive critic architecture implementing Bellman s approach to optimal control. However, we could only characterize the ESN performance as a critic in value function evaluation, which is just one of the pieces of the overall adaptive critic controller. The results were very convincing, and the simplicity of the implementation was unparalleled.

Principe, Jose C.↗

Use of Design of Experiments in Determining Neural Network Architectures for Loss of Control Detection

We describe empirical methods for selecting a neural network architecture to implement belief state inference on generic commercial transport aircraft. We highlight a case study on the planning, execution, and analysis of a set of experiments to determine the configurations of a conditional variational autoencoder (CVAE). Our main contribution is the application of a structured method that can be used for machine learning in many aerospace applications. This method optimizes the structure and training parameters of a neural network for belief state inference, using Design of Experiments (DOE) statistical methodologies. The motivation for this specific DOE analysis was to identify the appropriate hyperparameters for measuring the CVAE reconstruction probability and latent space, such that the measurements can be used to infer qualitative state changes for the aircraft. We demonstrate that this process yields information about a trained neural network’s utility for this specific application, along with a quantifiable range of certainty. We execute 84 experiments using loss-of-control flight maneuver data from the NASA T-2 aircraft, demonstrating that this empirical process allows us to construct cheap and simple models with specific attributes amenable to belief state inference in aerospace applications.

Loss of Control↗

Use of Design of Experiments in Determining Neural Network Architectures for Loss of Control Detection

We describe empirical methods for selecting a neural network architecture to implement belief state inference on generic commercial transport aircraft. We highlight a case study on the planning, execution, and analysis of a set of experiments to determine the configurations of a conditional variational autoencoder (CVAE). Our main contribution is the application of a structured method that can be used for machine learning in many aerospace applications. This method optimizes the structure and training parameters of a neural network for belief state inference, using Design of Experiments (DOE) statistical methodologies. The motivation for this specific DOE analysis was to identify the appropriate hyperparameters for measuring the CVAE reconstruction probability and latent space, such that the measurements can be used to infer qualitative state changes for the aircraft. We demonstrate that this process yields information about a trained neural network’s utility for this specific application, along with a quantifiable range of certainty. We execute 84 experiments using loss-of-control flight maneuver data from the NASA T 2 aircraft, demonstrating that this empirical process allows us to construct cheap and simple models with specific attributes amenable to belief state inference in aerospace applications.

Loss of Control↗

Use of Design of Experiments in Determining Neural Network Architectures for Loss of Control Detection

Abstract—We describe empirical methods for selecting a neural network architecture to implement belief state inference on generic commercial transport aircraft. We highlight a case study on the planning, execution, and analysis of a set of experiments to determine the configurations of a conditional variational autoencoder (CVAE). Our main contribution is the application of a structured method that can be used for machine learning in many aerospace applications. This method optimizes the structure and training parameters of a neural network for belief state inference, using Design of Experiments (DOE) statistical methodologies. The motivation for this specific DOE analysis was to identify the appropriate hyperparameters for measuring the CVAE reconstruction probability and latent space, such that the measurements can be used to infer qualitative state changes for the aircraft. We demonstrate that this process yields information about a trained neural network’s utility for this specific application, along with a quantifiable range of certainty. We execute 84 experiments using loss-of-control flight maneuver data from the NASA T-2 aircraft, demonstrating that this empirical process allows us to construct cheap and simple models with specific attributes amenable to belief state inference in aerospace applications.

neural networks↗

Variational Autoencoders for Learning Nonlinear Dynamics of Physical Systems

We develop data-driven methods for incorporating physical information for priors to learn parsimonious representations of nonlinear systems arising from parameterized PDEs and mechanics. Our approach is based on Variational Autoencoders (VAEs) for learning nonlinear state space models from observations. We develop ways to incorporate geometric and topological priors through general manifold latent space representations. We investigate the performance of our methods for learning low dimensional representations for the nonlinear Burgers equation and constrained mechanical systems.

97 MATHEMATICS AND COMPUTING↗

Super resolution for root imaging

Premise High‐resolution cameras are very helpful for plant phenotyping as their images enable tasks such as target vs. background discrimination and the measurement and analysis of fine above‐ground plant attributes. However, the acquisition of high‐resolution images of plant roots is more challenging than above‐ground data collection. An effective super‐resolution (SR) algorithm is therefore needed for overcoming the resolution limitations of sensors, reducing storage space requirements, and boosting the performance of subsequent analyses. Methods We propose an SR framework for enhancing images of plant roots using convolutional neural networks. We compare three alternatives for training the SR model: (i) training with non‐plant‐root images, (ii) training with plant‐root images, and (iii) pretraining the model with non‐plant‐root images and fine‐tuning with plant‐root images. The architectures of the SR models were based on two state‐of‐the‐art deep learning approaches: a fast SR convolutional neural network and an SR generative adversarial network. Results In our experiments, we observed that the SR models improved the quality of low‐resolution images of plant roots in an unseen data set in terms of the signal‐to‐noise ratio. We used a collection of publicly available data sets to demonstrate that the SR models outperform the basic bicubic interpolation, even when trained with non‐root data sets. Discussion The incorporation of a deep learning–based SR model in the imaging process enhances the quality of low‐resolution images of plant roots. We demonstrate that SR preprocessing boosts the performance of a machine learning system trained to separate plant roots from their background. Our segmentation experiments also show that high performance on this task can be achieved independently of the signal‐to‐noise ratio. We therefore conclude that the quality of the image enhancement depends on the desired application.

Ruiz‐Munoz, Jose F.↗

Long–short-term memory encoder–decoder with regularized hidden dynamics for fault detection in industrial processes

The ability of recurrent neural networks (RNN) to model nonlinear dynamics of high dimensional process data has enabled data-driven RNN-based fault detection algorithms. Previous studies have focused on detecting faults by identifying the discrepancies in data distribution between the faulty and normal data, as reflected in prediction errors generated by RNN models. However, in industrial processes, variations in data distribution can also result from changes in normal control setpoints and compensatory control adjustments in response to disturbances, making it hard to differentiate between normal and faulty conditions. This paper proposes a fault detection method utilizing a long short-term memory (LSTM) encoder–decoder structure with regularized hidden dynamics and reversible instance normalization (RevIN) to compactly represent high-dimensional measurements for effective monitoring. During training, the hidden states of the model are regularized to form a low-dimensional latent space representation of the original multivariate time series data. As a result, the prediction errors of the latent states can be used to monitor the abnormal dynamic variations, while the reconstruction errors of the measured variables are used to monitor the abnormal static variations. Furthermore, the proposed indices can reflect operating conditions, even when the distribution of test data changes, which helps distinguish faults from normal adjustments and disturbances that controllers can settle. Here, data from numerical simulation and the Tennessee Eastman process are used to illustrate the effectiveness of the proposed fault detection method.

42 ENGINEERING↗

Enhancing the Operational Resilience of Advanced Reactors with Digital Twins by Recurrent Neural Networks

Because of a lack of operational data and uncertainty in evaluation model for abnormal and accident scenarios, the established operating procedures can be biased in characterizing the reactor states and ensuring operational resilience. To reduce uncertainty associated with actual plant conditions, digital twin (DT) technology is suggested to support operator’s decision-making by effectively extracting and using knowledge of the current and future plant states from the knowledge base. This study first builds a knowledge base based on the characterization of issue space and the simulation tool. Next, this study discusses diagnosis and prognosis DTs for enhancing operational resilience by recovering the complete states of reactors and by predicting the future reactor behaviors. Finally, the decision-making module of the control system can determine the optimal control strategy that meets operational goals during loss-of-flow scenarios. To demonstrate and evaluate the DTs capability for supporting the operations of nuclear reactors, this study develops and assesses both the diagnosis and prognosis DTs in a nearly autonomous management and control system for an Experimental Breeder Reactor-II simulator during different loss-of-flow scenarios.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

..delta..-Learning of High-Fidelity Electronic Structure Using Graph Neural Networks with Modified Node-Level Features

In this work, we present a ..delta..-learning approach for predicting the eigenvalues calculated with the hybrid functional HSE06 (..epsilon..nkHSE) for a set of metal and nitrogen doped graphene catalysts (MNCs) from Perdew-Burke-Ernzerhof (PBE) inputs. The model presented here incorporates electronic scalar features along with structural information in a graph neural network (GNN). In particular, the PBE eigenvalues for different bands and k-points and orbital-resolved projectors are combined with the applied potential as node-level features along with structural information within the Atomistic Line Graph Neural Network (ALIGNN) architecture. These features enable flexibility for systems with electrified interfaces, such as in electrocatalysts and achieves mean absolute error (MAE) of less than 0.1 eV. The machine learning model reported here achieves a strong generalization to left-out adsorbates (MAE = 0.074 eV) and leave-one-chemical-space-out (MAE = 0.08 eV) and completely left-out metals (MAE = 0.072 eV), confirming the robustness of the machine learning (ML) model in predicting ..epsilon..nkHSE.

36 MATERIALS SCIENCE↗

Machine‐learning‐based construction of barrier functions and models for safe model predictive control

Abstract In this paper, we propose a control Lyapunov‐barrier function‐based model predictive control method utilizing a feed‐forward neural network specified control barrier function (CBF) and a recurrent neural network (RNN) predictive model to stabilize nonlinear processes with input constraints, and to guarantee that safety requirements are met for all times. The nonlinear system is first modeled using RNN techniques, and a CBF is characterized by constructing a feed‐forward neural network (FNN) model with unique structures and properties. The FNN model for the CBF is trained based on data samples collected from safe and unsafe operating regions, and the resulting FNN model is verified to demonstrate that the safety properties of the CBF are satisfied. Given sufficiently small bounded modeling errors for both the FNN and the RNN models, the proposed control system is able to guarantee closed‐loop stability while preventing the closed‐loop states from entering unsafe regions in state‐space under sample‐and‐hold control action implementation. We provide the theoretical analysis for bounded unsafe sets in state‐space, and demonstrate the effectiveness of the proposed control strategy using a nonlinear chemical process example with a bounded unsafe region.

Chen, Scarlett↗

Distributed deep reinforcement learning for simulation control

Abstract Several applications in the scientific simulation of physical systems can be formulated as control/optimization problems. The computational models for such systems generally contain hyperparameters, which control solution fidelity and computational expense. The tuning of these parameters is non-trivial and the general approach is to manually ‘spot-check’ for good combinations. This is because optimal hyperparameter configuration search becomes intractable when the parameter space is large and when they may vary dynamically. To address this issue, we present a framework based on deep reinforcement learning (RL) to train a deep neural network agent that controls a model solve by varying parameters dynamically. First, we validate our RL framework for the problem of controlling chaos in chaotic systems by dynamically changing the parameters of the system. Subsequently, we illustrate the capabilities of our framework for accelerating the convergence of a steady-state computational fluid dynamics solver by automatically adjusting the relaxation factors of the discretized Navier–Stokes equations during run-time. The results indicate that the run-time control of the relaxation factors by the learned policy leads to a significant reduction in the number of iterations for convergence compared to the random selection of the relaxation factors. Our results point to potential benefits for learning adaptive hyperparameter learning strategies across different geometries and boundary conditions with implications for reduced computational campaign expenses 4 4 Data and codes available at https://github.com/Romit-Maulik/PAR-RL . .

42 ENGINEERING↗

Many-body expansion based machine learning models for octahedral transition metal complexes

Abstract Graph-based machine learning (ML) models for material properties show great potential to accelerate virtual high-throughput screening of large chemical spaces. However, in their simplest forms, graph-based models do not include any 3D information and are unable to distinguish stereoisomers such as those arising from different orderings of ligands around a metal center in coordination complexes. In this work we present a modification to revised autocorrelation descriptors, a molecular graph featurization method, for predicting spin state dependent properties of octahedral transition metal complexes (TMCs). Inspired by analytical semi-empirical models for TMCs, the new modeling strategy is based on the many-body expansion (MBE) and allows one to tune the captured stereoisomer information by changing the truncation order of the MBE. We present the necessary modifications to include this approach in two commonly used ML methods, kernel ridge regression and feed-forward neural networks. On a test set composed of all possible isomers of binary TMCs, the best MBE models achieve mean absolute errors (MAEs) of 2.75 kcal mol −1 on spin-splitting energies and 0.26 eV on frontier orbital energy gaps, a 30%–40% reduction in error compared to models based on our previous approach. We also observe improved generalization to previously unseen ligands where the best-performing models exhibit MAEs of 4.00 kcal mol −1 (i.e. a 0.73 kcal mol −1 reduction) on the spin-splitting energies and 0.53 eV (i.e. a 0.10 eV reduction) on the frontier orbital energy gaps. Because the new approach incorporates insights from electronic structure theory, such as ligand additivity relationships, these models exhibit systematic generalization from homoleptic to heteroleptic complexes, allowing for efficient screening of TMC search spaces.

Meyer, Ralf (ORCID:0000000322360261)↗