Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Recurrent neural networks”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

A novel transfer learning framework for sorghum biomass prediction using UAV-based remote sensing data and genetic markers

Yield for biofuel crops is measured in terms of biomass, so measurements throughout the growing season are crucial in breeding programs, yet traditionally time- and labor-consuming since they involve destructive sampling. Modern remote sensing platforms, such as unmanned aerial vehicles (UAVs), can carry multiple sensors and collect numerous phenotypic traits with efficient, non-invasive field surveys. However, modeling the complex relationships between the observed phenotypic traits and biomass remains a challenging task, as the ground reference data are very limited for each genotype in the breeding experiment. In this study, a Long Short-Term Memory (LSTM) based Recurrent Neural Network (RNN) model is proposed for sorghum biomass prediction. The architecture is designed to exploit the time series remote sensing and weather data, as well as static genotypic information. As a large number of features have been derived from the remote sensing data, feature importance analysis is conducted to identify and remove redundant features. A strategy to extract representative information from high-dimensional genetic markers is proposed. To enhance generalization and minimize the need for ground reference data, transfer learning strategies are proposed for selecting the most informative training samples from the target domain. Consequently, a pre-trained model can be refined with limited training samples. Field experiments were conducted over a sorghum breeding trial planted in multiple years with more than 600 testcross hybrids. The results show that the proposed LSTM-based RNN model can achieve high accuracies for single year prediction. Further, with the proposed transfer learning strategies, a pre-trained model can be refined with limited training samples from the target domain and predict biomass with an accuracy comparable to that from a trained-from-scratch model for both multiple experiments within a given year and across multiple years.

36 MATERIALS SCIENCE↗

Long–short-term memory encoder–decoder with regularized hidden dynamics for fault detection in industrial processes

The ability of recurrent neural networks (RNN) to model nonlinear dynamics of high dimensional process data has enabled data-driven RNN-based fault detection algorithms. Previous studies have focused on detecting faults by identifying the discrepancies in data distribution between the faulty and normal data, as reflected in prediction errors generated by RNN models. However, in industrial processes, variations in data distribution can also result from changes in normal control setpoints and compensatory control adjustments in response to disturbances, making it hard to differentiate between normal and faulty conditions. This paper proposes a fault detection method utilizing a long short-term memory (LSTM) encoder–decoder structure with regularized hidden dynamics and reversible instance normalization (RevIN) to compactly represent high-dimensional measurements for effective monitoring. During training, the hidden states of the model are regularized to form a low-dimensional latent space representation of the original multivariate time series data. As a result, the prediction errors of the latent states can be used to monitor the abnormal dynamic variations, while the reconstruction errors of the measured variables are used to monitor the abnormal static variations. Furthermore, the proposed indices can reflect operating conditions, even when the distribution of test data changes, which helps distinguish faults from normal adjustments and disturbances that controllers can settle. Here, data from numerical simulation and the Tennessee Eastman process are used to illustrate the effectiveness of the proposed fault detection method.

42 ENGINEERING↗

xesn: Echo state networks powered by Xarray and Dask

Xesn is a Python package that allows scientists to easily design Echo State Networks (ESNs) for forecasting problems. ESNs are a Recurrent Neural Network architecture introduced by Jaeger (2001) that are part of a class of techniques termed Reservoir Computing. One defining characteristic of these techniques is that all internal weights are determined by a handful of global, scalar parameters, thereby avoiding problems during backpropagation and reducing training time significantly. Because this architecture is conceptually simple, many scientists implement ESNs from scratch, leading to questions about computational performance. Xesn offers a straightforward, standard implementation of ESNs that operates efficiently on CPU and GPU hardware. The package leverages optimization tools to automate the parameter selection process, so that scientists can reduce the time finding a good architecture and focus on using ESNs for their domain application. Importantly, the package flexibly handles forecasting tasks for out-of-core, multi-dimensional datasets, eliminating the need to write parallel programming code. Xesn was initially developed to handle the problem of forecasting weather dynamics, and so it integrates naturally with Python packages that have become familiar to weather and climate scientists such as Xarray (Hoyer & Hamman, 2017). However, the software is ultimately general enough to be utilized in other domains where ESNs have been useful, such as in signal processing (Jaeger & Haas, 2004).

97 MATHEMATICS AND COMPUTING↗

Online evolutionary neural architecture search for multivariate non-stationary time series forecasting

Time series forecasting (TSF) is one of the most important tasks in data science. TSF models are usually pre-trained with historical data and then applied on future unseen datapoints. However, real-world time series data is usually non-stationary and models trained offline usually face problems from data drift. Models trained and designed in an offline fashion can not quickly adapt to changes quickly or be deployed in real-time. To address these issues, this work presents the Online NeuroEvolution-based Neural Architecture Search (ONE-NAS) algorithm, which is a novel neural architecture search method capable of automatically designing and dynamically training recurrent neural networks (RNNs) for online forecasting tasks. Without any pre-training, ONE-NAS utilizes populations of RNNs that are continuously updated with new network structures and weights in response to new multivariate input data. ONE-NAS is tested on real-world, large-scale multivariate wind turbine data as well as the univariate Dow Jones Industrial Average (DJIA) dataset. These results demonstrate that ONE-NAS outperforms traditional statistical time series forecasting methods, including online linear regression, fixed long short-term memory (LSTM) and gated recurrent unit (GRU) models trained online, as well as state-of-the-art, online ARIMA strategies. Additionally, results show that utilizing multiple populations of RNNs which are periodically repopulated provide significant performance improvements, allowing this online neural network architecture design and training to be successful.

97 MATHEMATICS AND COMPUTING↗

Trainable Gene Regulation Networks with Applications to Drosophila Pattern Formation

This chapter will very briefly introduce and review some computational experiments in using trainable gene regulation network models to simulate and understand selected episodes in the development of the fruit fly, Drosophila melanogaster. For details the reader is referred to the papers introduced below. It will then introduce a new gene regulation network model which can describe promoter-level substructure in gene regulation. As described in chapter 2, gene regulation may be thought of as a combination of cis-acting regulation by the extended promoter of a gene (including all regulatory sequences) by way of the transcription complex, and of trans-acting regulation by the transcription factor products of other genes. If we simplify the cis-action by using a phenomenological model which can be tuned to data, such as a unit or other small portion of an artificial neural network, then the full transacting interaction between multiple genes during development can be modelled as a larger network which can again be tuned or trained to data. The larger network will in general need to have recurrent (feedback) connections since at least some real gene regulation networks do. This is the basic modeling approach taken, which describes how a set of recurrent neural networks can be used as a modeling language for multiple developmental processes including gene regulation within a single cell, cell-cell communication, and cell division. Such network models have been called "gene circuits", "gene regulation networks", or "genetic regulatory networks", sometimes without distinguishing the models from the actual modeled systems.

Mjolsness, Eric↗

Dynamical Large Deviations of Two-Dimensional Kinetically Constrained Models Using a Neural-Network State Ansatz

We use a neural-network ansatz originally designed for the variational optimization of quantum systems to study dynamical large deviations in classical ones. We use recurrent neural networks to describe the large deviations of the dynamical activity of model glasses, kinetically constrained models in two dimensions. We present the first finite size-scaling analysis of the large-deviation functions of the two-dimensional Fredrickson-Andersen model, and explore the spatial structure of the high-activity sector of the South-or-East model. These results provide a new route to the study of dynamical large-deviation functions, and highlight the broad applicability of the neural-network state ansatz across domains in physics.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Traffic Signal Control With Adaptive Online-Learning Scheme Using Multiple-Model Neural Networks

This article proposes a new traffic signal control algorithm to deal with unknown-traffic-system uncertainties and reduce delays in vehicle travel time. Unknown-traffic-system dynamics are approximated using a recurrent neural network (NN). To accurately identify the traffic system model, an online-learning scheme is developed to switch among a set of candidate NNs (i.e., multiple-model NNs) based on their estimation errors. Then, a bank of optimal signal-timing controllers is designed based on the online identification of the traffic system. Simulation studies have been carried out for the obtained control strategies using multiple-model NNs, and the desired results have been obtained. Moreover, compared with the widely used actuated traffic signal control schemes, it is shown that the proposed method can reduce vehicle travel delays and improve traffic system robustness.

99 GENERAL AND MISCELLANEOUS↗

Explainable machine learning of the underlying physics of high-energy particle collisions

We present an implementation of an explainable and physics-aware machine learning model capable of inferring the underlying physics of high-energy particle collisions using the information encoded in the energy-momentum four-vectors of the final state particles. We demonstrate the proof-of-concept of our White Box AI approach using a Generative Adversarial Network (GAN) which learns from a DGLAP-based parton shower Monte Carlo event generator. The constrained generator network architecture mimics the structure of a parton shower exhibiting similarities with Recurrent Neural Networks (RNNs). We show, for the first time, that our approach leads to a network that is able to learn not only the final distribution of particles, but also the underlying parton branching mechanism, i.e. the Altarelli-Parisi splitting function, the ordering variable of the shower, and the scaling behavior. While the current work is focused on perturbative physics of the parton shower, we foresee a broad range of applications of our framework to areas that are currently difficult to address from first principles in QCD. Examples include nonperturbative and collective effects, factorization breaking and the modification of the parton shower in heavy-ion, and electron-nucleus collisions.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Forward variable selection enables fast and accurate dynamic system identification with Karhunen-Loève decomposed Gaussian processes

A promising approach for scalable Gaussian processes (GPs) is the Karhunen-Loève (KL) decomposition, in which the GP kernel is represented by a set of basis functions which are the eigenfunctions of the kernel operator. Such decomposed kernels have the potential to be very fast, and do not depend on the selection of a reduced set of inducing points. However KL decompositions lead to high dimensionality, and variable selection thus becomes paramount. This paper reports a new method of forward variable selection, enabled by the ordered nature of the basis functions in the KL expansion of the Bayesian Smoothing Spline ANOVA kernel (BSS-ANOVA), coupled with fast Gibbs sampling in a fully Bayesian approach. It quickly and effectively limits the number of terms, yielding a method with competitive accuracies, training and inference times for tabular datasets of low feature set dimensionality. Theoretical computational complexities are O ( N P 2 ) in training and O ( P ) per point in inference, where N is the number of instances and P the number of expansion terms. The inference speed and accuracy makes the method especially useful for dynamic systems identification, by modeling the dynamics in the tangent space as a static problem, then integrating the learned dynamics using a high-order scheme. The methods are demonstrated on two dynamic datasets: a ‘Susceptible, Infected, Recovered’ (SIR) toy problem, along with the experimental ‘Cascaded Tanks’ benchmark dataset. Comparisons on the static prediction of time derivatives are made with a random forest (RF), a residual neural network (ResNet), and the Orthogonal Additive Kernel (OAK) inducing points scalable GP, while for the timeseries prediction comparisons are made with LSTM and GRU recurrent neural networks (RNNs) along with the SINDy package.

Hayes, Kyle↗

Spatio-Temporal Anomaly Detection with Graph Networks for Data Quality Monitoring of the Hadron Calorimeter

The Compact Muon Solenoid (CMS) experiment is a general-purpose detector for high-energy collision at the Large Hadron Collider (LHC) at CERN. It employs an online data quality monitoring (DQM) system to promptly spot and diagnose particle data acquisition problems to avoid data quality loss. In this study, we present a semi-supervised spatio-temporal anomaly detection (AD) monitoring system for the physics particle reading channels of the Hadron Calorimeter (HCAL) of the CMS using three-dimensional digi-occupancy map data of the DQM. We propose the GraphSTAD system, which employs convolutional and graph neural networks to learn local spatial characteristics induced by particles traversing the detector and the global behavior owing to shared backend circuit connections and housing boxes of the channels, respectively. Recurrent neural networks capture the temporal evolution of the extracted spatial features. We validate the accuracy of the proposed AD system in capturing diverse channel fault types using the LHC collision data sets. The GraphSTAD system achieves production-level accuracy and is being integrated into the CMS core production system for real-time monitoring of the HCAL. We provide a quantitative performance comparison with alternative benchmark models to demonstrate the promising leverage of the presented system.

43 PARTICLE ACCELERATORS↗

Deep Learning Systems for Increased Safeguards Surveillance Review Productivity

Nuclear safeguards inspectors expend significant time and maintain intense focus in reviewing video surveillance for safeguards relevant events. To increase efficiency and reduce the time burden of safeguards inspectors performing surveillance data review, this paper presents a novel deep learning (DL) systems concept to integrate generalized DL models into the safeguards surveillance review workflow. The Agency is investigating several DL algorithms for object recognition, localization, tracking, and flagging relevant activities. The project team is working closely with nuclear safeguards inspectors to identify review use cases (based on specific safeguards objectives) and collect their associated requirements. We focused on CANDU and LWR Nuclear Power Plants (NPPs) and their associated dry storage areas as these present a particularly heavy burden on the inspector surveillance review process due to the number of these facilities under safeguards worldwide. Initial DL algorithm results on safeguards data are promising. Using a convolutional neural network, the team attained a mean average precision (mAP) of 92.9% identifying and localizing spent fuel (SF) casks from a 475 surveillance image dataset. Further, the team had initial success in training a recurrent neural network to identify reactor area activities in video clips, successfully indicating when SF casks enter or exit a pool. We discuss how such DL algorithms would be integrated into the Next Generation Surveillance Review (NGSR) software application. Another issue impacting review productivity is the long time inspectors may have to wait when running these algorithms in NGSR. We present a concept to pre-process remotely collected surveillance data with DL models as the data arrives to IAEA headquarters so that results are already available when starting a new review in NGSR. The proposed DL system concept shows a pathway and workflow for increasing an inspector’s surveillance review productivity by quickly and accurately identifying declared and undeclared safeguards relevant objects and activities in large quantities of surveillance imagery data.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

Neural Networks for Rapid Design and Analysis

Artificial neural networks have been employed for rapid and efficient dynamics and control analysis of flexible systems. Specifically, feedforward neural networks are designed to approximate nonlinear dynamic components over prescribed input ranges, and are used in simulations as a means to speed up the overall time response analysis process. To capture the recursive nature of dynamic components with artificial neural networks, recurrent networks, which use state feedback with the appropriate number of time delays, as inputs to the networks, are employed. Once properly trained, neural networks can give very good approximations to nonlinear dynamic components, and by their judicious use in simulations, allow the analyst the potential to speed up the analysis process considerably. To illustrate this potential speed up, an existing simulation model of a spacecraft reaction wheel system is executed, first conventionally, and then with an artificial neural network in place.

Sparks, Dean W., Jr.↗

Technical note: Using long short-term memory models to fill data gaps in hydrological monitoring networks

Abstract. Quantifying the spatiotemporal dynamics in subsurface hydrological flows over a long time window usually employs a network of monitoring wells. However, such observations are often spatially sparse with potential temporal gaps due to poor quality or instrument failure. In this study, we explore the ability of recurrent neural networks to fill gaps in a spatially distributed time-series dataset. We use a well network that monitors the dynamic and heterogeneous hydrologic exchanges between the Columbia River and its adjacent groundwater aquifer at the U.S. Department of Energy's Hanford site. This 10-year-long dataset contains hourly temperature, specific conductance, and groundwater table elevation measurements from 42 wells with gaps of various lengths. We employ a long short-term memory (LSTM) model to capture the temporal variations in the observed system behaviors needed for gap filling. The performance of the LSTM-based gap-filling method was evaluated against a traditional autoregressive integrated moving average (ARIMA) method in terms of error statistics and accuracy in capturing the temporal patterns of river corridor wells with various dynamics signatures. Our study demonstrates that the ARIMA models yield better average error statistics, although they tend to have larger errors during time windows with abrupt changes or high-frequency (daily and subdaily) variations. The LSTM-based models excel in capturing both high-frequency and low-frequency (monthly and seasonal) dynamics. However, the inclusion of high-frequency fluctuations may also lead to overly dynamic predictions in time windows that lack such fluctuations. The LSTM can take advantage of the spatial information from neighboring wells to improve the gap-filling accuracy, especially for long gaps in system states that vary at subdaily scales. While LSTM models require substantial training data and have limited extrapolation power beyond the conditions represented in the training data, they afford great flexibility to account for the spatial correlations, temporal correlations, and nonlinearity in data without a priori assumptions. Thus, LSTMs provide effective alternatives to fill in data gaps in spatially distributed time-series observations characterized by multiple dominant frequencies of variability, which are essential for advancing our understanding of dynamic complex systems.

54 ENVIRONMENTAL SCIENCES↗

TorchBraid: High-Performance Layer-Parallel Training of Deep Neural Networks with MPI and GPU Acceleration

TorchBraid is a high-performance implementation of layer-parallel training for deep neural networks (DNNs) supporting MPI-based parallelism and GPU acceleration. Layer-parallel training has been developed to overcome the serialization inherent in forward and backward propagation of DNNs that limits utilization of computational resources in the strong scaling limit. To achieve this, TorchBraid integrates the PyTorch neural network framework with the state-of-the-art XBraid time-parallel library. Furthermore, this article presents the use and performance of TorchBraid, in addition to solutions for overcoming the algorithmic challenges inherent in combining automatic differentiation with layer-parallel. Results are presented with and without GPU acceleration for the Tiny ImageNet and MNIST image classification data sets, as well as recurrent neural networks. Overall, TorchBraid enables fast training of DNNs, both in a strong and weak scaling context. In addition to the TorchBraid software, several new advances in applying layer-parallel algorithms are detailed. Integration of layer-parallel with data-parallel algorithms is presented for the first time, showing the computational advantages of the combination. Standard deep learning techniques, like batch-normalization, are developed for layer-parallel training. Finally, a new approach combining layer-parallel with spatial coarsening in order to accelerate training for 3D image classification shows roughly a 10× speedup over serial execution.

Layer-parallel↗

Data-driven based coordinated smart inverter control for distributed energy resources

Smart inverters (SI) for distributed energy resources (DER) are becoming popular since they have the ability to stabilize as well as restore the voltage and frequency of power systems. Aiming at establishing the mathematical models combined with SI control methods, multiple optimization methods are developed. However, the computational complexity of solving such a mathematical model with various uncertainties limits the real-time application of the SI control. To conquer this challenge, a data-driven-based SI control approach is developed to achieve coordinated control in the high penetration DER system. First, an optimization problem for maximizing the active power generation and minimizing the power loss is designed using the Volt/VAR control. To reduce the time consumption, the recurrent neural network (RNN) is proposed to model the relationship between the uncertainties and control actions during the offline site. The RNN with different sub-structures such as the long short-term memory cell and gated recurrent unit cell are included to enrich the diversity of features. In the last stage, different experiment comparisons, including multiple uncertainties maps and stateof- art machine learning methods, are conducted to verify the effectiveness of the proposed method based on the IEEE 123 bus power system. The results demonstrate that the proposed method can effectively achieve a rapid and coordinated control with a lower error rate.

Qiu, Wei↗

Watch and learn—a generalized approach for transferrable learning in deep neural networks via physical principles

Transfer learning refers to the use of knowledge gained while solving a machine learning task and applying it to the solution of a closely related problem. Such an approach has enabled scientific breakthroughs in computer vision and natural language processing where the weights learned in state-of-the-art models can be used to initialize models for other tasks which dramatically improve their performance and save computational time. Here we demonstrate an unsupervised learning approach augmented with basic physical principles that achieves fully transferrable learning for problems in statistical physics across different physical regimes. By coupling a sequence model based on a recurrent neural network to an extensive deep neural network, we are able to learn the equilibrium probability distributions and inter-particle interaction models of classical statistical mechanical systems. Our approach, distribution-consistent learning, DCL, is a general strategy that works for a variety of canonical statistical mechanical models (Ising and Potts) as well as disordered interaction potentials. Using data collected from a single set of observation conditions, DCL successfully extrapolates across all temperatures, thermodynamic phases, and can be applied to different length-scales. This constitutes a fully transferrable physics-based learning in a generalizable approach.

97 MATHEMATICS AND COMPUTING↗

Upscaling Reactive Transport and Clogging in Shale Microcracks by Deep Learning

Fracture networks in shales exhibit multiscale features. A rock system may contain a few main fractures and thousands of microcracks, whose length and aperture are orders of magnitude smaller than the former. It is computationally prohibitive to resolve all the fractures explicitly for such multiscale fracture networks. One traditional approach is to model the small-scale features (e.g., microcracks in shales) as an effective medium. Although this fracture-matrix conceptualization significantly reduces the problem complexity, there are classes of physical processes that cannot be accurately upscaled by effective medium approximations, e.g., microcrack clogging during mineral reactions. Here, we employ deep learning in place of effective medium theory to upscale physical processes in small-scale features. Specifically, we consider reactive transport in a fracture-microcrack network where microcracks can be clogged by precipitation. A deep learning multiscale algorithm is developed, in which the microcracks are upscaled as a wall boundary condition of the main fractures. The wall boundary condition is constructed by recurrent neural networks, which take concentration histories as input and predict the solute transport from main fractures to microcracks. The deep learning multiscale algorithm is firstly employed in specific scenarios, then a general model is developed which can work under various conditions. The new approach is validated against fully resolved simulations and an analytical solution, providing a reliable and efficient solution for problems that cannot be upscaled by effective medium models.

58 GEOSCIENCES↗