SEARCH · Engineering Papers
Results for “Recurrent neural networks”
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Context Modulation Enables Model Robustness to Weight Noise in Recurrent Spiking Neural Networks
Explore the source record for details and available documents.
Deep Learning Based Superconducting Radio-Frequency Cavity Fault Classification at Jefferson Laboratory
This work investigates the efficacy of deep learning (DL) for classifying C100 superconducting radio-frequency (SRF) cavity faults in the Continuous Electron Beam Accelerator Facility (CEBAF) at Jefferson Lab. CEBAF is a large, high-power continuous wave recirculating linac that utilizes 418 SRF cavities to accelerate electrons up to 12 GeV. Recent upgrades to CEBAF include installation of 11 new cryomodules (88 cavities) equipped with a low-level RF system that records RF time-series data from each cavity at the onset of an RF failure. Typically, subject matter experts (SME) analyze this data to determine the fault type and identify the cavity of origin. This information is subsequently utilized to identify failure trends and to implement corrective measures on the offending cavity. Manual inspection of large-scale, time-series data, generated by frequent system failures is tedious and time consuming, and thereby motivates the use of machine learning (ML) to automate the task. This study extends work on a previously developed system based on traditional ML methods (Tennant and Carpenter and Powers and Shabalina Solopova and Vidyaratne and Iftekharuddin, Phys. Rev. Accel. Beams, 2020, 23, 114601), and investigates the effectiveness of deep learning approaches. The transition to a DL model is driven by the goal of developing a system with sufficiently fast inference that it could be used to predict a fault event and take actionable information before the onset (on the order of a few hundred milliseconds). Because features are learned, rather than explicitly computed, DL offers a potential advantage over traditional ML. Specifically, two seminal DL architecture types are explored: deep recurrent neural networks (RNN) and deep convolutional neural networks (CNN). We provide a detailed analysis on the performance of individual models using an RF waveform dataset built from past operational runs of CEBAF. In particular, the performance of RNN models incorporating long short-term memory (LSTM) are analyzed along with the CNN performance. Furthermore, comparing these DL models with a state-of-the-art fault ML model shows that DL architectures obtain similar performance for cavity identification, do not perform quite as well for fault classification, but provide an advantage in inference speed.
Detecting Spacecraft Anomalies Using LSTMs and Nonparametric Dynamic Thresholding
As spacecraft send back increasing amounts of telemetry data, improved anomaly detection systems are needed to lessen the monitoring burden placed on operations engineers and reduce operational risk. Current spacecraft monitoring systems only target a subset of anomaly types and often require costly expert knowledge to develop and maintain due to challenges involving scale and complexity. We demonstrate the effectiveness of Long Short-Term Memory (LSTMs) networks, a type of Recurrent Neural Network (RNN), in overcoming these issues using expert-labeled telemetry anomaly data from the Soil Moisture Active Passive (SMAP) satellite and the Mars Science Laboratory (MSL) rover, Curiosity. We also propose a complementary unsupervised and nonparametric anomaly thresholding approach developed during a pilot implementation of an anomaly detection system for SMAP, and offer false positive mitigation strategies along with other key improvements and lessons learned during development.
Physics-informed machine learning modeling for predictive control using noisy data
Due to the occurrence of over-fitting at the learning phase, the modeling of chemical processes via artificial neural networks (ANN) by using corrupted data (i.e., noisy data) is an ongoing challenge. Therefore, this work investigates the effect of both Gaussian and non-Gaussian noise on the performance of process-structure based recurrent neural networks (RNN) models, which take the form of partially-connected RNN models in this work, that are used to approximate a class of multi-input-multi-outputs nonlinear systems. Furthermore, two different techniques, specifically Monte Carlo dropout and co-teaching, are utilized in the development of partially-connected RNN models. Here, these two techniques are employed to reduce the over-fitting in ANNs when noisy data is used in the training process and, hence, to improve the open-loop accuracy as well as the closed-loop performance under a Lyapunov-based model predictive controller (MPC). Aspen Plus Dynamics, a well-known high-fidelity process simulator, is used to simulate a large-scale chemical process application in order to demonstrate the anticipated improvements in both open-loop approximation and closed-loop controller performance in the presence of Gaussian and non-Gaussian noise in the data set using physics-informed RNNs.
Learning-Accelerated ADMM for Distributed DC Optimal Power Flow
We suggest a novel data-driven method to accelerate the convergence of Alternating Direction Method of Multipliers (ADMM) for solving distributed DC optimal power flow (DC-OPF) where lines are shared between independent network partitions. Using previous observations of ADMM trajectories for a given system under varying load, the method trains a recurrent neural network (RNN) to predict the converged values of dual and consensus variables. Given a new realization of system load, a small number of initial ADMM iterations is taken as input to infer the converged values and directly inject them into the iteration. We empirically demonstrate that the online injection of these values into the ADMM iteration accelerates convergence by a significant factor for partitioned 14-, 118-and 2848-bus test systems under differing load scenarios. The proposed method has several advantages: it maintains the security of private decision variables inherent in consensus ADMM; inference is fast and so may be used in online settings; RNN-generated predictions can dramatically improve time to convergence but, by construction, can never result in infeasible ADMM subproblems; it can be easily integrated into existing software implementations. While we focus on the ADMM formulation of distributed DC-OPF in this paper, the ideas presented are naturally extended to other distributed optimization problems.
Learning-Accelerated ADMM for Distributed DC Optimal Power Flow
We propose a novel data-driven method to accelerate the convergence of Alternating Direction Method of Multipliers (ADMM) for solving distributed DC optimal power flow (DC-OPF) where lines are shared between independent network partitions. Using previous observations of ADMM trajectories for a given system under varying load, the method trains a recurrent neural network (RNN) to predict the converged values of dual and consensus variables. Given a new realization of system load, a small number of initial ADMM iterations is taken as input to infer the converged values and directly inject them into the iteration. We empirically demonstrate that the online injection of these values into the ADMM iteration accelerates convergence by a significant factor for partitioned 14-, 118- and 2848-bus test systems under differing load scenarios. The proposed method has several advantages: it maintains the security of private decision variables inherent in consensus ADMM; inference is fast and so may be used in online settings; RNN-generated predictions can dramatically improve time to convergence but, by construction, can never result in infeasible ADMM subproblems; it can be easily integrated into existing software implementations. While we focus on the ADMM formulation of distributed DC-OPF in this paper, the ideas presented are naturally extended to other distributed optimization problems.
Using machine learning for particle track identification in the CLAS12 detector
Particle track reconstruction is the most computationally intensive process in nuclear physics experiments. Traditional algorithms use a combinatorial approach that exhaustively tests track measurements ("hits") to identify those that form an actual particle trajectory. In this article, we describe the development of four machine learning (ML) models that assist the tracking algorithm by identifying valid track candidates from the measurements in drift chambers. Several types of machine learning models were tested, including: Convolutional Neural Networks (CNN), Multi-Layer Perceptrons (MLP), Extremely Randomized Trees (ERT) and Recurrent Neural Networks (RNN). As a result of this work, an MLP network classifier was implemented as part of the CLAS12 reconstruction software to provide the tracking code with recommended track candidates. The resulting software achieved accuracy of greater than 99% and resulted in an end-to-end speedup of 35% compared to existing algorithms.
A Survey: Handling Irregularities in Neural Network Acceleration with FPGAs
In the last decade, Artificial Intelligence (AI) through Deep Neural Networks (DNNs) has penetrated virtually every aspect of science, technology, and business. Many types of DNNs have been and continue to be developed, including Convolutional Neural Networks (CNNs), Recurrent Neural Net- works (RNNs), and Graph Neural Networks (GNNs). The overall problem for all of these Neural Networks (NNs) is that their target applications generally pose stringent constraints on latency and throughput, while also having strict accuracy requirements. There have been many previous efforts in creating hardware to accelerate NNs. The problem designers face is that optimal NN models typically have significant irregularities, making them hardware-unfriendly. In this paper, we first define the problems in NN acceleration by characterizing common irregularities in NN processing into 4 types; then we summarize the existing works that handle the four types of irregularities efficiently using hardware, especially FPGAs; finally, we provide a new vision of next-generation FPGA-based NN acceleration: that the emerging heterogeneity in the next-generation FPGAs is the key to achieving higher performance.
Predictive Skill of Deep Learning Models Trained on Limited Sequence Data
In this report we investigate the utility of one-dimensional convolutional neural network (CNN) models in epidemiological forecasting. Deep learning models, especially variants of recurrent neural networks (RNNs) have been studied for influenza forecasting, and have achieved higher forecasting skill compared to conventional models such as ARIMA models. In this study, we adapt two neural networks that employ one-dimensional temporal convolutional layers as a primary building block temporal convolutional networks and simple neural attentive meta-learner for epidemiological forecasting and test them with influenza data from the US collected over 2010-2019. We find that epidemiological forecasting with CNNs is feasible, and their forecasting skill is comparable to, and at times, superior to, RNNs. Thus CNNs and RNNs bring the power of nonlinear transformations to purely data-driven epidemiological models, a capability that heretofore has been limited to more elaborate mechanistic/compartmental disease models.
Development of a Machine-Learned Cruise Guide Indicator for Rotorcraft
This paper presents a machine-learned virtual cruise guide indicator (vCGI) for Chinook helicopters. Two temporal neural networks were trained and evaluated on measured data from 55 flight tests, one for the fore rotor and another for the aft rotor, to predict a vCGI value, which protects 23 components from fatigue damage during steady-state conditions. Three different classes of machine learning architectures were evaluated for prediction of the vCGI from time sequences: a temporal convolutional neural network with 1D dilated causal convolutions, a long short-term memory recurrent neural network, and an attention-based transformer architecture. The final average model accuracy on unseen flight data is currently greater than 93% for CGI values which could result in fatigue damage and 90% for normal operation CGI values. Model accuracy was improved through a series of advancements in:(1) selection of optimal training data using temporal collective variables and unsupervised learning, (2) dataset augmentation with maximum-entropy temporal collective variables, and (3) implementation of a mixture-of-experts classification- regression approach using an adversarial classification approach to assign maneuver labels. The results are presented for each advancement in model development along with lessons learned in training machine learning models on real- world, time-dependent rotorcraft data.
Deployment of Dynamic Neural Network Optimization to Minimize Heat Rate During Ramping for Coal Power Plants (Final Technical Report)
Much success was achieved throughout the course of this project. A successful implementation of Dynamic Neural Network Optimization (D-NNO) was coupled with Adaptive Predictive Controls (APC) and a novel hardware installation comprised of an advanced sensor network (ASN) measuring mass-weighted averages of flue gas constituents above the horizontal superheater of a coal-fired utility boiler. From 2019 through 2023 (including an extension due to COVID delays), the team was able to prototype, evaluate, deploy, iterate, and ultimately finalize an advanced closed-loop control D-NNO system which demonstrated the ability to: •improve unit efficiency ~2.0% relative to unoptimized operation (represented as total fuel fired per MWh generated) •improve unit NOx emission rates 10%+ beyond static optimization baselines •improve unit temperature stability as much as 58% and on average 12% •improve operating load stability as much as 35% The culmination of this project has generated an advanced methodology of deploying specially designed recurrent neural networks (long short-term memory, gated recurrent unit, encoder-decoder networks, transformers, etc.), customized trajectory planning and closed-loop optimization modules capable of adapting to live electric grid responses and demands, self-tuning and adaptive expert controls constantly adjusting prediction parameters to real-time unit behavior, and a hardware/software package able to reliably calculate net unit heat rate (NUHR) in real-time using flue gas constituents, machine learning, and known combustion relationships. Through this real-time NUHR value, immediate feedback on system adjustments relative to operating efficiency was available, allowing for rapid improvements to system performance. In addition to development and deployment of the advanced D-NNO system, the approach methodology has been readily commercialized through the project platform Griffin Open Systems, LLC, the D-NNO software platform host. Similar methodologies to those developed by this project have already been deployed at 5 other units across the United States, with another 6 implementations scheduled, and more expected. Over the course of the project, multiple academic papers were submitted and accepted for publication within esteemed academic journals, and PhD students were trained and graduated, as well as undergraduate students becoming involved and participating to project objectives.
Utilizing Physics-Informed Synthetic Data to Train a Digital Twin for Predicting Reactor Operations
Understanding techniques to strengthen the nuclear nonproliferation regime is crucial in reducing the creation of nuclear weaponry on the basis of advancements in the nuclear energy industry. Prior to construction of a nuclear power plant, it is necessary to understand the proliferation potential of the plant’s reactor. Digital twins serve as a unique solution to recognizing reactor behavior indicative of nuclear proliferation. The following research conducted serves as a validation to training a digital twin on synthetic data fabricated via means of reactor physics simulations based on parameters of Idaho State University’s AGN-201 reactor. The synthetic data is utilized to train a long short-term memory (LSTM) recurrent neural network model. The accuracy of the predicted data is measured against real operational data to verify the reliability of the synthetic data creation methods and if these methods should be used in the future to inform inspectors of a reactor’s proliferation capabilities.
Variational data augmentation for a learning-based granular predictive model of power outages
As the trend in climate change continues, extreme weather events are expected to occur with increasing frequency and severity and pose a significant threat to the electric power infrastructure. Regardless of the efforts a utility puts towards hardening the grid, storm-induced damage to the utility assets such as cables and distributed energy resources (DERs) that are particularly vulnerable to such events is unavoidable. Access to a highly granular, in space and time, outage forecasting tool with long lead times (i.e., days ahead) will enhance the efficiency of service restoration efforts. Here, in this study, we propose to develop and implement a multi-model framework as an operational tool based on a granular and multi-day outage forecasting model using operational numerical weather prediction model forecasts and detailed component outage information. An innovative two-layered recurrent neural network, i.e., a long-short-term-memory (LSTM)-based variational autoencoder (VAE) framework and a sliding window are used to address the uneven distribution of different types of weather events and make better use of the time-series data. Case studies are performed to demonstrate the performance of the new framework.
Towards a robust out-of-the-box neural network model for genomic data
The accurate prediction of biological features from genomic data is paramount for precision medicine and sustainable agriculture. For decades, neural network models have been widely popular in fields like computer vision, astrophysics and targeted marketing given their prediction accuracy and their robust performance under big data settings. Yet neural network models have not made a successful transition into the medical and biological world due to the ubiquitous characteristics of biological data such as modest sample sizes, sparsity, and extreme heterogeneity. Here, we investigate the robustness, generalization potential and prediction accuracy of widely used convolutional neural network and natural language processing models with a variety of heterogeneous genomic datasets. Mainly, recurrent neural network models outperform convolutional neural network models in terms of prediction accuracy, overfitting and transferability across the datasets under study. While the perspective of a robust out-of-the-box neural network model is out of reach, we identify certain model characteristics that translate well across datasets and could serve as a baseline model for translational researchers.
Reduced Order Model of Transactive Bidding Loads
Transactive energy (TE) has been identified to provide better grid efficiency and reliability by market-based transactive exchanges between energy producers and energy consumers. Simulations of TE systems are crucial to evaluate the benefits and impacts of different transactive mechanisms. However, such simulations can be time consuming due to the information exchange between various participants and complex co-simulation environments. In this paper, we develop a reduced order model to speed up the simulation of transactive systems in TE simulation platform (TESP) while achieving very low error between the reduced order and full model results. Specifically, the developed reduced order model consists of an aggregate responsive load agent which utilizes two Recurrent Neural Networks (RNNs) with Long Short-Term Memory units (LSTMs) to enable transactive elements to collectively participate in the TE system. The proposed aggregate responsive load (ARL) agent is able to produce similar transactive behaviors to the full simulation model while achieving significant simulation time reduction. Finally, we also show that the developed model enables generalization of simulation results across different dates and across different number of loads included in the simulations.
Physics-Informed Machine Learning Model for Ceramic Matrix Composite Creep
A physics-informed recurrent neural network (RNN) based surrogate model is developed to emulate the nonlinear, time-dependent constitutive behavior of ceramic matrix composites (CMCs) driven by matrix damage and constituent creep at the microscale. Physics-informed constraints are introduced into the surrogate model through regularization to ground the prediction in physics and improve its predictive capabilities. Training data is generated using the high-fidelity generalized method of cells (HFGMC) approach which calls appropriate creep and damage models for each of the constituents. This coupling permits simulating the nonlinear behavior of CMCs based on constituent response at the microscale along with microstructural features such as fiber and porosity volume fraction and fiber radius. The microscale repeating unit cell is loaded under creep fatigue conditions to replicate the material loading experienced in a turbine engine. Therefore, the RNN-based surrogate model is tasked with predicting, as a function of variable input stress sequence, temperature, and microstructural features, the resulting strain history response while satisfying physical constraints related to creep rate, isochoric inelastic deformation, and strain energy density. The trained surrogate model is shown to effectively match the strain history over quantified distributions of microstructural features and relevant loading regimes and temperatures. Neural network based surrogate models can offer efficient alternatives to running computationally intensive multiscale material models to simulate the nonlinear response of large structural models. Therefore, the presented work provides evidence towards the feasibility of developing, training, and running such models for CMCs with complex microstructures, nonlinear time-dependent material response, and under non-monotonic loading conditions.
Reconstruction of fast neutron direction in segmented organic detectors using deep learning
A method for reconstructing the direction of a fast neutron source using a segmented organic scintillator-based detector and deep learning model is proposed and analyzed. Here, the model is based on recurrent neural network, which can be trained by a sequence of data obtained from an event recorded in the detector and suitably pre-processed. The performance of deep learning-based model is compared with the conventional double-scatter detection algorithm in reconstructing the direction of a fast neutron source. With the deep learning model, the uncertainty in source direction of 0.301 rad is achieved with 100 neutron detection events in a segmented cubic organic scintillator detector with a side length of 46 mm. To reconstruct the source direction with the same angular resolution as the double-scatter algorithm, the deep learning method requires 75% fewer events. Application of this method could augment the operation of segmented detectors operated in the neutron scatter camera configuration for applications such as special nuclear material detection.