Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Recurrent neural network”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Learning-Accelerated ADMM for Distributed DC Optimal Power Flow

We suggest a novel data-driven method to accelerate the convergence of Alternating Direction Method of Multipliers (ADMM) for solving distributed DC optimal power flow (DC-OPF) where lines are shared between independent network partitions. Using previous observations of ADMM trajectories for a given system under varying load, the method trains a recurrent neural network (RNN) to predict the converged values of dual and consensus variables. Given a new realization of system load, a small number of initial ADMM iterations is taken as input to infer the converged values and directly inject them into the iteration. We empirically demonstrate that the online injection of these values into the ADMM iteration accelerates convergence by a significant factor for partitioned 14-, 118-and 2848-bus test systems under differing load scenarios. The proposed method has several advantages: it maintains the security of private decision variables inherent in consensus ADMM; inference is fast and so may be used in online settings; RNN-generated predictions can dramatically improve time to convergence but, by construction, can never result in infeasible ADMM subproblems; it can be easily integrated into existing software implementations. While we focus on the ADMM formulation of distributed DC-OPF in this paper, the ideas presented are naturally extended to other distributed optimization problems.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Learning-Accelerated ADMM for Distributed DC Optimal Power Flow

We propose a novel data-driven method to accelerate the convergence of Alternating Direction Method of Multipliers (ADMM) for solving distributed DC optimal power flow (DC-OPF) where lines are shared between independent network partitions. Using previous observations of ADMM trajectories for a given system under varying load, the method trains a recurrent neural network (RNN) to predict the converged values of dual and consensus variables. Given a new realization of system load, a small number of initial ADMM iterations is taken as input to infer the converged values and directly inject them into the iteration. We empirically demonstrate that the online injection of these values into the ADMM iteration accelerates convergence by a significant factor for partitioned 14-, 118- and 2848-bus test systems under differing load scenarios. The proposed method has several advantages: it maintains the security of private decision variables inherent in consensus ADMM; inference is fast and so may be used in online settings; RNN-generated predictions can dramatically improve time to convergence but, by construction, can never result in infeasible ADMM subproblems; it can be easily integrated into existing software implementations. While we focus on the ADMM formulation of distributed DC-OPF in this paper, the ideas presented are naturally extended to other distributed optimization problems.

alternating direction method of multipliers↗

Using machine learning for particle track identification in the CLAS12 detector

Particle track reconstruction is the most computationally intensive process in nuclear physics experiments. Traditional algorithms use a combinatorial approach that exhaustively tests track measurements ("hits") to identify those that form an actual particle trajectory. In this article, we describe the development of four machine learning (ML) models that assist the tracking algorithm by identifying valid track candidates from the measurements in drift chambers. Several types of machine learning models were tested, including: Convolutional Neural Networks (CNN), Multi-Layer Perceptrons (MLP), Extremely Randomized Trees (ERT) and Recurrent Neural Networks (RNN). As a result of this work, an MLP network classifier was implemented as part of the CLAS12 reconstruction software to provide the tracking code with recommended track candidates. The resulting software achieved accuracy of greater than 99% and resulted in an end-to-end speedup of 35% compared to existing algorithms.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

A Survey: Handling Irregularities in Neural Network Acceleration with FPGAs

In the last decade, Artificial Intelligence (AI) through Deep Neural Networks (DNNs) has penetrated virtually every aspect of science, technology, and business. Many types of DNNs have been and continue to be developed, including Convolutional Neural Networks (CNNs), Recurrent Neural Net- works (RNNs), and Graph Neural Networks (GNNs). The overall problem for all of these Neural Networks (NNs) is that their target applications generally pose stringent constraints on latency and throughput, while also having strict accuracy requirements. There have been many previous efforts in creating hardware to accelerate NNs. The problem designers face is that optimal NN models typically have significant irregularities, making them hardware-unfriendly. In this paper, we first define the problems in NN acceleration by characterizing common irregularities in NN processing into 4 types; then we summarize the existing works that handle the four types of irregularities efficiently using hardware, especially FPGAs; finally, we provide a new vision of next-generation FPGA-based NN acceleration: that the emerging heterogeneity in the next-generation FPGAs is the key to achieving higher performance.

Geng, Tong↗

Predictive Skill of Deep Learning Models Trained on Limited Sequence Data

In this report we investigate the utility of one-dimensional convolutional neural network (CNN) models in epidemiological forecasting. Deep learning models, especially variants of recurrent neural networks (RNNs) have been studied for influenza forecasting, and have achieved higher forecasting skill compared to conventional models such as ARIMA models. In this study, we adapt two neural networks that employ one-dimensional temporal convolutional layers as a primary building block temporal convolutional networks and simple neural attentive meta-learner for epidemiological forecasting and test them with influenza data from the US collected over 2010-2019. We find that epidemiological forecasting with CNNs is feasible, and their forecasting skill is comparable to, and at times, superior to, RNNs. Thus CNNs and RNNs bring the power of nonlinear transformations to purely data-driven epidemiological models, a capability that heretofore has been limited to more elaborate mechanistic/compartmental disease models.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Development of a Machine-Learned Cruise Guide Indicator for Rotorcraft

This paper presents a machine-learned virtual cruise guide indicator (vCGI) for Chinook helicopters. Two temporal neural networks were trained and evaluated on measured data from 55 flight tests, one for the fore rotor and another for the aft rotor, to predict a vCGI value, which protects 23 components from fatigue damage during steady-state conditions. Three different classes of machine learning architectures were evaluated for prediction of the vCGI from time sequences: a temporal convolutional neural network with 1D dilated causal convolutions, a long short-term memory recurrent neural network, and an attention-based transformer architecture. The final average model accuracy on unseen flight data is currently greater than 93% for CGI values which could result in fatigue damage and 90% for normal operation CGI values. Model accuracy was improved through a series of advancements in:(1) selection of optimal training data using temporal collective variables and unsupervised learning, (2) dataset augmentation with maximum-entropy temporal collective variables, and (3) implementation of a mixture-of-experts classification- regression approach using an adversarial classification approach to assign maneuver labels. The results are presented for each advancement in model development along with lessons learned in training machine learning models on real- world, time-dependent rotorcraft data.

Boyer, Mathew↗

Deployment of Dynamic Neural Network Optimization to Minimize Heat Rate During Ramping for Coal Power Plants (Final Technical Report)

Much success was achieved throughout the course of this project. A successful implementation of Dynamic Neural Network Optimization (D-NNO) was coupled with Adaptive Predictive Controls (APC) and a novel hardware installation comprised of an advanced sensor network (ASN) measuring mass-weighted averages of flue gas constituents above the horizontal superheater of a coal-fired utility boiler. From 2019 through 2023 (including an extension due to COVID delays), the team was able to prototype, evaluate, deploy, iterate, and ultimately finalize an advanced closed-loop control D-NNO system which demonstrated the ability to: •improve unit efficiency ~2.0% relative to unoptimized operation (represented as total fuel fired per MWh generated) •improve unit NOx emission rates 10%+ beyond static optimization baselines •improve unit temperature stability as much as 58% and on average 12% •improve operating load stability as much as 35% The culmination of this project has generated an advanced methodology of deploying specially designed recurrent neural networks (long short-term memory, gated recurrent unit, encoder-decoder networks, transformers, etc.), customized trajectory planning and closed-loop optimization modules capable of adapting to live electric grid responses and demands, self-tuning and adaptive expert controls constantly adjusting prediction parameters to real-time unit behavior, and a hardware/software package able to reliably calculate net unit heat rate (NUHR) in real-time using flue gas constituents, machine learning, and known combustion relationships. Through this real-time NUHR value, immediate feedback on system adjustments relative to operating efficiency was available, allowing for rapid improvements to system performance. In addition to development and deployment of the advanced D-NNO system, the approach methodology has been readily commercialized through the project platform Griffin Open Systems, LLC, the D-NNO software platform host. Similar methodologies to those developed by this project have already been deployed at 5 other units across the United States, with another 6 implementations scheduled, and more expected. Over the course of the project, multiple academic papers were submitted and accepted for publication within esteemed academic journals, and PhD students were trained and graduated, as well as undergraduate students becoming involved and participating to project objectives.

01 COAL, LIGNITE, AND PEAT↗

Utilizing Physics-Informed Synthetic Data to Train a Digital Twin for Predicting Reactor Operations

Understanding techniques to strengthen the nuclear nonproliferation regime is crucial in reducing the creation of nuclear weaponry on the basis of advancements in the nuclear energy industry. Prior to construction of a nuclear power plant, it is necessary to understand the proliferation potential of the plant’s reactor. Digital twins serve as a unique solution to recognizing reactor behavior indicative of nuclear proliferation. The following research conducted serves as a validation to training a digital twin on synthetic data fabricated via means of reactor physics simulations based on parameters of Idaho State University’s AGN-201 reactor. The synthetic data is utilized to train a long short-term memory (LSTM) recurrent neural network model. The accuracy of the predicted data is measured against real operational data to verify the reliability of the synthetic data creation methods and if these methods should be used in the future to inform inspectors of a reactor’s proliferation capabilities.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

Variational data augmentation for a learning-based granular predictive model of power outages

As the trend in climate change continues, extreme weather events are expected to occur with increasing frequency and severity and pose a significant threat to the electric power infrastructure. Regardless of the efforts a utility puts towards hardening the grid, storm-induced damage to the utility assets such as cables and distributed energy resources (DERs) that are particularly vulnerable to such events is unavoidable. Access to a highly granular, in space and time, outage forecasting tool with long lead times (i.e., days ahead) will enhance the efficiency of service restoration efforts. Here, in this study, we propose to develop and implement a multi-model framework as an operational tool based on a granular and multi-day outage forecasting model using operational numerical weather prediction model forecasts and detailed component outage information. An innovative two-layered recurrent neural network, i.e., a long-short-term-memory (LSTM)-based variational autoencoder (VAE) framework and a sliding window are used to address the uneven distribution of different types of weather events and make better use of the time-series data. Case studies are performed to demonstrate the performance of the new framework.

54 ENVIRONMENTAL SCIENCES↗

Towards a robust out-of-the-box neural network model for genomic data

The accurate prediction of biological features from genomic data is paramount for precision medicine and sustainable agriculture. For decades, neural network models have been widely popular in fields like computer vision, astrophysics and targeted marketing given their prediction accuracy and their robust performance under big data settings. Yet neural network models have not made a successful transition into the medical and biological world due to the ubiquitous characteristics of biological data such as modest sample sizes, sparsity, and extreme heterogeneity. Here, we investigate the robustness, generalization potential and prediction accuracy of widely used convolutional neural network and natural language processing models with a variety of heterogeneous genomic datasets. Mainly, recurrent neural network models outperform convolutional neural network models in terms of prediction accuracy, overfitting and transferability across the datasets under study. While the perspective of a robust out-of-the-box neural network model is out of reach, we identify certain model characteristics that translate well across datasets and could serve as a baseline model for translational researchers.

59 BASIC BIOLOGICAL SCIENCES↗

Reduced Order Model of Transactive Bidding Loads

Transactive energy (TE) has been identified to provide better grid efficiency and reliability by market-based transactive exchanges between energy producers and energy consumers. Simulations of TE systems are crucial to evaluate the benefits and impacts of different transactive mechanisms. However, such simulations can be time consuming due to the information exchange between various participants and complex co-simulation environments. In this paper, we develop a reduced order model to speed up the simulation of transactive systems in TE simulation platform (TESP) while achieving very low error between the reduced order and full model results. Specifically, the developed reduced order model consists of an aggregate responsive load agent which utilizes two Recurrent Neural Networks (RNNs) with Long Short-Term Memory units (LSTMs) to enable transactive elements to collectively participate in the TE system. The proposed aggregate responsive load (ARL) agent is able to produce similar transactive behaviors to the full simulation model while achieving significant simulation time reduction. Finally, we also show that the developed model enables generalization of simulation results across different dates and across different number of loads included in the simulations.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Physics-Informed Machine Learning Model for Ceramic Matrix Composite Creep

A physics-informed recurrent neural network (RNN) based surrogate model is developed to emulate the nonlinear, time-dependent constitutive behavior of ceramic matrix composites (CMCs) driven by matrix damage and constituent creep at the microscale. Physics-informed constraints are introduced into the surrogate model through regularization to ground the prediction in physics and improve its predictive capabilities. Training data is generated using the high-fidelity generalized method of cells (HFGMC) approach which calls appropriate creep and damage models for each of the constituents. This coupling permits simulating the nonlinear behavior of CMCs based on constituent response at the microscale along with microstructural features such as fiber and porosity volume fraction and fiber radius. The microscale repeating unit cell is loaded under creep fatigue conditions to replicate the material loading experienced in a turbine engine. Therefore, the RNN-based surrogate model is tasked with predicting, as a function of variable input stress sequence, temperature, and microstructural features, the resulting strain history response while satisfying physical constraints related to creep rate, isochoric inelastic deformation, and strain energy density. The trained surrogate model is shown to effectively match the strain history over quantified distributions of microstructural features and relevant loading regimes and temperatures. Neural network based surrogate models can offer efficient alternatives to running computationally intensive multiscale material models to simulate the nonlinear response of large structural models. Therefore, the presented work provides evidence towards the feasibility of developing, training, and running such models for CMCs with complex microstructures, nonlinear time-dependent material response, and under non-monotonic loading conditions.

ceramic matrix composites↗

Reconstruction of fast neutron direction in segmented organic detectors using deep learning

A method for reconstructing the direction of a fast neutron source using a segmented organic scintillator-based detector and deep learning model is proposed and analyzed. Here, the model is based on recurrent neural network, which can be trained by a sequence of data obtained from an event recorded in the detector and suitably pre-processed. The performance of deep learning-based model is compared with the conventional double-scatter detection algorithm in reconstructing the direction of a fast neutron source. With the deep learning model, the uncertainty in source direction of 0.301 rad is achieved with 100 neutron detection events in a segmented cubic organic scintillator detector with a side length of 46 mm. To reconstruct the source direction with the same angular resolution as the double-scatter algorithm, the deep learning method requires 75% fewer events. Application of this method could augment the operation of segmented detectors operated in the neutron scatter camera configuration for applications such as special nuclear material detection.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

A novel transfer learning framework for sorghum biomass prediction using UAV-based remote sensing data and genetic markers

Yield for biofuel crops is measured in terms of biomass, so measurements throughout the growing season are crucial in breeding programs, yet traditionally time- and labor-consuming since they involve destructive sampling. Modern remote sensing platforms, such as unmanned aerial vehicles (UAVs), can carry multiple sensors and collect numerous phenotypic traits with efficient, non-invasive field surveys. However, modeling the complex relationships between the observed phenotypic traits and biomass remains a challenging task, as the ground reference data are very limited for each genotype in the breeding experiment. In this study, a Long Short-Term Memory (LSTM) based Recurrent Neural Network (RNN) model is proposed for sorghum biomass prediction. The architecture is designed to exploit the time series remote sensing and weather data, as well as static genotypic information. As a large number of features have been derived from the remote sensing data, feature importance analysis is conducted to identify and remove redundant features. A strategy to extract representative information from high-dimensional genetic markers is proposed. To enhance generalization and minimize the need for ground reference data, transfer learning strategies are proposed for selecting the most informative training samples from the target domain. Consequently, a pre-trained model can be refined with limited training samples. Field experiments were conducted over a sorghum breeding trial planted in multiple years with more than 600 testcross hybrids. The results show that the proposed LSTM-based RNN model can achieve high accuracies for single year prediction. Further, with the proposed transfer learning strategies, a pre-trained model can be refined with limited training samples from the target domain and predict biomass with an accuracy comparable to that from a trained-from-scratch model for both multiple experiments within a given year and across multiple years.

36 MATERIALS SCIENCE↗

Long–short-term memory encoder–decoder with regularized hidden dynamics for fault detection in industrial processes

The ability of recurrent neural networks (RNN) to model nonlinear dynamics of high dimensional process data has enabled data-driven RNN-based fault detection algorithms. Previous studies have focused on detecting faults by identifying the discrepancies in data distribution between the faulty and normal data, as reflected in prediction errors generated by RNN models. However, in industrial processes, variations in data distribution can also result from changes in normal control setpoints and compensatory control adjustments in response to disturbances, making it hard to differentiate between normal and faulty conditions. This paper proposes a fault detection method utilizing a long short-term memory (LSTM) encoder–decoder structure with regularized hidden dynamics and reversible instance normalization (RevIN) to compactly represent high-dimensional measurements for effective monitoring. During training, the hidden states of the model are regularized to form a low-dimensional latent space representation of the original multivariate time series data. As a result, the prediction errors of the latent states can be used to monitor the abnormal dynamic variations, while the reconstruction errors of the measured variables are used to monitor the abnormal static variations. Furthermore, the proposed indices can reflect operating conditions, even when the distribution of test data changes, which helps distinguish faults from normal adjustments and disturbances that controllers can settle. Here, data from numerical simulation and the Tennessee Eastman process are used to illustrate the effectiveness of the proposed fault detection method.

42 ENGINEERING↗

xesn: Echo state networks powered by Xarray and Dask

Xesn is a Python package that allows scientists to easily design Echo State Networks (ESNs) for forecasting problems. ESNs are a Recurrent Neural Network architecture introduced by Jaeger (2001) that are part of a class of techniques termed Reservoir Computing. One defining characteristic of these techniques is that all internal weights are determined by a handful of global, scalar parameters, thereby avoiding problems during backpropagation and reducing training time significantly. Because this architecture is conceptually simple, many scientists implement ESNs from scratch, leading to questions about computational performance. Xesn offers a straightforward, standard implementation of ESNs that operates efficiently on CPU and GPU hardware. The package leverages optimization tools to automate the parameter selection process, so that scientists can reduce the time finding a good architecture and focus on using ESNs for their domain application. Importantly, the package flexibly handles forecasting tasks for out-of-core, multi-dimensional datasets, eliminating the need to write parallel programming code. Xesn was initially developed to handle the problem of forecasting weather dynamics, and so it integrates naturally with Python packages that have become familiar to weather and climate scientists such as Xarray (Hoyer & Hamman, 2017). However, the software is ultimately general enough to be utilized in other domains where ESNs have been useful, such as in signal processing (Jaeger & Haas, 2004).

97 MATHEMATICS AND COMPUTING↗

Online evolutionary neural architecture search for multivariate non-stationary time series forecasting

Time series forecasting (TSF) is one of the most important tasks in data science. TSF models are usually pre-trained with historical data and then applied on future unseen datapoints. However, real-world time series data is usually non-stationary and models trained offline usually face problems from data drift. Models trained and designed in an offline fashion can not quickly adapt to changes quickly or be deployed in real-time. To address these issues, this work presents the Online NeuroEvolution-based Neural Architecture Search (ONE-NAS) algorithm, which is a novel neural architecture search method capable of automatically designing and dynamically training recurrent neural networks (RNNs) for online forecasting tasks. Without any pre-training, ONE-NAS utilizes populations of RNNs that are continuously updated with new network structures and weights in response to new multivariate input data. ONE-NAS is tested on real-world, large-scale multivariate wind turbine data as well as the univariate Dow Jones Industrial Average (DJIA) dataset. These results demonstrate that ONE-NAS outperforms traditional statistical time series forecasting methods, including online linear regression, fixed long short-term memory (LSTM) and gated recurrent unit (GRU) models trained online, as well as state-of-the-art, online ARIMA strategies. Additionally, results show that utilizing multiple populations of RNNs which are periodically repopulated provide significant performance improvements, allowing this online neural network architecture design and training to be successful.

97 MATHEMATICS AND COMPUTING↗