Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Recurrent neural network”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

TorchBraid: High-Performance Layer-Parallel Training of Deep Neural Networks with MPI and GPU Acceleration

TorchBraid is a high-performance implementation of layer-parallel training for deep neural networks (DNNs) supporting MPI-based parallelism and GPU acceleration. Layer-parallel training has been developed to overcome the serialization inherent in forward and backward propagation of DNNs that limits utilization of computational resources in the strong scaling limit. To achieve this, TorchBraid integrates the PyTorch neural network framework with the state-of-the-art XBraid time-parallel library. Furthermore, this article presents the use and performance of TorchBraid, in addition to solutions for overcoming the algorithmic challenges inherent in combining automatic differentiation with layer-parallel. Results are presented with and without GPU acceleration for the Tiny ImageNet and MNIST image classification data sets, as well as recurrent neural networks. Overall, TorchBraid enables fast training of DNNs, both in a strong and weak scaling context. In addition to the TorchBraid software, several new advances in applying layer-parallel algorithms are detailed. Integration of layer-parallel with data-parallel algorithms is presented for the first time, showing the computational advantages of the combination. Standard deep learning techniques, like batch-normalization, are developed for layer-parallel training. Finally, a new approach combining layer-parallel with spatial coarsening in order to accelerate training for 3D image classification shows roughly a 10× speedup over serial execution.

Layer-parallel↗

Data-driven based coordinated smart inverter control for distributed energy resources

Smart inverters (SI) for distributed energy resources (DER) are becoming popular since they have the ability to stabilize as well as restore the voltage and frequency of power systems. Aiming at establishing the mathematical models combined with SI control methods, multiple optimization methods are developed. However, the computational complexity of solving such a mathematical model with various uncertainties limits the real-time application of the SI control. To conquer this challenge, a data-driven-based SI control approach is developed to achieve coordinated control in the high penetration DER system. First, an optimization problem for maximizing the active power generation and minimizing the power loss is designed using the Volt/VAR control. To reduce the time consumption, the recurrent neural network (RNN) is proposed to model the relationship between the uncertainties and control actions during the offline site. The RNN with different sub-structures such as the long short-term memory cell and gated recurrent unit cell are included to enrich the diversity of features. In the last stage, different experiment comparisons, including multiple uncertainties maps and stateof- art machine learning methods, are conducted to verify the effectiveness of the proposed method based on the IEEE 123 bus power system. The results demonstrate that the proposed method can effectively achieve a rapid and coordinated control with a lower error rate.

Qiu, Wei↗

Watch and learn—a generalized approach for transferrable learning in deep neural networks via physical principles

Transfer learning refers to the use of knowledge gained while solving a machine learning task and applying it to the solution of a closely related problem. Such an approach has enabled scientific breakthroughs in computer vision and natural language processing where the weights learned in state-of-the-art models can be used to initialize models for other tasks which dramatically improve their performance and save computational time. Here we demonstrate an unsupervised learning approach augmented with basic physical principles that achieves fully transferrable learning for problems in statistical physics across different physical regimes. By coupling a sequence model based on a recurrent neural network to an extensive deep neural network, we are able to learn the equilibrium probability distributions and inter-particle interaction models of classical statistical mechanical systems. Our approach, distribution-consistent learning, DCL, is a general strategy that works for a variety of canonical statistical mechanical models (Ising and Potts) as well as disordered interaction potentials. Using data collected from a single set of observation conditions, DCL successfully extrapolates across all temperatures, thermodynamic phases, and can be applied to different length-scales. This constitutes a fully transferrable physics-based learning in a generalizable approach.

97 MATHEMATICS AND COMPUTING↗

Upscaling Reactive Transport and Clogging in Shale Microcracks by Deep Learning

Fracture networks in shales exhibit multiscale features. A rock system may contain a few main fractures and thousands of microcracks, whose length and aperture are orders of magnitude smaller than the former. It is computationally prohibitive to resolve all the fractures explicitly for such multiscale fracture networks. One traditional approach is to model the small-scale features (e.g., microcracks in shales) as an effective medium. Although this fracture-matrix conceptualization significantly reduces the problem complexity, there are classes of physical processes that cannot be accurately upscaled by effective medium approximations, e.g., microcrack clogging during mineral reactions. Here, we employ deep learning in place of effective medium theory to upscale physical processes in small-scale features. Specifically, we consider reactive transport in a fracture-microcrack network where microcracks can be clogged by precipitation. A deep learning multiscale algorithm is developed, in which the microcracks are upscaled as a wall boundary condition of the main fractures. The wall boundary condition is constructed by recurrent neural networks, which take concentration histories as input and predict the solute transport from main fractures to microcracks. The deep learning multiscale algorithm is firstly employed in specific scenarios, then a general model is developed which can work under various conditions. The new approach is validated against fully resolved simulations and an analytical solution, providing a reliable and efficient solution for problems that cannot be upscaled by effective medium models.

58 GEOSCIENCES↗

Neural Network-Based Electric Vehicle Range Prediction for Smart Charging Optimization

Range prediction is a standard feature in most modern road vehicles, allowing drivers to make informed decisions about when to refuel. Most vehicles make range predictions through data- or model-driven means, monitoring the average fuel consumption rate or using a tuned vehicle model to predict fuel consumption. The uncertainty of future driving conditions makes the range prediction problem challenging, particularly for less pervasive battery electric vehicles (BEV). Most contemporary machine learning-based methods attempt to forecast the battery SOC discharge profile to predict vehicle range. In this work, we propose a novel approach using two recurrent neural networks (RNNs) to predict the remaining range of BEVs and the minimum charge required to safely complete a trip. Each RNN has two outputs that can be used for statistical analysis to account for uncertainties; the first loss function leads to mean and variance estimation (MVE), while the second results in bounded interval estimation (BIE). These outputs of the proposed RNNs are then used to predict the probability of a vehicle completing a given trip without charging, or if charging is needed, the remaining range and minimum charging required to finish the trip with high probability. Training data was generated using a low-order physics model to estimate vehicle energy consumption from historical drive cycle data collected from medium-duty last-mile delivery vehicles. Here, the proposed method demonstrated high accuracy in the presence of day-to-day route variability, with the root-mean-square error (RMSE) below 6% for both RNN models.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Deep Learning Predicts Stress–Strain Relations of Granular Materials Based on Triaxial Testing Data

This study presents an AI-based constitutive modelling framework wherein the prediction model directly learns from triaxial testing data by combining discrete element modelling (DEM) and deep learning. A constitutive learning strategy is proposed based on the generally accepted frame-indifference assumption in constructing material constitutive models. The low-dimensional principal stress-strain sequence pairs, measured from discrete element modelling of triaxial testing, are used to train recurrent neural networks, and then the predicted principal stress sequence is augmented to other high-dimensional or general stress tensor via coordinate transformation. Through detailed hyperparameter investigations, it is found that long short-term memory (LSTM) and gated recurrent unit (GRU) networks have similar prediction performance in constitutive modelling problems, and both satisfactorily predict the stress responses of granular materials subjected to a given unseen strain path. Furthermore, the unique merits and ongoing challenges of data-driven constitutive models for granular materials are discussed.

42 ENGINEERING↗

Neural-network learning of SPOD latent dynamics

Here, we aim to reconstruct the latent space dynamics of high dimensional, quasi-stationary systems using model order reduction via the spectral proper orthogonal decomposition (SPOD). The proposed method is based on three fundamental steps: in the first, once that the mean flow field has been subtracted from the realizations (also referred to as snapshots), we compress the data from a high-dimensional representation to a lower dimensional one by constructing the SPOD latent space; in the second, we build the time-dependent coefficients by projecting the snapshots containing the fluctuations onto the SPOD basis and we learn their evolution in time with the aid of recurrent neural networks; in the third, we reconstruct the high-dimensional data from the learnt lower -dimensional representation. The proposed method is demonstrated on two different test cases, namely, a compressible jet flow, and a geophysical problem known as the Madden-Julian Oscillation. An extensive comparison between SPOD and the equivalent POD-based counterpart is provided and differences between the two approaches are highlighted. The numerical results suggest that the proposed model is able to provide low rank predictions of complex statistically stationary data and to provide insights into the evolution of phenomena characterized by specific range of frequencies. The comparison between POD and SPOD surrogate strategies highlights the need for further work on the characterization of the interplay of error between data reduction techniques and neural network forecasts.

97 MATHEMATICS AND COMPUTING↗

Using Neural Networks for Low Energy Reconstruction and Neutron Identification in the MicroBooNE LArTPC

Identifying and reconstructing final-state neutrons from neutrino interactions in Liquid Argon Time Projection Chambers (LArTPCs) will enhance future oscillation measurements by recovering missing energy and improving neutrino interaction channel identification. However, neutrons are challenging to reconstruct as the majority leave only small, isolated charge signatures known as blips. Here we present initial efforts to identify neutrons in the MicroBooNE LArTPC with low energy protons from neutron-argon inelastic interactions that present as blips below the traditional tracking threshold in the TPC. Unlike for tracks, there is no algorithmic method to determine direction for blips since they span only a few wires. Therefore, we developed and trained a Recurrent Neural Network (RNN) to reconstruct the directionality of proton-induced blips, allowing us to separate signal from background by selecting blips that point back to the neutrino vertex. The model achieves a preliminary average angular resolution of 17 degrees when tested on a simulated sample of protons over 6 MeV in kinetic energy. This novel tool will enhance neutron detection in LArTPCs and expand a broad range of other low-energy physics searches such as for solar and supernova neutrinos.

Silva, Liani Isabel [Unlisted, US]↗

Evapotranspiration partitioning estimates from 8 methods from 47 NEON sites, 2019-2021

This dataset provides daily estimates of evapotranspiration (ET) and the transpiration-to-evapotranspiration ratio (T/ET) across 47 terrestrial National Ecological Observatory Network (NEON) sites spanning diverse environmental and biome conditions in the United States across three years of data (2019-2021). Daily ET is reported in both energy units (MJ m⁻² day⁻¹) and equivalent water depth (mm day⁻¹), assuming a constant latent heat of vaporization of 2.45 MJ/kg. The primary method uses a hybrid recurrent neural network–Penman–Monteith framework (RNN-PM), which integrates physically based surface energy balance constraints with data-driven learning to partition ET into transpiration and evaporation components. Model inputs include in situ meteorological observations (air temperature, vapor pressure deficit, wind speed, and radiation) combined with satellite-derived land surface temperature, leaf area index, and soil moisture. For benchmarking and uncertainty assessment, T/ET estimates from seven additional models are included: Priestley-Taylor Jet Propulsion Laboratory (PT-JPL), Penman-Monteith (P-M), Two-Source Energy Balance (TSEB), Support Vector Regression (SVR), and Categorical Boosting (CatBoost), among others—spanning empirical, machine-learning, and process-based approaches (see methods section or linked publication for detailed descriptions). Data Package Contents: The dataset a csv files containing daily ET and T/ET estimates for each site and model, along with associated metadata files these variables. Data can be accessed using common spreadsheet software (e.g., Microsoft Excel, LibreOffice) or programming environments such as R or Python. Together, these data support cross-site comparisons of ecosystem water use, evaluation of ET partitioning methods, and development of improved land–atmosphere exchange models.

EARTH SCIENCE > ATMOSPHERE↗

From Coexpression to Coregulation: An Approach to Inferring Transcriptional Regulation Among Gene Classes from Large-Scale Expression Data

We provide preliminary evidence that existing algorithms for inferring small-scale gene regulation networks from gene expression data can be adapted to large-scale gene expression data coming from hybridization microarrays. The essential steps are (I) clustering many genes by their expression time-course data into a minimal set of clusters of co-expressed genes, (2) theoretically modeling the various conditions under which the time-courses are measured using a continuous-time analog recurrent neural network for the cluster mean time-courses, (3) fitting such a regulatory model to the cluster mean time courses by simulated annealing with weight decay, and (4) analysing several such fits for commonalities in the circuit parameter sets including the connection matrices. This procedure can be used to assess the adequacy of existing and future gene expression time-course data sets for determining transcriptional regulatory relationships such as coregulation.

Mjolsness, Eric↗

Construction of generalized quasilinear diffusion coefficient using neural networks with physical restrictions

The quasilinear diffusion coefficient (D QL ) derived from our machine learning framework shows comparable trends with the ground truth D QL obtained from GENRAY-CQL3D simulations. Additionally, for the strong absorption cases, the radial current drive profiles generated using the D QL from our model exhibit consistent behavior with those obtained from the original simulation. These findings indicate the potential of our surrogate modeling approach with physical restrictions to replicate key wave–plasma interaction characteristics while reducing computational costs. Traditionally, calculating D QL for wave–particle interactions relies on computationally intensive wave simulations coupled with Fokker–Planck solvers. To address this challenge, we developed a machine learning-based surrogate model with physical restrictions derived from cold plasma theory and bounce-averaged damping effects. First, we establish the propagation domain of Lower Hybrid Waves in the (N∥, ρ) space by identifying the accessibility limit and determining the upper and lower bounds of N∥ using the Potential Power Deposition (PPD) method. Subsequently, leveraging a database constructed using Latin hypercube sampling alongside the underlying physical restrictions (e.g. PPD), machine learning methods including U-Net and Recurrent Neural Networks are employed to design a physics-restricted machine learning framework capable of reconstructing D QL .

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Monitoring of Temperature Measurements for Different Flow Regimes in Water and Galinstan with Long Short-Term Memory Networks and Transfer Learning of Sensors

Temperature sensing is one of the most common measurements of a nuclear reactor monitoring system. The coolant fluid flow in a reactor core depends on the reactor power state. We investigated the monitoring and estimation of the thermocouple time series using machine learning for a range of flow regimes. Measurement data were obtained, in two separate experiments, in a flow loop filled with water and with liquid metal Galinstan. We developed long short-term memory (LSTM) recurrent neural networks (RNNs) for sensor predictions by training on the sensor’s own prior history, and transfer learning LSTM (TL-LSTM) by training on a correlated sensor’s prior history. Sensor cross-correlations were identified by calculating the Pearson correlation coefficient of the time series. The accuracy of LSTM and TL-LSTM predictions of temperature was studied as a function of Reynolds number (Re). The root-mean-square error (RMSE) for the test segment of time series of each sensor was shown to linearly increase with Re for both water and Galinstan fluids. Using linear correlations, we estimated the range of values of Re for which RMSE is smaller than the thermocouple measurement uncertainty. For both water and Galinstan fluids, we showed that both LSTM and TL-LSTM provide reliable estimations of temperature for typical flow regimes in a nuclear reactor. The LSTM runtime was shown to be substantially smaller than the data acquisition rate, which allows for performing estimation and validation of sensor measurements in real time.

Pantopoulou, Stella↗

One-Step Ahead Prediction of Thermal Mixing Tee Sensors with Long Short Term Memory (LSTM) Neural Networks

High-temperature advanced reactors under development, such as sodium fast reactors (SFR) and molten salt cooled reactors (MSCR), are expected to offer lower levelized cost of energy (LCOE) compared to existing light water reactor (LWR’s). In the existing light water reactors (LWR’s), operation and maintenance (O&M) expenses constitute the largest fraction of the total operating cost. Some of the O&M costs are related maintenance of sensors which can fail due to exposure to harsh environment in a reactor. The O&M costs of Advanced Reactor (AR)’s are expected to constitute a significant fraction of the total cost as well, because of high temperature and radiation level in AR are likely to cause material fatigue and premature failure of sensors and components. The O&M costs in AR’s could be reduced through integration of advanced informatics of performance-related sensors into a digital twin designed for reactor monitoring. For example, machine learning (ML) could be employed for real-time validation and correction of performance-related sensors, and reducing the number of performance-related physical sensor units through virtual sensing. As part of the effort, we investigate real-time validation of thermal hydraulic sensors through one-step ahead forecasting of sensor values using long short-term memory (LSTM) recurrent neural networks (RNN). The sensors are installed in a flow loop containing a thermal mixing tee, which is a common experimental model to study thermal fatigue in a thermal hydraulic loop. In addition, nonlinear transients generated in a thermal mixing tee constitute a good challenge data set for training and validation of ML algorithms. Sensors in this study include thermocouples, flow meters, and optical fibers for distributed temperature sensing. In one experiment, measurement data sets were obtained for a loop was filled with water, and in another experiment, measurements were performed on a loop filled with liquid metal Galinstan. We have also conducted preliminary investigation of one-step ahead prediction of fiber optics-based distributed temperature sensing with LSTM networks. In predicting fiber-based temperature measurements, we treated each gauge pitch of the fiber as an independent sensor. Accuracy of one-step ahead forecasting was estimated by calculating root mean square error (RMSE) for the test segment of time series of each sensor. RMSE’s for temperature sensors in water loop were, for the most part, lower than for the same sensors in Galinstan loop. The RMSE’s for flow meters were similar for both loops. The RMSE’s for distributed temperature measured with the fiber optic sensor were similar to those of the point sensors. Results of this study demonstrated the capability of LSTM one-step ahead forecasting with RMSE comparable to uncertainty in sensor measurements.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Sequence-to-sequence neural networks for short-term electrical load forecasting in commercial office buildings

The U.S. power grid is transforming to become smarter, cleaner, and more effi- cient. This is leading to the addition of significant distributed variable renew- able generation. Due to the variable nature of renewable generation, the short- and long-term supply-demand imbalances are less predictable, and conventional approaches to mitigating the imbalance will not be efficient or cost-effective. To address this challenge, transactive control technologies have been proposed which balance energy generation and consumption with market activity and in- frastructural limitations. Transactive control requires the ability of individual end-use loads to express flexibility as a function of a transactive signal (e.g., price). Empirical gray- and black-box models have been widely used to express flexibility, and although these approaches are generally easy to construct and simple to use, they do not capture the non-linear behavior that some end-use loads represent . Machine learning approaches have been proposed to address this limitation. Although deep learning approaches for forecasting end-use loads have been explored, certain aspects of the application of deep models to load forecasting are not well understood. These aspects include how much training data is required, and how models should be structured and trained. To that end, this work explores how to approach applying deep recurrent neural networks to short-term electrical load forecasting with a case study of four commercial office buildings. We identify data requirements for training accurate models of whole building electricity use conditioned on outdoor temperature, provide insight into model hyperparameter sensitivity, and demonstrate how readily models can be generalized to unseen buildings.

Skomski, Elliott↗

Neural Network Approaches for Mobile Spectroscopic Gamma-Ray Source Detection

Artificial neural networks (ANNs) for performing spectroscopic gamma-ray source identification have been previously introduced, primarily for applications in controlled laboratory settings. To understand the utility of these methods in scenarios and environments more relevant to nuclear safety and security, this work examines the use of ANNs for mobile detection, which involves highly variable gamma-ray background, low signal-to-noise ratio measurements, and low false alarm rates. Simulated data from a 2” × 4” × 16” NaI(Tl) detector are used in this work for demonstrating these concepts, and the minimum detectable activity (MDA) is used as a performance metric in assessing model performance.In addition to examining simultaneous detection and identification, binary spectral anomaly detection using autoencoders is introduced in this work, and benchmarked using detection methods based on Non-negative Matrix Factorization (NMF) and Principal Component Analysis (PCA). On average, the autoencoder provides a 12% and 23% improvement over NMF- and PCA-based detection methods, respectively. Additionally, source identification using ANNs is extended to leverage temporal dynamics by means of recurrent neural networks, and these time-dependent models outperform their time-independent counterparts by 17% for the analysis examined here. The paper concludes with a discussion on tradeoffs between the ANN-based approaches and the benchmark methods examined here.

Bilton, Kyle J. (ORCID:0000000184553689)↗

Deep neural operators can predict the real-time response of floating offshore structures under irregular waves

The use of neural operators in a digital twin model of an offshore floating structure holds the potential for a significant shift in the prediction of structural responses and health monitoring, offering valuable real-time control insights. In this work, we investigate the effectiveness of three neural operators, namely the deep operator network (DeepONet), the Fourier neural operator (FNO), and the Wavelet neural operator (WNO), to accurately capture the responses of a floating structure under six different sea state codes (3 − 8) based on the wave characteristics described by the World Meteorological Organization (WMO). To further enhance the accuracy of the vanilla architecture of the neural operators, novel extensions, such as wavelet-DeepONet and self-adaptive WNO, are proposed in this paper. The results demonstrate that these high-precision neural operators can deliver structural responses more efficiently, up to two orders of magnitude faster than a dynamic analysis using conventional numerical solvers. Additionally, compared to gated recurrent units (GRUs), a commonly used recurrent neural network for time-series estimation, neural operators are both more accurate and efficient, especially in situations with limited data availability. Taken together, our study shows that FNO outperforms all other operators for approximating the mapping of one input functional space to the output space as well as for responses that have small bandwidth of the frequency spectrum. Conversely, DeepONet, with historical states, proves most accurate in learning the mapping of multiple input functions to the output space and capturing responses within a broad frequency spectrum.

97 MATHEMATICS AND COMPUTING↗

DroughtCast: A Machine Learning Forecast of the United States Drought Monitor

Drought is one of the most ecologically and economically devastating natural phenomena affecting the United States, causing the U.S. economy billions of dollars in damage, and driving widespread degradation of ecosystem health. Many drought indices are implemented to monitor the current extent and status of drought so stakeholders such as farmers and local governments can appropriately respond. Methods toforecast drought conditions weeks to months in advance are less common but would provide a more effective early warning system to enhance drought response, mitigation, and adaptation planning. To resolve this issue, we introduce DroughtCast, a machine learning framework for forecasting the United States Drought Monitor (USDM). DroughtCast operates on the knowledge that recent anomalies in hydrology and meteorology drive future changes in drought conditions. We use simulated meteorology and satellite observed soil moisture as inputs into a recurrent neural network to accurately forecast the USDM between 1 and 12 weeks into the future. Our analysis shows that precipitation, soil moisture, and temperature are the most important input variables when forecasting future drought conditions. Additionally, a case study of the 2017 Northern Plains Flash Drought shows that DroughtCast was able to forecast a very extreme drought event up to 12 weeks before its onset. Given the favorable forecasting skill of the model, DroughtCast may provide a promising tool for land managers and local governments in preparing for and mitigating the effects of drought.

Machine Learning↗

Inter-well connectivity detection in CO 2 WAG projects using statistical recurrent unit models

Routine well-wise injection and production measurements contain significant information on subsurface structure and properties. Data-driven technology that interprets surface data into subsurface structure or properties can assist operators in making informed decisions by providing a better understanding of field assets. Our machine-learning framework is built on the statistical recurrent unit (SRU) model and interprets well-based injection/production data into inter-well connectivity without relying on a geologic model. We test it on synthetic and field-scale CO 2 EOR projects utilizing the water-alternating-gas (WAG) process. SRU is a special type of recurrent neural network (RNN) that allows for better characterization of temporal trends, by learning various statistics of the input at different time scales. In our application, the complete states (injection rate, pressure and cumulative injection) at injectors and pressure states at producers are fed to SRU as the input and the phase rates at producers are treated as the output. Once the SRU is trained and validated, it is then used to assess the connectivity of each injector to any producer using permutation variable importance method, wherein inputs corresponding to an injector are shuffled and the increase in prediction error at a given producer is recorded as the importance (connectivity metric) of the injector to the producer. This method is tested in both synthetic and field-scale cases. The validation of the proposed data-driven inter-well connectivity assessment is performed using synthetic data from simulation models where inter-well connectivity can be easily measured using the streamline-based flux allocation. The SRU model is shown to offer excellent prediction performance on the synthetic case. Despite significant measurement noise and frequent well shut-ins imposed in the field-scale case, the SRU model offers good prediction accuracy, the overall relative error of the phase production rates at most producers ranges from 10% to 30%. It is shown that the dominant connections identified by the data-driven method and streamline method are in close agreement. This significantly improves confidence in our data-driven procedure. The novelty of this work is that it is purely data-driven method and can directly interpret routine surface measurements to intuitive subsurface knowledge. Furthermore, the streamline-based validation procedure provides physics-based backing to the results obtained from data analytics. This study results in a reliable and efficient data analytics framework that is well-suited for large field applications.

42 ENGINEERING↗