Application of Recurrent Neural Network to Modeling Earth's Global Electron Density
Explore the source record for details and available documents.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
Because of a lack of operation data during abnormal and accident scenarios, along with the existence of uncertainty in the evaluation model for transient and accident analysis, the established abnormal and emergency operating procedures can be biased in characterizing the reactor states and ensuring operational resilience. To improve state awareness and ensure operational flexibility for minimizing effects on the system due to anomaly, digital twin (DT) technology is suggested to support operator's decision-making by effectively extracting and using knowledge of the current and future plant states from the knowledge base. To demonstrate DT's capability for recovering the complete states of reactors and for predicting the future reactor behaviors, this paper develops and assesses both the diagnosis and prognosis DTs in a nearly autonomous management and control system for an Experimental Breeder Reactor-II simulator during different loss-of-flow scenarios.
Abstract not provided.
A method of identifying molecular parameters in a complex mixture may include receiving a set of combined transition frequencies and analyzing the set of combined transition frequencies using a first trained artificial neural network to generate a plurality of separated transition frequency sets. Each of the plurality of separated frequency sets may be analyzed using a second trained artificial neural network to generate a respective set of estimated spectral parameters. The method may include identifying a set of molecular parameters corresponding to the set of separated transition frequencies.
Not provided.
Abstract We examine the zero-temperature Metropolis Monte Carlo (MC) algorithm as a tool for training a neural network by minimizing a loss function. We find that, as expected on theoretical grounds and shown empirically by other authors, Metropolis MC can train a neural net with an accuracy comparable to that of gradient descent (GD), if not necessarily as quickly. The Metropolis algorithm does not fail automatically when the number of parameters of a neural network is large. It can fail when a neural network’s structure or neuron activations are strongly heterogenous, and we introduce an adaptive Monte Carlo algorithm (aMC) to overcome these limitations. The intrinsic stochasticity and numerical stability of the MC method allow aMC to train deep neural networks and recurrent neural networks in which the gradient is too small or too large to allow training by GD. MC methods offer a complement to gradient-based methods for training neural networks, allowing access to a distinct set of network architectures and principles.
A data-driven model predictive control (MPC) was developed to enable the self-regulating capability of heat pipe (HP) nuclear microreactors. The MPC can proactively respond to potential disturbances of HP microreactors using three approaches for system identifications: linear state-space model, feedforward neural network, and recurrent neural networks with long short-term memory units. We present numerical results of data-driven MPCs to control the temperatures of selected HPs in a 37-HP test article. Our results show qualitatively that all data-driven MPCs produced similar control actions, while quantitatively, with artificial neural nets (especially feedforward neural nets), MPC can better follow drastic changes in setpoints with small errors.
To enable the self-regulating capability of heat pipe (HP) microreactors, an anticipatory control strategy through model predictive control (MPC) could proactively respond to potential disturbances and deviations in operating setpoints. However, a key factor prohibiting the widespread adoption of MPCs in nuclear applications is the effort and computational costs associated with learning and calibrating first-principles-based process models when the target system is complex and when there are gaps between modeled and target reactor systems. In this paper, we demonstrate data-driven MPC using three approaches for modeling the system dynamics, including a linear state-space model, feedforward neural network, and recurrent neural networks long short-term memory. We present the development and validation process of each model and compare the performance of data-driven MPCs in controlling the temperatures of selected HPs at the evaporator and condenser regions in a 37-HP-monolith system. Our results show that, qualitatively, all data-driven MPCs are producing similar control actions, while quantitatively, with artificial neural nets (especially feedforward neural nets), MPC can better follow drastic changes in setpoints with smallest errors.
Developmental second-order recurrent neural networks of special type modified to enhance stability in face of inputs beyond range of inputs on which trained. Second-order recurrent neural networks contain product feedback units and can be trained, by use of example inputs and outputs, to act as finite-state automatons. Particular second-order recurrent neural networks in question learn grammars in sense they are trained to generate binary responses to input training sequences of ones and zeros, each sequence being marked "legal" or "illegal" according to grammar to be learned.
The present study demonstrates the efficacy of a recurrent artificial neural network to provide a high fidelity time-dependent nonlinear reduced-order model (ROM) for flutter/limit-cycle oscillation (LCO) modeling. An artificial neural network is a relatively straightforward nonlinear method for modeling an input-output relationship from a set of known data, for which we use the radial basis function (RBF) with its parameters determined through a training process. The resulting RBF neural network, however, is only static and is not yet adequate for an application to problems of dynamic nature. The recurrent neural network method [1] is applied to construct a reduced order model resulting from a series of high-fidelity time-dependent data of aero-elastic simulations. Once the RBF neural network ROM is constructed properly, an accurate approximate solution can be obtained at a fraction of the cost of a full-order computation. The method derived during the study has been validated for predicting nonlinear aerodynamic forces in transonic flow and is capable of accurate flutter/LCO simulations. The obtained results indicate that the present recurrent RBF neural network is accurate and efficient for nonlinear aero-elastic system analysis
Recurrent neural networks (RNNs) have recently been extensively applied to model the time evolution in fluid dynamics, weather predictions, and even chaotic systems due to their ability to capture temporal dependencies and sequential patterns in data. Here we present an RNN model based on convolutional neural networks for modeling the nonlinear nonadiabatic dynamics of hybrid quantum-classical systems. The dynamical evolution of the hybrid systems is governed by equations of motion for classical degrees of freedom and von Neumann equation for electrons. The Physics-Aware Recurrent Convolution (PARC) neural network structure incorporates a differentiator-integrator architecture that inductively models the spatiotemporal dynamics of generic physical systems. Here, we apply our RNN approach to learn the space-time evolution of a one-dimensional semiclassical Holstein model after an interaction quench. For shallow quenches (small changes in electron-lattice coupling), the deterministic dynamics can be accurately captured using a single-CNN-based recurrent network. In contrast, deep quenches induce chaotic evolution, making long-term trajectory prediction significantly more challenging. Nonetheless, we demonstrate that the PARC-CNN architecture can effectively learn the statistical climate of the Holstein model under deep-quench conditions.
Growing vaccine hesitancy is contributing to the decline in immunization rates for highly contagious, vaccine-preventable childhood diseases. Therefore, there has been a significant interest in understanding how hesitancy is spreading at higher spatio-temporal resolutions, enabling more targeted interventions. Motivated by this, we study the problem of prediction of vaccine hesitancy at the ZIP Code level, referred to as the VaxHesitancy problem. A significant challenge for this problem is the lack of high-resolution data that indicates hesitancy. Here, we develop a hybrid VaxHesSTL framework that combines a Graph Neural Network (GNN) and a Recurrent Neural Network (RNN) to address the VaxHesitancy problem. The GNN uses a ZIP Code-level network to capture spatial signals from neighboring areas, while the RNN models the temporal dynamics present in the data. We train and evaluate VaxHesSTL using a large dataset, namely the All-Payer Claims Databases (APCD), for Virginia, consisting of insurance claims from over five million individuals for six years. We find that an aggregated contact network or graph, developed from a detailed activity-based population network, plays an important role in the performance of VaxHesSTL, compared to graph models based solely on spatial proximity. Experiments demonstrate that VaxHesSTL outperforms a range of state-of-the-art baselines, which rely solely on historical time series data without accounting for spatial relationships. Since hesitancy data at higher spatial resolution is often unavailable or hard to get, we incorporate an active learning approach with our VaxHesSTL framework to optimize the training set without compromising the prediction performance. We find that hesitancy data for only 18% of ZIP Codes selected by active learning allows us to forecast hesitancy for all the ZIP Codes in the Virginia.
We report a recurrent neural network (RNN) based model is developed as a surrogate to predict nonlinear plastic response under multiaxial loading. The RNN-based model is trained and tested on stress versus strain curves generated using a numerical solution based on the classical radial return method. Besides simply learning the basic constitutive relationship, a novel approach is taken to enforce certain physical conditions. Specifically, regularization is employed to maintain non-negative plastic power density throughout the loading history thereby ensuring monotonically increasing plastic work and thermodynamic consistency. Enforcing physics in this manner permits coupling of the data-driven RNN approach with physics-based knowledge and laws. This has the effect of reducing the necessary amount of data and ensuring known physical laws are not violated. Since, once trained, the model need not perform the expensive task of solving nonlinear equations, its efficiency is orders of magnitude greater than its numerical counterpart. The RNN-based model has been trained on varied sets of data and the accuracy on test datasets validated. The developed model is general and robust and has widespread application such as in the simulation of metal forming, large scale plasticity, and part life prediction.
Recurrent networks of polynomial threshold elements with random symmetric interactions are studied. Precise asymptotic estimates are derived for the expected number of fixed points as a function of the margin of stability. In particular, it is shown that there is a critical range of margins of stability (depending on the degree of polynomial interaction) such that the expected number of fixed points with margins below the critical range grows exponentially with the number of nodes in the network, while the expected number of fixed points with margins above the critical range decreases exponentially with the number of nodes in the network. The random energy model is also briefly examined and links with higher order neural networks and higher order spin glass models made explicit.
Charged track reconstruction is a critical task in nuclear physics experiments, enabling the identification and analysis of particles produced in high-energy collisions. Machine learning (ML) has emerged as a powerful tool for this purpose, addressing the challenges posed by complex detector geometries, high event multiplicities, and noisy data. Traditional methods rely on pattern recognition algorithms like the Kalman filter, but ML techniques, such as neural networks, graph neural networks (GNNs), and recurrent neural networks (RNNs), offer improved accuracy and scalability. By learning from simulated and real detector data, ML models can identify and classify tracks, predict trajectories, and handle ambiguities caused by overlapping or missing hits. Moreover, ML-based approaches can process data in near-real-time, enhancing the efficiency of experiments at large-scale facilities like the Large Hadron Collider (LHC) and Jefferson Lab (JLAB). As detector technologies and computational resources evolve, ML-driven charged track reconstruction continues to push the boundaries of precision and discovery in nuclear physics. In these proceedings, we highlight advancements in charged track identification leveraging Artificial Intelligence within the CLAS12 detector, achieving a notable enhancement in experimental statistics compared to traditional methods. Additionally, we showcase real-time event reconstruction capabilities, including the inference of charged particle properties, such as momentum, direction, and species identification, at speeds matching data acquisition rates. These innovations enable the extraction of physics observables directly from the experiment in real-time.
The overall objective of the project was to develop a computationally feasible and user-friendly mechanized process to integrate traditional probabilistic risk assessment (PRA) and dynamic PRA (DPRA) results. Starting with the systematic identification of items in an existing PRA that need dynamic augmentation, the project used a generic 4-loop pressurized reactor (PWR) and 3-loop PWR as example plants. Station blackout (SBO) and large break loss of coolant accident SBLOCA) were selected as the example initiating events. Using the traditional event-tree (ET)/fault-tree (FT) methodology augmented by dynamic evet tree approach, the potential consequences of the initiating events were simulated with RELAP-3D and MELCOR/RASCAL codes to cover Level 1 through Level 3 of PRA. RAVEN and ADAPT software were used to generate Level 1 simulations with RELAP-3D and Level 2/3 simulations with MELCOR (Level 2)/RASCAL (Level 3), respectively. Example branching conditions (BCs) for SBO included AC power recovery time, valve repair failure time, reactor coolant pump leak time/break size and emergency power supply duration to a total of 9. Example BCs for LOCA included off-site power recovery time, diesel generator power recovery time, auxiliary feed water system operation time, safety relief valve failure to open upon demand, reactor coolant pump seal break time and size to a total of 21. Each RELAP-3D simulation (9,587 scenarios) was labelled OK or Core Damage based on the maximum allowed peak clad temperature (2,100oF). Each MELCOR simulation (4610 scenarios) was labeled as Bin over 10rem or Bin 0-10rem based on the dose at the site boundary. The scenarios were clustered based on the criteria above using the mean shift methodology. Classical PRA (CPRA) and DPRA results were compared to identify the ET sequences that need DPRA augmentation. Several approaches were proposed for the incorporation of these sequences into CPRA using clustering with the mean shift methodology, restructuring the CPRA ETs by adding new BCs/sequences, and using the concept of a limit surface. Procedures for decision making regarding the possible consequences of an initiating event (e.g. core damage or not, site evacuation or not) were developed using a convolutional neural network (CNN), a recurrent neural network (RNN) and a transformer neural network (TNN). The project has led to two PhD degrees, three archival journal papers and five refereed conference proceedings.
Meteorological prediction is crucial for various sectors, including agriculture, navigation, daily life, disaster prevention, and scientific research. However, traditional numerical weather prediction (NWP) models are constrained by their high computational resource requirements, while the accuracy of deep learning models remains suboptimal. In response to these challenges, we propose a novel deep learning-based model, the Spatiotemporal Fusion Model (STFM), designed to enhance the accuracy of meteorological predictions. Our model leverages Fifth-Generation ECMWF Reanalysis (ERA5) data and introduces two key components: a spatiotemporal encoder module and a spatiotemporal fusion module. The spatiotemporal encoder integrates the strengths of convolutional neural networks (CNNs) and recurrent neural networks (RNNs), effectively capturing both spatial and temporal dependencies. Meanwhile, the spatiotemporal fusion module employs a dual attention mechanism, decomposing spatial attention into global static attention and channel dynamic attention. This approach ensures comprehensive extraction of spatial features from meteorological data. The combination of these modules significantly improves prediction performance. Experimental results demonstrate that STFM excels in extracting spatiotemporal features from reanalysis data, yielding predictions that closely align with observed values. In comparative studies, STFM outperformed other models, achieving a 7% improvement in ground and high-altitude temperature predictions, a 5% enhancement in the prediction of the u/v components of 10 m wind speed, and an increase in the accuracy of potential height and relative humidity predictions by 3% and 1%, respectively. This enhanced performance highlights STFM’s potential to advance the accuracy and reliability of meteorological forecasting.
Background: The number of applications of deep learning algorithms in bioinformatics is increasing as they usually achieve superior performance over classical approaches, especially, when bigger training datasets are available. In deep learning applications, discrete data, e.g. words or n-grams in language, or amino acids or nucleotides in bioinformatics, are generally represented as a continuous vector through an embedding matrix. Recently, learning this embedding matrix directly from the data as part of the continuous iteration of the model to optimize the target prediction – a process called ‘end-to-end learning’ – has led to state-of-the-art results in many fields. Although usage of embeddings is well described in the bioinformatics literature, the potential of end-to-end learning for single amino acids, as compared to more classical manually-curated encoding strategies, has not been systematically addressed. To this end, we compared classical encoding matrices, namely one-hot, VHSE8 and BLOSUM62, to end-to-end learning of amino acid embeddings for two different prediction tasks using three widely used architectures, namely recurrent neural networks (RNN), convolutional neural networks (CNN), and the hybrid CNN-RNN. Results: By using different deep learning architectures, we show that end-to-end learning is on par with classical encodings for embeddings of the same dimension even when limited training data is available, and might allow for a reduction in the embedding dimension without performance loss, which is critical when deploying the models to devices with limited computational capacities. We found that the embedding dimension is a major factor in controlling the model performance. Surprisingly, we observed that deep learning models are capable of learning from random vectors of appropriate dimension. Conclusion: Our study shows that end-to-end learning is a flexible and powerful method for amino acid encoding. Further, due to the flexibility of deep learning systems, amino acid encoding schemes should be benchmarked against random vectors of the same dimension to disentangle the information content provided by the encoding scheme from the distinguishability effect provided by the scheme.