Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Recurrent neural network”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

Adaptive Neuron Model: An architecture for the rapid learning of nonlinear topological transformations

A method for the rapid learning of nonlinear mappings and topological transformations using a dynamically reconfigurable artificial neural network is presented. This fully-recurrent Adaptive Neuron Model (ANM) network was applied to the highly degenerate inverse kinematics problem in robotics, and its performance evaluation is bench-marked. Once trained, the resulting neuromorphic architecture was implemented in custom analog neural network hardware and the parameters capturing the functional transformation downloaded onto the system. This neuroprocessor, capable of 10(exp 9) ops/sec, was interfaced directly to a three degree of freedom Heathkit robotic manipulator. Calculation of the hardware feed-forward pass for this mapping was benchmarked at approximately 10 microsec.

Tawel, Raoul↗

Controlling basins of attraction in a neural network-based telemetry monitor

The size of the basins of attraction around fixed points in recurrent neural nets (NNs) can be modified by a training process. Controlling these attractive regions by presenting training data with various amount of noise added to the prototype signal vectors is discussed. Application of this technique to signal processing results in a classification system whose sensitivity can be controlled. This new technique is applied to the classification of temporal sequences in telemetry data.

Bell, Benjamin↗

Improving Subsurface Stress Characterization for Carbon Dioxide Storage Projects by Incorporating Machine Learning Techniques

The overall objective of this project is to develop a framework for reliable characterization and prediction of the state of stress in the overburden and underburden (including the basement) in CO 2 storage reservoirs using machine learning and integrated geomechanics and geophysical methods. Specifically, we propose to develop workflow encompassing of technologies and/or methods to predict stress and pressure changes due to CO 2 injection in an active tertiary recovery site and their impacts on subtle fault activation, fractures and occurrence of microseismic events and compare responses to field observations. In this project, we anticipate using dataset from the Farnsworth field Unit (FWU) which is operated by Purdure Petroleum. A novel elastic-waveform VSP inversion technique will be used to estimate high-resolution spatial and temporal changes of elastic moduli in CO 2 storage reservoirs, which will be combined with velocity-stress relationship derived from laboratory tests to obtain subsurface pressure and stress. Clustered microseismic data will be jointly inverted for improved focal mechanisms. Least-squares reverse-time migration of microseismic waveform data will be performed to directly image fracture/fault zones. Additionally, a deep neural network machine learning technique with convolutional and recurrent layers will be used for learning the spectro-temporal structures in microseismic waveforms. The results of this geotechnical data analysis will be integrated to develop a high-resolution 3D mechanical earth model extending from the overburden sealing formations to the underburden including the basement. Mechanical properties will be derived through integration of mechanical logs, tests, available results from chemo-mechanical laboratory tests, and elastic inversion of seismic data using a combination of Bayesian and stochastic methods as well as machine learning technique. Failure features (faults/fractures) will be represented and/or modeled based on seismic and core data analysis. A transient hydrodynamic-geomechanical model will be developed through coupling with the calibrated FWU reservoir simulation model. The full physics coupled model will be used to train a reduced order proxy model using machine learning algorithm for estimating stress which will then be used with appropriate constitutive relationships and forward seismological models to simulate pressure changes and induced microseismicity. An advanced optimization framework will be developed to perform a history match to minimize error between field observations and simulated. The history matched proxy model will be verified against the full-physics equivalent. The field observations that will be used in the coupled model calibration process include pressure/stress inverted from VSP, moment magnitude from microseismic analysis, real time downhole pressure measurements, production and injection data. Parameter sensitivity and uncertainty analysis will be performed to characterize the impact of model parameter uncertainty on stress estimates. The proposed project will have significant impact on future field implementation of the proposed technology. Because the project field site is an ongoing CO 2 EOR development, the value of the new technology will be demonstrated in an operational context and evaluated as a viable risk mitigation strategy. Cost/benefit will be evaluated together with the various commercial incentives for CO 2 sequestration available to oil and gas operators. The extensive available dataset and ongoing data acquisition under the SWP Phase III work plan provides flexibility for investigation of multiple approaches and reduces technical risk.

58 GEOSCIENCES↗

Efficiently modeling neural networks on massively parallel computers

Neural networks are a very useful tool for analyzing and modeling complex real world systems. Applying neural network simulations to real world problems generally involves large amounts of data and massive amounts of computation. To efficiently handle the computational requirements of large problems, we have implemented at Los Alamos a highly efficient neural network compiler for serial computers, vector computers, vector parallel computers, and fine grain SIMD computers such as the CM-2 connection machine. This paper describes the mapping used by the compiler to implement feed-forward backpropagation neural networks for a SIMD (Single Instruction Multiple Data) architecture parallel computer. Thinking Machines Corporation has benchmarked our code at 1.3 billion interconnects per second (approximately 3 gigaflops) on a 64,000 processor CM-2 connection machine (Singer 1990). This mapping is applicable to other SIMD computers and can be implemented on MIMD computers such as the CM-5 connection machine. Our mapping has virtually no communications overhead with the exception of the communications required for a global summation across the processors (which has a sub-linear runtime growth on the order of O(log(number of processors)). We can efficiently model very large neural networks which have many neurons and interconnects and our mapping can extend to arbitrarily large networks (within memory limitations) by merging the memory space of separate processors with fast adjacent processor interprocessor communications. This paper will consider the simulation of only feed forward neural network although this method is extendable to recurrent networks.

Farber, Robert M.↗

On Neural Architectures for Astronomical Time-series Classification with Application to Variable Stars

Despite the utility of neural networks (NNs) for astronomical time-series classification, the proliferation of learning architectures applied to diverse data sets has thus far hampered a direct intercomparison of different approaches. Here we perform the first comprehensive study of variants of NN-based learning and inference for astronomical time series, aiming to provide the community with an overview on relative performance and, hopefully, a set of best-in-class choices for practical implementations. In both supervised and self-supervised contexts, we study the effects of different time-series-compatible layer choices, namely the dilated temporal convolutional neural network (dTCNs), long-short term memory NNs, gated recurrent units and temporal convolutional NNs (tCNNs). Additionally, we also study the efficacy and performance of encoder-decoder (i.e., autoencoder) networks compared to direct classification networks, different pathways to include auxiliary (non-time-series) metadata, and different approaches to incorporate multi-passband data (i.e., multiple time series per source). Performance—applied to a sample of 17,604 variable stars (VSs) from the MAssive Compact Halo Objects (MACHO) survey across 10 imbalanced classes—is measured in training convergence time, classification accuracy, reconstruction error, and generated latent variables. We find that networks with recurrent NNs generally outperform dTCNs and, in many scenarios, yield to similar accuracy as tCNNs. In learning time and memory requirements, convolution-based layers perform better. We conclude by discussing the advantages and limitations of deep architectures for VS classification, with a particular eye toward next-generation surveys such as the Legacy Survey of Space and Time, the Roman Space Telescope, and Zwicky Transient Facility.

79 ASTRONOMY AND ASTROPHYSICS↗

Data source authentication of synchrophasor measurement devices based on 1D-CNN and GRU

Synchrophasor measurement devices (SMDs) have been widely deployed to support real-time monitoring and control of power systems. In the meantime, data spoofing has emerged in recent years. Therefore, it is of great importance to study data authentication algorithms for detecting and defending the data spoofing effectively. Here, a one-dimensional convolutional neural network (1D-CNN) is utilized to extract temporal signatures hidden in frequency, voltage angle and amplitude data; then the gated recurrent unit (GRU) employs these temporal signatures for data source authentication. In case studies, the performances of different algorithms are tested in large-scale power systems with numerous SMDs for the first time, and comparisons among different algorithms show that the proposed algorithm can achieve a higher accuracy of data source authentication with a shorter time window.

47 OTHER INSTRUMENTATION↗

Modular, Hierarchical Learning By Artificial Neural Networks

Modular and hierarchical approach to supervised learning by artificial neural networks leads to neural networks more structured than neural networks in which all neurons fully interconnected. These networks utilize general feedforward flow of information and sparse recurrent connections to achieve dynamical effects. The modular organization, sparsity of modular units and connections, and fact that learning is much more circumscribed are all attractive features for designing neural-network hardware. Learning streamlined by imitating some aspects of biological neural networks.

Baldi, Pierre F.↗

Quantitative assessment of distant recurrence risk in early stage breast cancer using a nonlinear combination of pathological, clinical and imaging variables

Abstract Use of genomic assays to determine distant recurrence risk in patients with early stage breast cancer has expanded and is now included in the American Joint Committee on Cancer staging manual. Algorithmic alternatives using standard clinical and pathology information may provide equivalent benefit in settings where genomic tests, such as OncotypeDx, are unavailable. We developed an artificial neural network (ANN) model to nonlinearly estimate risk of distant cancer recurrence. In addition to clinical and pathological variables, we enhanced our model using intraoperatively determined global mammographic breast density (MBD) and local breast density (LBD). LBD was measured with optical spectral imaging capable of sensing regional concentrations of tissue constituents. A cohort of 56 ER+ patients with an OncotypeDx score was evaluated. We demonstrated that combining MBD/LBD measurements with clinical and pathological variables improves distant recurrence risk prediction accuracy, with high correlation ( r = 0.98) to the OncotypeDx recurrence score.

Nichols, Brandon S.↗

Keras2c: A library for converting Keras neural networks to real-time compatible C

With the growth of machine learning models and neural networks in measurement and control systems comes the need to deploy these models in a way that is compatible with existing systems. Existing options for deploying neural networks either introduce very high latency, require expensive and time consuming work to integrate into existing code bases, or only support a very limited subset of model types. We have therefore developed a new method called Keras2c, which is a simple library for converting Keras/TensorFlow neural network models into real-time compatible C code. It supports a wide range of Keras layers and model types including multidimensional convolutions, recurrent layers, multi-input/output models, and shared layers. Keras2c re-implements the core components of Keras/TensorFlow required for predictive forward passes through neural networks in pure C, relying only on standard library functions considered safe for real-time use. The core functionality consists of ~ 1500 lines of code, making it lightweight and easy to integrate into existing codebases. Keras2c has been successfully tested in experiments and is currently in use on the plasma control system at the DIII-D National Fusion Facility at General Atomics in San Diego.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

On neural networks in identification and control of dynamic systems

This paper presents a discussion of the applicability of neural networks in the identification and control of dynamic systems. Emphasis is placed on the understanding of how the neural networks handle linear systems and how the new approach is related to conventional system identification and control methods. Extensions of the approach to nonlinear systems are then made. The paper explains the fundamental concepts of neural networks in their simplest terms. Among the topics discussed are feed forward and recurrent networks in relation to the standard state-space and observer models, linear and nonlinear auto-regressive models, linear, predictors, one-step ahead control, and model reference adaptive control for linear and nonlinear systems. Numerical examples are presented to illustrate the application of these important concepts.

Phan, Minh↗

A neural network with modular hierarchical learning

This invention provides a new hierarchical approach for supervised neural learning of time dependent trajectories. The modular hierarchical methodology leads to architectures which are more structured than fully interconnected networks. The networks utilize a general feedforward flow of information and sparse recurrent connections to achieve dynamic effects. The advantages include the sparsity of units and connections, the modular organization. A further advantage is that the learning is much more circumscribed learning than in fully interconnected systems. The present invention is embodied by a neural network including a plurality of neural modules each having a pre-established performance capability wherein each neural module has an output outputting present results of the performance capability and an input for changing the present results of the performance capabilitiy. For pattern recognition applications, the performance capability may be an oscillation capability producing a repeating wave pattern as the present results. In the preferred embodiment, each of the plurality of neural modules includes a pre-established capability portion and a performance adjustment portion connected to control the pre-established capability portion.

Baldi, Pierre F.↗

Core reactivity estimation in space reactors using recurrent dynamic networks

A recurrent multilayer perceptron network topology is used in the identification of nonlinear dynamic systems from only the input/output measurements. The identification is performed in the discrete time domain, with the learning algorithm being a modified form of the back propagation (BP) rule. The recurrent dynamic network (RDN) developed is applied for the total core reactivity prediction of a spacecraft reactor from only neutronic power level measurements. Results indicate that the RDN can reproduce the nonlinear response of the reactor while keeping the number of nodes roughly equal to the relative order of the system. As accuracy requirements are increased, the number of required nodes also increases, however, the order of the RDN necessary to obtain such results is still in the same order of magnitude as the order of the mathematical model of the system. It is believed that use of the recurrent MLP structure with a variety of different learning algorithms may prove useful in utilizing artificial neural networks for recognition, classification, and prediction of dynamic systems.

Parlos, Alexander G.↗

GRUMDN: A Multi-Task Model for Predicting Human Patterns-of-Life from Stay Transition Data

Understanding human patterns-of-life (PoL) is essential towards ensuring safe and secure indoor facility environment as well as outdoor urban environment. Prediction of human movement in between places of interest is vital in understanding human PoL. Movement between spaces maybe represented and detected in one of the two forms: 1) trajectories: locations measured at regular time intervals by mobile sensors, bluetooth or GPS sensors; or 2) stay transitions: semantic PoI (points of interest) and stay duration data measurable by eventbased sensors that collect data when a check-in or check-out event is detected. Stay transition data provides a more compressed data format compared to trajectories data, especially in situations with longer stay durations, while preserving the information necessary for PoL analysis. Now as introduced briefly in the paper, our deployed end application (Digital Twin of a facility with non-player characters, besides the interactive user in virtual reality) needed a well-performing and validated AI/ML model for simulating high quality stay transitions behavior. In this study we thus primarily present our findings with developing and validating that model, which is a multi-task neural network for stay transition prediction. The neural network consists of two heads, for corresponding two tasks of stay category prediction and stay duration prediction. We evaluated gated recurrent units and multi-layer perceptrons of varying network sizes for stay category prediction; while mixture density networks, noisy generator-only networks, and generative adversarial networks of varying network sizes for stay duration prediction. We have then evaluated four multi-task models, constructed by combining these specialized models, on their ability to predict stay transition data. We tested our models on datasets from two different cases: 1) a simulation-generated dataset of indoor movement within the HFIR (high flux isotope reactor) nuclear reactor facility at Oak Ridge National Laboratory (ORNL); and 2) the GeoLife human mobility dataset of outdoor urban movement available in literature. Our results indicate that GRUMDN, which combines gated recurrent units (GRU) for stay category prediction task, and mixture density networks (MDN) for stay duration prediction task, did overall outperform other multitask models and the current state-of-the-art.

Gunaratne, Chathika [ORNL] (ORCID:0000000225088745↗

Characterizing the acceleration time of laser-driven ion acceleration with data-informed neural networks

Peak ion energy is an important figure-of-merit in short-pulse, laser-driven ion acceleration and is dependent on an associated acceleration time. Standard metrics for these quantities depend on analytical results such as the self-similar fluid model or empirical models based on relatively small experimental and simulation datasets. In this work we attempt to use a data-informed neural network (NN) as a surrogate model for a large ensemble of PIC simulations to investigate an effective acceleration time. We explore the application of a stacked convolutional and recurrent NN architecture for improved regression by incorporating the time dependencies of the data into the training process. Of particular note is how pretraining a network on lower fidelity data, e.g. 1D analytical results, greatly improves the network's ability to learn more complex, higher fidelity data. Finally, the dependency of the acceleration time on various laser and plasma parameters is explored.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

NeuroCoreX: An Open-Source FPGA-Based Spiking Neural Network Emulator with On-Chip Learning

Spiking Neural Networks (SNNs) are computational models inspired by the event-driven communication and connectivity patterns of biological neural circuits. They enable high energy efficiency and natural support for diverse architectures ranging from layered networks to small-world and graphstructured topologies. In this work, we introduce NeuroCoreX, an open-source, FPGA-based spiking neural network emulator that provides real-time, on-chip learning and flexible network organization. NeuroCoreX supports both feedforward sensory inputs streamed directly from sensors or PCs via UART and recurrent on-chip connectivity, enabling simultaneous processing and learning from external stimuli and internal network dynamics-capabilities rarely available in existing FPGA SNN platforms. The system implements a Leaky Integrate-and-Fire (LIF) neuron model with current-based synapses and supports pair-based STDP learning on both feedforward and recurrent synapses. A lightweight Python interface enables interactive configuration, live monitoring, weight read-back, and experiment control. Importantly, NeuroCoreX is tightly integrated with the SuperNeuroMAT simulator, allowing SNN models to be transferred seamlessly from software to hardware for hardware-in-the-loop development. By combining real-time plasticity, flexible connectivity, and an open-source VHDL implementation, NeuroCoreX provides an extensible and accessible platform for neuromorphic research, algorithm-hardware co-design, and energy-efficient edge intelligence.

Gautam, Ashish [ORNL]↗

Algorithm for Training a Recurrent Multilayer Perceptron

An improved algorithm has been devised for training a recurrent multilayer perceptron (RMLP) for optimal performance in predicting the behavior of a complex, dynamic, and noisy system multiple time steps into the future. [An RMLP is a computational neural network with self-feedback and cross-talk (both delayed by one time step) among neurons in hidden layers]. Like other neural-network-training algorithms, this algorithm adjusts network biases and synaptic-connection weights according to a gradient-descent rule. The distinguishing feature of this algorithm is a combination of global feedback (the use of predictions as well as the current output value in computing the gradient at each time step) and recursiveness. The recursive aspect of the algorithm lies in the inclusion of the gradient of predictions at each time step with respect to the predictions at the preceding time step; this recursion enables the RMLP to learn the dynamics. It has been conjectured that carrying the recursion to even earlier time steps would enable the RMLP to represent a noisier, more complex system.

Parlos, Alexander G.↗

Learning macroscopic internal variables and history dependence from microscopic models

This paper concerns the study of history dependent phenomena in heterogeneous materials in a two-scale setting where the material is specified at a fine microscopic scale of heterogeneities that is much smaller than the coarse macroscopic scale of application. Here, we specifically study a polycrystalline medium where each grain is governed by crystal plasticity while the solid is subjected to macroscopic dynamic loads. The theory of homogenization allows us to solve the macroscale problem directly with a constitutive relation that is defined implicitly by the solution of the microscale problem. However, the homogenization leads to a highly complex history dependence at the macroscale, one that can be quite different from that at the microscale. In this paper, we examine the use of machine-learning, and especially deep neural networks, to harness data generated by repeatedly solving the finer scale model to: (i) gain insights into the history dependence and the macroscopic internal variables that govern the overall response; and (ii) to create a computationally efficient surrogate of its solution operator, that can directly be used at the coarser scale with no further modeling. We do so by introducing a recurrent neural operator (RNO), and show that: (i) the architecture and the learned internal variables can provide insight into the physics of the macroscopic problem; and (ii) that the RNO can provide multiscale, specifically FE 2 , accuracy at a cost comparable to a conventional empirical constitutive relation.

36 MATERIALS SCIENCE↗

Development of a data-driven neural network model for electron thermal transport in NSTX

A data-driven electron thermal transport neural network (ETT-NN) model, trained on TRANSP interpretative analysis results of National Spherical Torus Experiment (NSTX), was developed to enable faster and more accurate ETT computation for spherical tokamaks (STs). The model incorporates both convolutional NNs and recurrent NNs, allowing it to simultaneously account for the spatial and temporal non-localities and multi-scale features of turbulent transport, which have been considered only in a limited manner in conventional models. The model was validated through interpretative analysis and predictive simulations using Tokamak Reactor Integrated Automated Suite for Simulation and Computation, demonstrating relatively high accuracy. Additionally, parameter scans were performed on test discharges known to exhibit specific turbulent modes, such as microtearing mode, trapped electron mode, kinetic ballooning mode, and electron temperature gradient mode. The scanning results revealed that the ETT-NN model exhibits the same trends as those observed in conventional gyrokinetic simulations or theories, while also capturing the global nature of turbulent transport, indicating that the data-driven model accurately reflects the underlying physical characteristics. Furthermore, due to the dimensionless nature of the model, we can feasibly expand its applicability by incorporating data from other devices and uncovering the characteristics of ETT in STs in the future.

NSTX↗