Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Recurrent neural network”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

284 records · Page 16

Deep-Learning based Reconstruction of the Shower Maximum $X_{\mathrm{max}}$ using the Water-Cherenkov Detectors of the Pierre Auger Observatory

The atmospheric depth of the air shower maximum X max is an observable commonly used for the determination of the nuclear mass composition of ultra-high energy cosmic rays. Direct measurements of X max are performed using observations of the longitudinal shower development with fluorescence telescopes. At the same time, several methods have been proposed for an indirect estimation of X max from the characteristics of the shower particles registered with surface detector arrays. In this paper, we present a deep neural network (DNN) for the estimation of X max . The reconstruction relies on the signals induced by shower particles in the ground based water-Cherenkov detectors of the Pierre Auger Observatory. The network architecture features recurrent long short-term memory layers to process the temporal structure of signals and hexagonal convolutions to exploit the symmetry of the surface detector array. We evaluate the performance of the network using air showers simulated with three different hadronic interaction models. Thereafter, we account for long-term detector effects and calibrate the reconstructed X max using fluorescence measurements. Finally, we show that the event-by-event resolution in the reconstruction of the shower maximum improves with increasing shower energy and reaches less than 25 g/cm 2 at energies above 2 × 10 19 eV.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Spatio–Temporal Machine Learning for Regional to Continental Scale Terrestrial Hydrology

Integrated hydrologic models can simulate coupled surface and subsurface processes but are computationally expensive to run at high resolutions over large domains. Here we develop a novel deep learning model to emulate subsurface flows simulated by the integrated ParFlow–CLM model across the contiguous US. We compare convolutional neural networks like ResNet and UNet run autoregressively against our novel architecture called the Forced SpatioTemporal RNN (FSTR). The FSTR model incorporates separate encoding of initial conditions, static parameters, and meteorological forcings, which are fused in a recurrent loop to produce spatiotemporal predictions of groundwater. We evaluate the model architectures on their ability to reproduce 4D pressure heads, water table depths, and surface soil moisture over the contiguous US at 1 km resolution and daily time steps over the course of a full water year. The FSTR model shows superior performance to the baseline models, producing stable simulations that capture both seasonal and event–scale dynamics across a wide array of hydroclimatic regimes. The emulators provide over 1,000× speedup compared to the original physical model, which will enable new capabilities like uncertainty quantification and data assimilation for integrated hydrologic modeling that were not previously possible. Our results demonstrate the promise of using specialized deep learning architectures like FSTR for emulating complex process–based models without sacrificing fidelity.

54 ENVIRONMENTAL SCIENCES↗

Spatio-Temporal Deep Graph Network for Event Detection, Localization, and Classification in Cyber-Physical Electric Distribution System

This work proposes a deep graph learning framework to identify, locate, and classify power, cyber, and cyber power events at the distribution system level. The proposed algorithm jointly exploits spatial, temporal, and node-level cyber and physical data features. The developed graph neural network, together with a deep autoencoder, utilizes physical measurements from distribution level phasor measurement units and cyber data from communication network logs. The spatial structure of the synchrophasor measurements and network is incorporated through a weighted adjacency matrix. The temporal structure is incorporated by defining a spatial operation in the gated recurrent unit. This spatio-temporal learning element resides inside a power event detection, localization, and classification module that provides the degree of confidence for an event label. To accurately pinpoint the location of an event to the nearest bus equipped with a measurement unit, a combination of squared error and proximity score is utilized. Also included is a cyber event detection module that employs heteroskedasticity to analyze the significance of various cyber features during different types of attacks. Finally, a dual-bit cyber-power decision table determines the nature of the event. The proposed method is validated on two distribution systems modeled in OPAL-RT/Hypersim with limited phasor measurement units for different possible physical and cyber events. Further analyses include comparison with other state-of-the-art methods and validation in the presence of measurement noise. As a result, our method outperforms existing approaches and achieves an average detection accuracy of 97.97%, F1-score of 96.88%, precision of 96.53%, and recall of 98.57%.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Physics-constrained Deep Recurrent Neural Models of Building Thermal Dynamics

We develop physics-constrained and control-oriented predictive deep learning models for the thermal dynamics of a real-world commercial office building. The proposed method is based on the systematic encoding of physics-based prior knowledge into a structured recurrent neural architecture. Specifically, our model mimics the structure of the building thermal dynamics model and leverages penalty methods to model inequality constraints. Additionally, we use constrained matrix parameterization based on the Perron-Frobenius theorem to bound the eigenvalues of the learned network weights. We interpret the stable eigenvalues as dissipativeness of the learned building thermal model. We demonstrate the effectiveness of the proposed approach on a dataset obtained from an office building with $20$ thermal zones.

Building Energy, structured neural networks↗

A general weight matrix formulation using optimal control

Classical methods from optimal control theory are used in deriving general forms for neural network weights. The network learning or application task is encoded in a performance index of a general structure. Consequently, different instances of this performance index lead to special cases of weight rules, including some well-known forms. Comparisons are made with the outer product rule, spectral methods, and recurrent back-propagation. Simulation results and comparisons are presented.

Farotimi, Oluseyi↗

A deep neural network regressor for phase constitution estimation in the high entropy alloy system Al-Co-Cr-Fe-Mn-Nb-Ni

High Entropy Alloys (HEAs) are composed of more than one principal element and constitute a major paradigm in metals research. The HEA space is vast and an exhaustive exploration is improbable. Therefore, a thorough estimation of the phases present in the HEA is of paramount importance for alloy design. Machine Learning presents a feasible and non-expensive method for predicting possible new HEAs on-the-fly. A deep neural network (DNN) model for the elemental system of: Mn, Ni, Fe, Al, Cr, Nb, and Co is developed using a dataset generated by high-throughput computational thermodynamic calculations using Thermo-Calc. The features list used for the neural network is developed based on literature and freely available databases. A feature significance analysis matches the reported HEAs phase constitution trends on elemental properties and further expands it by providing so far-overlooked features. The final regressor has a coefficient of determination ( r 2 ) greater than 0.96 for identifying the most recurrent phases and the functionality is tested by running optimization tasks that simulate those required in alloy design. The DNN developed constitutes an example of an emulator that can be used in fast, real-time materials discovery/design tasks.

36 MATERIALS SCIENCE↗

Functional expansion representations of artificial neural networks

In the past few years, significant interest has developed in using artificial neural networks to model and control nonlinear dynamical systems. While there exists many proposed schemes for accomplishing this and a wealth of supporting empirical results, most approaches to date tend to be ad hoc in nature and rely mainly on heuristic justifications. The purpose of this project was to further develop some analytical tools for representing nonlinear discrete-time input-output systems, which when applied to neural networks would give insight on architecture selection, pruning strategies, and learning algorithms. A long term goal is to determine in what sense, if any, a neural network can be used as a universal approximator for nonliner input-output maps with memory (i.e., realized by a dynamical system). This property is well known for the case of static or memoryless input-output maps. The general architecture under consideration in this project was a single-input, single-output recurrent feedforward network.

Gray, W. Steven↗

Bit-serial neuroprocessor architecture

A neuroprocessor architecture employs a combination of bit-serial and serial-parallel techniques for implementing the neurons of the neuroprocessor. The neuroprocessor architecture includes a neural module containing a pool of neurons, a global controller, a sigmoid activation ROM look-up-table, a plurality of neuron state registers, and a synaptic weight RAM. The neuroprocessor reduces the number of neurons required to perform the task by time multiplexing groups of neurons from a fixed pool of neurons to achieve the successive hidden layers of a recurrent network topology.

Tawel, Raoul↗

Learning to train neural networks for real-world control problems

Over the past three years, our group has concentrated on the application of neural network methods to the training of controllers for real-world systems. This presentation describes our approach, surveys what we have found to be important, mentions some contributions to the field, and shows some representative results. Topics discussed include: (1) executing model studies as rehearsal for experimental studies; (2) the importance of correct derivatives; (3) effective training with second-order (DEKF) methods; (4) the efficacy of time-lagged recurrent networks; (5) liberation from the tyranny of the control cycle using asynchronous truncated backpropagation through time; and (6) multistream training for robustness. Results from model studies of automotive idle speed control serve as examples for several of these topics.

Feldkamp, Lee A.↗

Physics informed neural network can retrieve rate and state friction parameters from acoustic monitoring of laboratory stick-slip experiments

Various machine learning (ML) and deep learning (DL) techniques have been recently applied to the forecasting of laboratory earthquakes from friction experiments. The magnitude and timing of shear failures in stick-slip cycles are predicted using features extracted from the recorded ultrasonic or acoustic emission (AE) signals. In addition, the Rate and State Friction (RSF) constitutive laws are extensively used to model the frictional behavior of faults. In this work, we use data from shear experiments coupled with passive acoustic (variance, kurtosis, and AE rate) interleaved with active source ultrasonic monitoring (transmitted wave amplitude) to develop physics-informed neural network (PINN) models incorporating the RSF law and AE rate generation equation with wave amplitude serving as a proxy for friction state variable. This PINN framework allows learning RSF parameters from stick-slip experiments rather than measuring them through a series of velocity step experiments. We observe that when the stick-slip cycles are irregular, the PINN models outperform the data-driven DL models. Transfer learning (TL) PINN models are also developed by pre-training on data collected at one normal stress level followed by forecasting shear failures and retrieving RSF parameters at other stress levels (i.e., with different recurrence intervals) after retraining on a limited amount of new data. Our findings suggest that TL models perform better compared to standalone models. Both standalone and TL PINN-estimated RSF parameters and their ground truth values show excellent agreements thus demonstrating that RSF parameters can be retrieved from laboratory stick-slip experiments using the corresponding acoustic data and that the transmitted wave amplitude provides a good representation of the evolving frictional state during stick-slips.

58 GEOSCIENCES↗

Novel symmetry-preserving neural network model for phylogenetic inference

Abstract Motivation Scientists world-wide are putting together massive efforts to understand how the biodiversity that we see on Earth evolved from single-cell organisms at the origin of life and this diversification process is represented through the Tree of Life. Low sampling rates and high heterogeneity in the rate of evolution across sites and lineages produce a phenomenon denoted “long branch attraction” (LBA) in which long nonsister lineages are estimated to be sisters regardless of their true evolutionary relationship. LBA has been a pervasive problem in phylogenetic inference affecting different types of methodologies from distance-based to likelihood-based. Results Here, we present a novel neural network model that outperforms standard phylogenetic methods and other neural network implementations under LBA settings. Furthermore, unlike existing neural network models in phylogenetics, our model naturally accounts for the tree isomorphisms via permutation invariant functions which ultimately result in lower memory and allows the seamless extension to larger trees. Availability and implementation We implement our novel theory on an open-source publicly available GitHub repository: https://github.com/crsl4/nn-phylogenetics.

59 BASIC BIOLOGICAL SCIENCES↗

Lagrangian Characterization of Surface Transport From the Equatorial Atlantic to the Caribbean Sea Using Climatological Lagrangian Coherent Structures and Self‐Organizing Maps

Abstract This study presents an assessment of the transport of suspended material by surface ocean currents, which have a critical role in determining the connectivity and distribution of living and non‐living material. Lagrangian experiments reveal pathways from the Equatorial Atlantic to 10 strategic regions within the Caribbean Sea, determined by considering the space‐time variability of climatological Lagrangian Coherent Structures, which act as recurrent attracting pathways and transport barriers. Due to windage or Stokes drift, wind forcing is a significant factor in determining the spatial locations where particles cluster and the time needed to reach the Caribbean from the Equatorial Atlantic. Pathways shift westward within the Caribbean and take less time to arrive with increasing wind influence. Depending on the wind effect, the particles show higher confluence in different areas of the Caribbean. A case study is presented for the Mexican Caribbean nearshore area, isolated from ocean‐current trajectories. Here, wind weakens the transport barrier responsible for this isolation and causes particle confluence toward that region. Spatial patterns of the Eulerian velocity identified through Self‐Organizing Maps, with time dependence given by their best matching units, can reproduce the characteristic Lagrangian patterns of surface current climate variability. Our study demonstrates the application of tools from dynamical systems and unsupervised neural networks to understand Lagrangian patterns and identify the processes that drive them. These findings improve our understanding of transport mechanisms of suspended material by surface ocean currents in the Western Atlantic and the Caribbean Sea, which is essential for managing and conserving marine ecosystems.

Allende‐Arandía, Ma. Eugenia↗

Robust deep learning framework for constitutive relations modeling

Modeling the full-range deformation behaviors of materials under complex loading and materials conditions is a significant challenge for constitutive relations (CRs) modeling. Here, we propose a general encoder-decoder deep learning framework that can model high-dimensional stress-strain data and complex loading histories with robustness and universal capability. The framework employs an encoder to project high-dimensional input information (e.g., loading history, loading conditions, and materials information) to a lower-dimensional hidden space and a decoder to map the hidden representation to the stress of interest. We evaluated various encoder architectures, including gated recurrent unit (GRU), GRU with attention, temporal convolutional network (TCN), and the Transformer encoder, on two complex stress-strain datasets that were designed to include a wide range of complex loading histories and loading conditions. All architectures achieved excellent test results with an root-mean-square error (RMSE) below 1 MPa. Additionally, we analyzed the capability of the different architectures to make predictions on out-of-domain applications, with an uncertainty estimation based on deep ensembles. The proposed approach provides a robust alternative to empirical/semi-empirical models for CRs modeling, offering the potential for more accurate and efficient materials design and optimization.

36 MATERIALS SCIENCE↗

Large-scale deep learning for metastasis detection in pathology reports

Objectives No existing algorithm can reliably identify metastasis from pathology reports across multiple cancer types and the entire US population. In this study, we develop a deep learning model that automatically detects patients with metastatic cancer by using pathology reports from many laboratories and of multiple cancer types. Materials and Methods We use 60 471 unstructured pathology reports from 4 Surveillance, Epidemiology, and End Results (SEER) registries. The reports were coded into 1 of 3 labels: metastasis negative, metastases positive, or metastasis undetermined. We utilize a task-specific deep neural network trained from scratch and compare its performance with a widely used large language model (LLM). Results Our deep learning architecture trained on task-specific data outperforms a general-purpose LLM, with a recall of 0.894 compared to 0.824. We quantified model uncertainty and used it to defer reports for human review. We found that retaining 72.9% of reports increased recall from 0.894 to 0.969. Discussion A smaller deep learning architecture trained on task-specific data outperforms a general LLM. Equally critical to model performance is the incorporation of uncertainty quantification, achieved here through an abstention mechanism. Conclusions This study’s finding demonstrate the feasibility of developing algorithms to automatically identify metastatic cancer cases from unstructured pathology reports.

machine learning↗