Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Recurrent neural network”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Earthquake Nowcasting with Deep Learning

We review previous approaches to nowcasting earthquakes and introduce new approaches based on deep learning using three distinct models based on recurrent neural networks and transformers. We discuss different choices for observables and measures presenting promising initial results for a region of Southern California from 1950–2020. Earthquake activity is predicted as a function of 0.1-degree spatial bins for time periods varying from two weeks to four years. The overall quality is measured by the Nash Sutcliffe efficiency comparing the deviation of nowcast and observation with the variance over time in each spatial region. The software is available as open source together with the preprocessed data from the USGS.

Fox, Geoffrey Charles (ORCID:0000000310171391)↗

Tracing and Forecasting Metabolic Indices of Cancer Patients Using Patient-Specific Deep Learning Models

We develop a patient-specific dynamical system model from the time series data of the cancer patient’s metabolic panel taken during the period of cancer treatment and recovery. The model consists of a pair of stacked long short-term memory (LSTM) recurrent neural networks and a fully connected neural network in each unit. It is intended to be used by physicians to trace back and look forward at the patient’s metabolic indices, to identify potential adverse events, and to make short-term predictions. When the model is used in making short-term predictions, the relative error in every index is less than 10% in the L ∞ norm and less than 6.3% in the L 1 norm in the validation process. Once a master model is built, the patient-specific model can be calibrated through transfer learning. As an example, we obtain patient-specific models for four more cancer patients through transfer learning, which all exhibit reduced training time and a comparable level of accuracy. This study demonstrates that this modeling approach is reliable and can deliver clinically acceptable physiological models for tracking and forecasting patients’ metabolic indices.

60 APPLIED LIFE SCIENCES↗

An Evolve-Then-Correct Reduced Order Model for Hidden Fluid Dynamics

In this paper, we put forth an evolve-then-correct reduced order modeling approach that combines intrusive and nonintrusive models to take hidden physical processes into account. Specifically, we split the underlying dynamics into known and unknown components. In the known part, we first utilize an intrusive Galerkin method projected on a set of basis functions obtained by proper orthogonal decomposition. We then present two variants of correction formula based on the assumption that the observed data are a manifestation of all relevant processes. The first method uses a standard least-squares regression with a quadratic approximation and requires solving a rank-deficient linear system, while the second approach employs a recurrent neural network emulator to account for the correction term. We further enhance our approach by using an orthonormality conforming basis interpolation approach on a Grassmannian manifold to address off-design conditions. The proposed framework is illustrated here with the application of two-dimensional co-rotating vortex simulations under modeling uncertainty. The results demonstrate highly accurate predictions underlining the effectiveness of the evolve-then-correct approach toward real-time simulations, where the full process model is not known a priori.

long short-term memory↗

Data Quality Monitoring for the Hadron Calorimeters Using Transfer Learning for Anomaly Detection

The proliferation of sensors brings an immense volume of spatio-temporal (ST) data in many domains, including monitoring, diagnostics, and prognostics applications. Data curation is a time-consuming process for a large volume of data, making it challenging and expensive to deploy data analytics platforms in new environments. Transfer learning (TL) mechanisms promise to mitigate data sparsity and model complexity by utilizing pre-trained models for a new task. Despite the triumph of TL in fields like computer vision and natural language processing, efforts on complex ST models for anomaly detection (AD) applications are limited. In this study, we present the potential of TL within the context of high-dimensional ST AD with a hybrid autoencoder architecture, incorporating convolutional, graph, and recurrent neural networks. Motivated by the need for improved model accuracy and robustness, particularly in scenarios with limited training data on systems with thousands of sensors, this research investigates the transferability of models trained on different sections of the Hadron Calorimeter of the Compact Muon Solenoid experiment at CERN. The key contributions of the study include exploring TL’s potential and limitations within the context of encoder and decoder networks, revealing insights into model initialization and training configurations that enhance performance while substantially reducing trainable parameters and mitigating data contamination effects.

47 OTHER INSTRUMENTATION↗

NASA Tech Briefs, May 2005

Topics covered include: Fastener Starter; Multifunctional Deployment Hinges Rigidified by Ultraviolet; Temperature-Controlled Clamping and Releasing Mechanism; Long-Range Emergency Preemption of Traffic Lights; High-Efficiency Microwave Power Amplifier; Improvements of ModalMax High-Fidelity Piezoelectric Audio Device; Alumina or Semiconductor Ribbon Waveguides at 30 to 1,000 GHz; HEMT Frequency Doubler with Output at 300 GHz; Single-Chip FPGA Azimuth Pre-Filter for SAR; Autonomous Navigation by a Mobile Robot; Software Would Largely Automate Design of Kalman Filter; Predicting Flows of Rarefied Gases; Centralized Planning for Multiple Exploratory Robots; Electronic Router; Piezo-Operated Shutter Mechanism Moves 1.5 cm; Two SMA-Actuated Miniature Mechanisms; Vortobots; Ultrasonic/Sonic Jackhammer; Removing Pathogens Using Nano-Ceramic-Fiber Filters; Satellite-Derived Management Zones; Digital Equivalent Data System for XRF Labeling of Objects; Identifying Objects via Encased X-Ray-Fluorescent Materials - the Bar Code Inside; Vacuum Attachment for XRF Scanner; Simultaneous Conoscopic Holography and Raman Spectroscopy; Adding GaAs Monolayers to InAs Quantum-Dot Lasers on (001) InP; Vibrating Optical Fibers to Make Laser Speckle Disappear; Adaptive Filtering Using Recurrent Neural Networks; and Applying Standard Interfaces to a Process-Control Language.

Source record↗

Application of Machine Learning to Rotorcraft Health Monitoring

Machine learning is a powerful tool for data exploration and model building with large data sets. This project aimed to use machine learning techniques to explore the inherent structure of data from rotorcraft gear tests, relationships between features and damage states, and to build a system for predicting gear health for future rotorcraft transmission applications. Classical machine learning techniques are difficult, if not irresponsible to apply to time series data because many make the assumption of independence between samples. To overcome this, Hidden Markov Models were used to create a binary classifier for identifying scuffing transitions and Recurrent Neural Networks were used to leverage long distance relationships in predicting discrete damage states. When combined in a workflow, where the binary classifier acted as a filter for the fatigue monitor, the system was able to demonstrate accuracy in damage state prediction and scuffing identification. The time dependent nature of the data restricted data exploration to collecting and analyzing data from the model selection process. The limited amount of available data was unable to give useful information, and the division of training and testing sets tended to heavily influence the scores of the models across combinations of features and hyper-parameters. This work built a framework for tracking scuffing and fatigue on streaming data and demonstrates that machine learning has much to offer rotorcraft health monitoring by using Bayesian learning and deep learning methods to capture the time dependent nature of the data. Suggested future work is to implement the framework developed in this project using a larger variety of data sets to test the generalization capabilities of the models and allow for data exploration.

machine learning↗

Predictive Modeling for Differential Diagnosis and Mortality Risk Assessment

The prevalence of electronic health record (EHR) systems has brought prodigious biomedical informatics opportunity. Automated machine learning methods can effectively utilize such data and have become common tools for healthcare predictive modeling. Researches in medical informatics have explored the potential of deep learning and classical models in emergent care scenarios. In particular, predicting differential diagnoses for admissions have proven useful in decreasing unnecessary lab tests and improving inpatient triage decision-making. Moreover, identification of high-risk patients for in-hospital mortality is vitally important to maximize allocation of medical resources.The Medical Information Mart for Intensive Care (MIMIC-III) database, containing de-identified critical care inpatient was used in our study. This data set captures hospital patient laboratory measurements, pharmacologic prescriptions, diagnostic data and procedure event recordings. When considering adult patients and discounting admissions with ICU length of stay less than 24 hours, there were 37,787 unique admissions and 30,414 total patients. We examined the top 25 most prevalent ICD-9 group-level disease specificities in MIMIC-III using a multi-label classification model. In-hospital mortality was modeled as binary classification with 4,155 (13%) adult patients that expired, of which 3,138 (75.5%) were in the ICU setting. The metrics AUC, F1 score, sensitivity and specificity values calculated for each disease label measured prediction performance.The usage of ICD-9 group codes reduced feature dimension from 14,567 to 942 and greatly improved distribution of patient diagnostic categories. Disease temporal patterns were captured by considering the most frequently sampled 6 vital signs and 13 laboratory values. Missing data were imputed at each time-stamp. Time-series raw hourly average values were converted into 5 summary features (mean, standard deviation, number of observations, min & max values). Patient demographic variables such as age, gender, marital status and ethnicity were also factored into the modeling. Choi et al showed that contextual embedding of medical data, diagnostic and procedural codes alone can predict future diagnoses with sensitivity as high as 0.79. We utilized an embedding technique called word2vec which allowed sparse representations of medical history to be transformed into dense word vectors. The mappings captured contextual information by treating each admission as a sentence and learning the most likely neighboring words in a sliding window fashion. Binary and multi-label classification was achieved via collapse models, which do not consider temporal information, as well as recurrent neural networks with regularization, Softmax output layer activation together with categorical cross-entropy as the loss function.

US Army collaboration↗

Amino Acid Encoding for Deep Learning Applications

Background: The number of applications of deep learning algorithms in bioinformatics is increasing as they usually achieve superior performance over classical approaches, especially, when bigger training datasets are available. In deep learning applications, discrete data, e.g. words or n-grams in language, or amino acids or nucleotides in bioinformatics, are generally represented as a continuous vector through an embedding matrix. Recently, learning this embedding matrix directly from the data as part of the continuous iteration of the model to optimize the target prediction – a process called ‘end-to-end learning’ – has led to state-ofthe-art results in many fields. Although usage of embeddings is well described in the bioinformatics literature, the potential of end-to-end learning for single amino acids, as compared to more classical manually-curated encoding strategies, has not been systematically addressed. To this end, we compared classical encoding matrices, namely one-hot, VHSE8 and BLOSUM62, to end-to-end learning of amino acid embeddings for two different prediction tasks using three widely used architectures, namely recurrent neural networks (RNN), convolutional neural networks (CNN), and the hybrid CNN-RNN. Results: By using different deep learning architectures, we show that end-to-end learning is on par with classical encodings for embeddings of the same dimension even when limited training data is available, and might allow for a reduction in the embedding dimension without performance loss, which is critical when deploying the models to devices with limited computational capacities. We found that the embedding dimension is a major factor in controlling the model performance. Surprisingly, we observed that deep learning models are capable of learning from random vectors of appropriate dimension. Conclusion: Our study shows that end-to-end learning is a flexible and powerful method for amino acid encoding. Further, due to the flexibility of deep learning systems, amino acid encoding schemes should be benchmarked against random vectors of the same dimension to disentangle the information content provided by the encoding scheme from the distinguishability effect provided by the scheme.

Hesham ElAbd↗

Comparison Study of Machine Learning Techniques to Predict Flight Energy Consumption for Advanced Air Mobility

This paper addresses the need to predict the flight energy consumption of aerial vehicles in the presence of wind using machine learning techniques. The presented work is critical to achieving sustainable and efficient operations for Advanced Air Mobility (AAM) and to evaluating the readiness of the ground-supporting energy infrastructure, e.g., electric grid and AAM portals. The flight energy consumption is described using the "energy per meter" (EPM) metric. We present a comparison study of influential machine learning techniques in predicting EPM using real-world flight test data. We presented new results of using the Decision Tree, Random Forest, and linear regression techniques, along with our previous results using the Recurrent Neural Network and Feed Forward Neural Network techniques. The comparison results show that the Linear Regression method outperforms other methods on the basis of the Mean Squared Error and error variance.

Machine Learning↗

Predictive Workload Model for Air Traffic Controllers during UAM Operations

The effect of airspace factors on air traffic controller (ATC) workload has been an active area of study for almost three decades due to the importance of safety considerations necessary to design and maintain operations. Existing literature has examined several traffic-related (e.g., number of aircraft under control, loss of separation) contributors to ATC workload and proposed mathematical functions to best describe controller response. However, future air traffic continues to increase in complexity with the introduction of urban air mobility (UAM) – or the transportation of humans and cargo using electric vertical takeoff and landing (eVTOL) aircraft. UAM aims to alleviate congestion for existing ground transportation systems and improve mobility within urban centers and other high-demand locations. This shift in the traditional airspace paradigm necessitates an evolved understanding of model use and development for ATC workload prediction. This study aimed to develop an ATC workload forecasting model based on human-in-the-loop (HITL) simulation data for UAM operations at large airports. Data collected from the HITL simulation served as the training and testing data for a Long Short-Term Memory recurrent neural network and enabled time-series forecasting of ATC workload from traffic characteristics. Results demonstrated the potential of LSTM models for forecasting ATC workload 40 minutes into the future and highlighted important considerations for future development.

predictive model↗

Predictive Workload Model for Air Traffic Controllers during UAM Operations

The effect of airspace factors on air traffic controller (ATC) workload has been an active area of study for almost three decades due to the importance of safety considerations necessary to design and maintain operations. Existing literature has examined several traffic-related (e.g., number of aircraft under control, loss of separation) contributors to ATC workload and proposed mathematical functions to best describe controller response. However, future air traffic continues to increase in complexity with the introduction of urban air mobility (UAM) – or the transportation of humans and cargo using electric vertical takeoff and landing (eVTOL) aircraft. UAM aims to alleviate congestion for existing ground transportation systems and improve mobility within urban centers and other high-demand locations. This shift in the traditional airspace paradigm necessitates an evolved understanding of model use and development for ATC workload prediction. This study aimed to develop an ATC workload forecasting model based on human-in-the-loop (HITL) simulation data for UAM operations at large airports. Data collected from the HITL simulation served as the training and testing data for a Long Short-Term Memory recurrent neural network and enabled time-series forecasting of ATC workload from traffic characteristics. Results demonstrated the potential of LSTM models for forecasting ATC workload 40 minutes into the future and highlighted important considerations for future development.

predictive model↗

Detecting Anomalous Computation with RNNs on GPU-Accelerated HPC Machines

This paper presents a workload classification framework that discriminates illicit computation from authorized workloads on GPU-accelerated HPC systems. As such systems become more and more powerful, they are exploited by attackers to run malicious and for-profit programs that typically require extremely high computing ability to be successful. Our classification framework leverages the distinctive signatures between illicit and authorized workloads, and explore machine learning methods to learn the workloads and classify them. The framework uses lightweight, non-intrusive workload profiling to collect model input data, and explores multiple machine learning methods, particularly recurrent neural network (RNN) that is suitable for online anomalous workload detection. Evaluation results on three generations of GPU machines demonstrate that the workload classification framework can tell apart the illicit authorized workloads with a high accuracy of over 95%.

Pengfei, Zou↗

CSB-RNN: A Faster-Than-Realtime RNN Acceleration Framework with Compressed Structured Blocks

Recurrent Neural Networks (RNN) is widely applied to temporal sequence analysis, where real-time performance is usually in demand. However, RNN suffers a heavy computational workload as the model comes with a large weight matrix. To alleviate the pain, model compression (pruning) schemes have been proposed for RNN that pruning the redundant (near-zero) weight-values. On the one hand, the non-structured pruning methods achieve a considerable pruning rate while bringing the computational irregularity, which is un-friendly to parallel-hardware. On the other hand, the existing structured pruning methods consider the hardware parallelism; However, they suffer a poor pruning rate due to the restrict constraints on pruning structure. This paper presents CSB-RNN, an optimized full-stack RNN framework with the novel compressed structured block (CSB) technique. The CSB-pruned RNN model comes with both fine-granularity that benefits the pruning rate and regular structure that facilitates the hardware-parallelism. Further, we propose a novel hardware architecture for inferencing the CSB-pruned model. Different from conventional parallel hardware, this architecture solves the block-workload imbalance issue and achieves an over 95% hardware utilization. With the experiments on 10 RNN models in 5 application domains, the CSB-RNN realizes 7×-20× lossless compression and up to 50× acceptable lossy-compression, which is 2×-7× to the prior art. With the addition of the novel hardware, the compressed-RNN inference reaches a super real-time latency of 10-400µs with FPGA implementation.

Shi, Runbin↗

Core-collapse contamination in photometric samples of Type Ia Supernovae

This is an exciting time for cosmology with type Ia supernovae (SNe Ia). The recentlyconcluded Dark Energy Survey SN programme (DES-SN) has obtained the largest anddeepest high-redshift cosmological SN Ia sample, and the Vera Rubin Observatory isexpected to observe at least one order of magnitude more SNe Ia in the next decade.In both these experiments, only a limited fraction (.10 per cent) of the SNe can bespectroscopically classified. This leaves us with large ‘photometric’ SN samples, withthe potential for significant contamination by core-collapse SNe that may bias SN Iacosmological measurements. This thesis demonstrates how this contamination can bemodelled and accounted for in current and future cosmological analyses.First, I present state-of-the-art simulations of the SN universe. These are designedto accurately model the population of SNe Ia, peculiar SNe Ia and core-collapse SNe, aswell as their host galaxies. To improve the diversity and quality of the simulated corecollapseSNe, I build a new library of core-collapse SN templates using spectroscopicand photometric (optical and near-ultraviolet) data of 67 core-collapse SNe from theliterature. I account for our incomplete knowledge of core-collapse SN properties bygenerating a set of SN simulations (rather than a single one), each exploring differentmodelling choices and template libraries. I then characterise selection effects in theDES-SN survey and incorporate them in the simulations, thus obtaining a series ofDES-like simulated SN samples that can be compared to the observed DES-SN data.The agreement between the simulations and data is excellent across many observed SNproperties, including Hubble residuals. These simulations are the first to reproduce theobserved photometric SN and host galaxy properties in high-redshift surveys with no fine-tuning of the input parameters.I use my simulation framework to train and test the performance of SuperNNova,a photometric SN classifier based on recurrent neural networks. I explore differenttraining and validation strategies and show that, across all the DES-SN simulationstested, SuperNNova reduces core-collapse SN contamination to 0.8–3.5 per cent. Ithen show that biases due to contamination on the equation-of-state of dark energy,w, are < 0:008 when using our reference SuperNNova model. This compares to anexpected statistical uncertainty on w from the DES-SN sample of 0:039, thus showing that contamination is not a limiting systematic for the cosmological analysis of theDES-SN sample.The results presented in this thesis are the foundation of the DES SN Ia cosmologicalanalysis; they also provide important implications for the future of SN cosmology,as they demonstrate that contamination is not expected to significantly degrade thecosmological figure of merit of the Rubin SN Ia analysis.

79 ASTRONOMY AND ASTROPHYSICS↗

Identifying Neutrino Final States and Energies in MicroBooNE with New Deep-Learning Based LArTPC Reconstruction Frameworks

MicroBooNE, a Liquid Argon Time Projection Chamber (LArTPC) located in the $\nu_{\mu}$-dominated Booster Neutrino Beam at Fermilab, has been studying $\nu_{e}$ charged-current (CC) interaction rates to shed light on the MiniBooNE low energy excess. The LArTPC technology employed by MicroBooNE provides the capability to image neutrino interactions with mm-scale precision. Computer vision and other machine learning techniques are promising tools for image processing that could boost efficiencies for selecting $\nu_{e}$-CC and other rare signals, reduce cosmic and beam-induced backgrounds, and improve the reconstruction of neutrino energies. The MicroBooNE experiment has been at the forefront of developing and testing such techniques for use in physics analyses. In this poster we overview deep-learning based reconstruction methods. We will showcase the use of a recurrent neural network to estimate neutrino energies and present a new reconstruction framework that uses convolutional neural networks to locate neutrino interaction vertices, tag pixels with track and shower labels, and perform particle identification on reconstructed clusters. We will present studies characterizing the performance of these new tools and demonstrate their effectiveness through their use in an inclusive $\nu_{e}$-CC event selection.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

MSU IETC LSTM Ethernet Decode (AN EDGE)

This research explores the ability of machine learning to perform signal separation of an Ethernet style encoded, full-duplex communication. Typical signal separation currently requires an active tap of the communication line, followed by a recombination and retransmission of the data. The purpose of this research is to study a passive approach to data acquisition from a full-duplex signal. The machine learning model used in this research is a long-short-term memory recurrent neural network (LSTM-RNN). The results show that the LSTM was largely successful in recreating the transmission signal from the measured data points, though the separated signals have not yet been tested using a decoding method.

Full Duplex Signals↗

Recurrent Convolutional Deep Neural Networks for Modeling Time-Resolved Wildfire Spread Behavior

The increasing incidence and severity of wildfires underscores the necessity of accurately predicting their behavior. While high-fidelity models derived from first principles offer physical accuracy, they are too computationally expensive for use in real-time fire response. Low-fidelity models sacrifice some physical accuracy and generalizability via the integration of empirical measurements, but enable real-time simulations for operational use in fire response. Machine learning techniques have demonstrated the ability to bridge these objectives by learning first-principles physics while achieving computational speedups. While deep learning approaches have demonstrated the ability to predict wildfire propagation over large time periods, time-resolved fire-spread predictions are needed for active fire management. Here, in this work, we evaluate the ability of deep learning approaches in accurately modeling the time-resolved dynamics of wildfires. We use an autoregressive process in which a convolutional recurrent deep learning model makes predictions that propagate a wildfire over 15 min increments. We apply the model to four simulated datasets of increasing complexity, containing both field fires with homogeneous fuel distribution as well as real-world topologies sampled from the California region of the United States. We show that even after 100 autoregressive predictions representing more than 24 h of simulated fire spread, the resulting models generate stable and realistic propagation dynamics, achieving a Jaccard score between 0.89 and 0.94 when predicting the resulting fire scar. The inference time of the deep learning models are examined and compared, and directions for future work are discussed.

54 ENVIRONMENTAL SCIENCES↗

Recurrent and convolutional neural networks for sequential multispectral optoacoustic tomography ( MSOT ) imaging

Abstract Multispectral optoacoustic tomography (MSOT) is a beneficial technique for diagnosing and analyzing biological samples since it provides meticulous details in anatomy and physiology. However, acquiring high through‐plane resolution volumetric MSOT is time‐consuming. Here, we propose a deep learning model based on hybrid recurrent and convolutional neural networks to generate sequential cross‐sectional images for an MSOT system. This system provides three modalities (MSOT, ultrasound, and optoacoustic imaging of a specific exogenous contrast agent) in a single scan. This study used ICG‐conjugated nanoworms particles (NWs‐ICG) as the contrast agent. Instead of acquiring seven images with a step size of 0.1 mm, we can receive two images with a step size of 0.6 mm as input for the proposed deep learning model. The deep learning model can generate five other images with a step size of 0.1 mm between these two input images meaning we can reduce acquisition time by approximately 71%.

Juhong, Aniwat↗