Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “convolutional neural network model”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

The predictive skill of convolutional neural networks models for disease forecasting

In this paper we investigate the utility of one-dimensional convolutional neural network (CNN) models in epidemiological forecasting. Deep learning models, in particular variants of recurrent neural networks (RNNs) have been studied for ILI (Influenza-Like Illness) forecasting, and have achieved a higher forecasting skill compared to conventional models such as ARIMA. In this study, we adapt two neural networks that employ one-dimensional temporal convolutional layers as a primary building block—temporal convolutional networks and simple neural attentive meta-learners—for epidemiological forecasting. We then test them with influenza data from the US collected over 2010-2019. We find that epidemiological forecasting with CNNs is feasible, and their forecasting skill is comparable to, and at times, superior to, plain RNNs. Thus CNNs and RNNs bring the power of nonlinear transformations to purely data-driven epidemiological models, a capability that heretofore has been limited to more elaborate mechanistic/compartmental disease models.

59 BASIC BIOLOGICAL SCIENCES↗

Uncertainty quantification of graph convolution neural network models of evolving processes

The application of neural network models to scientific machine learning tasks has proliferated in recent years. In particular, neural networks have proved to be adept at modeling processes with spatial–temporal complexity. Nevertheless, these highly parameterized models have garnered skepticism in their ability to produce outputs with quantified error bounds over the regimes of interest. Hence there is a need to find uncertainty quantification methods that are suitable for neural networks. In this work we present comparisons of the parametric uncertainty quantification of neural networks modeling complex spatial–temporal processes with Hamiltonian Monte Carlo and Stein variational gradient descent and its projected variant. Specifically we apply these methods to graph convolutional neural network models of evolving systems modeled with recurrent neural network and neural ordinary differential equations architectures. We show that Stein variational inference is a viable alternative to Monte Carlo methods with some clear advantages for complex neural network models. For our exemplars, Stein variational interference gave similar pushed forward uncertainty profiles through time compared to Hamiltonian Monte Carlo, albeit with generally more generous variance. As a result, projected Stein variational gradient descent also produced similar uncertainty profiles to the non-projected counterpart, but large reductions in the active weight space were confounded by the stability of the neural network predictions and the convoluted likelihood landscape.

36 MATERIALS SCIENCE↗

A Convolutional Neural Network Model for Battery Capacity Fade Curve Prediction using Early Life Data

Herein, early prediction of battery performance degradation trends can facilitate research of new materials and cell designs, rapid deployment of batteries in real-world applications, timely replacement of batteries in critical applications, and even the secondary use market. In this study, we design a convolutional neural network model to predict the entire battery capacity fade curve - a critical indicator of battery performance degradation - using first 100 cycles of data (~ three weeks of testing). We use the discharge voltage-capacity curves as input to the model and automate the feature extraction process through the convolutional layers of the network. Our approach can predict the per cycle capacity fade rate and rollover cycle (knee point) in the capacity fade curve, which indicate the onset of rapid capacity decay. On the publicly available graphite/LiFePO 4 battery dataset, optimized networks predict the capacity fade curves, rollover cycle, and end of life with 3.7% (worst-case), 19%, and 17% mean absolute percentage errors, respectively.

25 ENERGY STORAGE↗

Attention-based 3D – convolutional neural network model for mechanical property predictions using visible light images in metal additive manufacturing

Additive manufacturing (AM), while commonly used for rapid prototyping and creating components with complex geometries, has not been widely adopted for critical applications across the aerospace, automotive, defense, energy, and medical industries. This is, in part, due to the challenges of controlling flaws and uncertainty in the mechanical behavior of additively manufactured components. In recent years, there has been an increase in research aimed at predicting the final mechanical properties of additively manufactured components during the printing process. To address these issues, a 3D-CNN model was trained using low-cost in situ visible-light camera data, anomaly classifications, and the chosen process parameters to predict the ultimate tensile strength (UTS), yield strength (YS), total elongation (TE), and uniform elongation (UE). The 3D-CNN layers of the model employed attention mechanisms to prioritize features in the data, thereby improving prediction accuracy. Furthermore, the effect of each process parameter and anomaly class is investigated using attention-based dynamic sigmoid weighted gates to interpret the influence each class has on the final prediction. Different combinations of the in situ data were fed into the 3D-CNN, with varying amounts of image layers, to determine the ideal combination for predicting mechanical properties in situ. Here, the 3D-CNN model achieved mean absolute percentage errors (MAPE) below 5% for both UTS and YS while using only a single camera input and under half of the available image layers.

36 MATERIALS SCIENCE↗

Data-Driven Template Discovery Using Graph Convolutional Neural Networks

Modeling adversarial activities is a critical component of developing high-con?dence indicators of efforts to acquire, fabricate, proliferate, and/or deploy weapons of mass terror (WMTs). Current approaches to generating representative patterns of interest (a.k.a templates) from the real-world domains involve a Subject Matter Expert (SME)-guided manual process. The goal of Data-Driven Template Discovery (DDTD) is to use a (potentially small) set of SME generated templates to discover other previously unknown and interesting templates in an attributed graph. A template is an activity pattern describing a set of interactions among a group of nodes in the graph. The motivation behind DDTD is to expand the original set of templates, without having SMEs craft all the templates by hand. DDTD also provides seed templates to SMEs, to help them construct larger, high-?delity, and scenario-oriented templates. In these cases, obtaining a larger set of templates that are related (contain similar signals) to the original set is of great value. In this work, we propose to use Graph Convolutional Neural Networks (GCNs) to discover new templates that are heavily related to the original set. GCNs are a family of Neural Network (NN) architectures especially designed to work directly on graphs. In contrast to the traditional NNs, that require considerable amounts of labeled data, GCNs do not require a big labeled training set because they can directly leverage the graph structure instead. This property makes GCNs the perfect tool for creating activity templates.

Joaristi, Mikel↗

Recurrent convolutional neural networks for modeling nonadiabatic dynamics of quantum-classical systems

Recurrent neural networks (RNNs) have recently been extensively applied to model the time evolution in fluid dynamics, weather predictions, and even chaotic systems due to their ability to capture temporal dependencies and sequential patterns in data. Here we present an RNN model based on convolutional neural networks for modeling the nonlinear nonadiabatic dynamics of hybrid quantum-classical systems. The dynamical evolution of the hybrid systems is governed by equations of motion for classical degrees of freedom and von Neumann equation for electrons. The Physics-Aware Recurrent Convolution (PARC) neural network structure incorporates a differentiator-integrator architecture that inductively models the spatiotemporal dynamics of generic physical systems. Here, we apply our RNN approach to learn the space-time evolution of a one-dimensional semiclassical Holstein model after an interaction quench. For shallow quenches (small changes in electron-lattice coupling), the deterministic dynamics can be accurately captured using a single-CNN-based recurrent network. In contrast, deep quenches induce chaotic evolution, making long-term trajectory prediction significantly more challenging. Nonetheless, we demonstrate that the PARC-CNN architecture can effectively learn the statistical climate of the Holstein model under deep-quench conditions.

Holstein model↗

Deciphering enhancer sequence using thermodynamics-based models and convolutional neural networks

Abstract Deciphering the sequence-function relationship encoded in enhancers holds the key to interpreting non-coding variants and understanding mechanisms of transcriptomic variation. Several quantitative models exist for predicting enhancer function and underlying mechanisms; however, there has been no systematic comparison of these models characterizing their relative strengths and shortcomings. Here, we interrogated a rich data set of neuroectodermal enhancers in Drosophila, representing cis- and trans- sources of expression variation, with a suite of biophysical and machine learning models. We performed rigorous comparisons of thermodynamics-based models implementing different mechanisms of activation, repression and cooperativity. Moreover, we developed a convolutional neural network (CNN) model, called CoNSEPT, that learns enhancer ‘grammar’ in an unbiased manner. CoNSEPT is the first general-purpose CNN tool for predicting enhancer function in varying conditions, such as different cell types and experimental conditions, and we show that such complex models can suggest interpretable mechanisms. We found model-based evidence for mechanisms previously established for the studied system, including cooperative activation and short-range repression. The data also favored one hypothesized activation mechanism over another and suggested an intriguing role for a direct, distance-independent repression mechanism. Our modeling shows that while fundamentally different models can yield similar fits to data, they vary in their utility for mechanistic inference. CoNSEPT is freely available at: https://github.com/PayamDiba/CoNSEPT.

59 BASIC BIOLOGICAL SCIENCES↗

Characterizing Quantum Classifier Utility in Natural Language Processing Workflows

Quantum Natural Language Processing (QNLP) develops natural language processing (NLP) models for deployment on quantum computers. We explore feature and data prototype selection techniques to address challenges posed by encoding high dimensional features. Our study builds quantum circuit classifiers that includes classical feature pre-processing, quantum embedding and quantum model training. The quantum models are built on 4 or 6 qubits and the quantum neural network (QNN) uses the established bricklayer design. We compare the dependence of model performance (in terms of accuracy and F1 scores) on feature length, embedding gates and parameterized unitary design. We compare the performance of quantum machine learning models to classical convolution neural network model (CNN) on binary and multi-class classification tasks using two datasets of synthetic features and labels. The first is the ECP-CANDLE P3B3 dataset a corpus of synthetically generated cancer pathology reports. The second dataset is extracted from well-known benchmark dataset (MADELON) - features are generated with a combination of informative, repeated and uninformative features. Both datasets are used for binary classification and multi-class classification with 3 classes. We observe robust, accurate performance from all models on the binary classification tasks, but multiclass classification is a challenge for the quantum models-there is a notable decrease in accuracy when using 3 classes. Overall the performance is comparable in terms of recall and accuracy between QNNs and CNNs, even with large datasets. These results provide a point of comparison between quantum and classical models on real-world datasets.

Hamilton, Kathleen↗

Deep learning model to detect various synchrophasor data anomalies

High-density synchrophasors provide valuable information for power grid situational awareness, operation and control. Unfortunately, due to factors including communication instability and hardware failure, their data quality can be greatly deteriorated by anomalies. Since the anomalies can impact the performance of the synchrophasor applications, it is of paramount significance to propose a model to detect anomalies in synchrophasor. In this study, a convolutional neural network model is established to detect and classify the anomalies in the synchrophasor measurements. Additionally, four types of anomalies observed in actual synchrophasors including erroneous patterns, random spikes, missing points and high-frequency interferences are considered in this study. The proposed model is extensively evaluated via field-collected measurements from the synchrophasor network in Jiangsu grid, China. The superior performance of the proposed model indicates the great potential of using deep learning for the detection of abnormal synchrophasor measurements.

42 ENGINEERING↗

Overview of RFID Applications Utilizing Neural Networks

As Radio Frequency Identification (RFID) methods continue to evolve to higher levels of complexity, one form of machine learning is making its appearance. The use of Neural Networks (NN) in the RFID field is steadily increasing, and in the fields of localization and activity recognition, promising results are being shown from a variety of research. RFID applications fall primarily under two types of problems including regression and classification. We analyze RIFD localization techniques which fall under regression, and activity recognition which falls under classification. Many works don’t classify themselves as activity recognition methods, but because they fall under the classification category, we still consider them as activity recognition techniques. This research overviews the Neural Network models in the localization field based on whether they can perform independently of the environment in which they were tested. For activity recognition and accessory fields, the major methods involve tag-based and tag-free approaches. In conclusion, after the models are surveyed, a comparison study is given to examine what may be the cause for increased accuracy between different Neural Network models.

42 ENGINEERING↗

Phase behavior of continuous-space systems: A supervised machine learning approach

The phase behavior of complex fluids is a challenging problem for molecular simulations. Supervised machine learning (ML) methods have shown potential for identifying the phase boundaries of lattice models. In this work, we extend these ML methods to continuous-space systems. We propose a convolutional neural network model that utilizes grid-interpolated coordinates of molecules as input data of ML and optimizes the search for phase transitions with different filter sizes. We test the method for the phase diagram of two off-lattice models, namely, the Widom–Rowlinson model and a symmetric freely jointed polymer blend, for which results are available from standard molecular simulations techniques. The ML results show good agreement with results of previous simulation studies with the added advantage that there is no critical slowing down. We find that understanding intermediate structures near a phase transition and including them in the training set is important to obtain the phase boundary near the critical point. The method is quite general and easy to implement and could find wide application to study the phase behavior of complex fluids.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Gaussian Process Classification for Galaxy Blend Identification in LSST

Abstract A significant fraction of observed galaxies in the Rubin Observatory Legacy Survey of Space and Time (LSST) will overlap at least one other galaxy along the same line of sight, in a so-called “blend.” The current standard method of assessing blend likelihood in LSST images relies on counting up the number of intensity peaks in the smoothed image of a blend candidate, but the reliability of this procedure has not yet been comprehensively studied. Here we construct a realistic distribution of blended and unblended galaxies through high-fidelity simulations of LSST-like images, and from this we examine the blend classification accuracy of the standard peak-finding method. Furthermore, we develop a novel Gaussian process blend classifier model, and show that this classifier is competitive with both the peak finding method as well as with a convolutional neural network model. Finally, whereas the peak-finding method does not naturally assign probabilities to its classification estimates, the Gaussian process model does, and we show that the Gaussian process classification probabilities are generally reliable.

79 ASTRONOMY AND ASTROPHYSICS↗

Generalization of Deep-Learning Models for Classification of Local Distance Earthquakes and Explosions across Various Geologic Settings

Although accurately classifying signals from earthquakes and explosions at local distance (<250 km) remains an important task for seismic network operations, the growing volume of available seismic data presents a challenge for analysts using traditional source discrimination techniques. In recent years, deep-learning models have proven effective at discriminating between low-magnitude earthquakes and explosions measured at local distances, but it is not clear how well these models are capable of generalizing across different geological settings. To address the issue of generalization between regions, we train deep-learning models (convolutional neural networks [CNNs]) on time–frequency representations (scalograms) of three-component earthquake and explosion signals from eight different regions in the continental United States. We explore scenarios where models are trained on data from all regions, individual regions, or all but one region. We find that although CNN models trained on individual regions do not necessarily generalize well across different settings, models trained on multiple regions that include diverse path coverage generalize to new regions, with station-level accuracy of up to 90% or more for data sets from unseen regions. In general, CNN-based discrimination models significantly outperform models based on uncorrected P/S ratio (measured in the 10–18 Hz frequency band), even when CNN models are tested on data from entirely unseen regions.

58 GEOSCIENCES↗

Classification of animal sounds in a hyperdiverse rainforest using convolutional neural networks with data augmentation

To protect tropical forest biodiversity, we need to be able to detect it reliably, cheaply, and at scale. Automated detection of sound producing animals from passively recorded soundscapes via machine-learning approaches is a promising technique towards this goal, but it is constrained by the necessity of large training data sets. Using soundscapes from a tropical forest in Borneo and a Convolutional Neural Network model (CNN), we investigate i) the minimum viable training data set size for accurate prediction of call types (‘sonotypes’), and ii) the extent to which data augmentation and transfer learning can overcome the issue of small and imbalanced training data sets. We found that even relatively high sample sizes (>80 per sonotype) lead to mediocre accuracy, which however improved significantly with data augmentation and transfer learning, including at extremely small sample sizes (3 per sonotype), regardless of taxonomic group or call characteristics. Neither transfer learning nor data augmentation alone achieved high accuracy. Our results suggest that transfer learning and data augmentation could make the use of CNNs to classify species’ vocalizations feasible even for small soundscape-based projects with many rare species. Retraining our open-source model requires only basic programming skills which makes it possible for individual conservation initiatives to match their local context, in order to enable more evidence-informed management of biodiversity.

54 ENVIRONMENTAL SCIENCES↗

Sparse-Data Deep Learning Strategies for Radiographic Non-Destructive Testing

Radiography is an imaging technique used in a variety of applications, such as medical diagnosis, airport security, and nondestructive testing. We present a deep learning system for extracting information from radiographic images. We perform various prediction tasks using our system, including material classification and regression on the dimensions of a given object that is being radiographed. Our system is designed to address the sparse-data issue for radiographic nondestructive testing applications. It uses a radiographic simulation tool for synthetic data augmentation, and it uses transfer learning with a pre-trained convolutional neural network model. Using this system, our preliminary results indicate that the object geometry regression task saw an improvement of 70% in the R-squared value when using a multi-regime model. In addition, we increase the performance of the object material classification tasks by utilizing data from different imaging systems. In particular, using neutron imaging improved the material classification accuracy by 20% when compared to x-ray imaging.

convolutional neural networks↗