Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Neural network embeddings”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Combining protein sequences and structures with transformers and equivariant graph neural networks to predict protein function

Abstract Motivation Millions of protein sequences have been generated by numerous genome and transcriptome sequencing projects. However, experimentally determining the function of the proteins is still a time consuming, low-throughput, and expensive process, leading to a large protein sequence-function gap. Therefore, it is important to develop computational methods to accurately predict protein function to fill the gap. Even though many methods have been developed to use protein sequences as input to predict function, much fewer methods leverage protein structures in protein function prediction because there was lack of accurate protein structures for most proteins until recently. Results We developed TransFun—a method using a transformer-based protein language model and 3D-equivariant graph neural networks to distill information from both protein sequences and structures to predict protein function. It extracts feature embeddings from protein sequences using a pre-trained protein language model (ESM) via transfer learning and combines them with 3D structures of proteins predicted by AlphaFold2 through equivariant graph neural networks. Benchmarked on the CAFA3 test dataset and a new test dataset, TransFun outperforms several state-of-the-art methods, indicating that the language model and 3D-equivariant graph neural networks are effective methods to leverage protein sequences and structures to improve protein function prediction. Combining TransFun predictions and sequence similarity-based predictions can further increase prediction accuracy. Availability and implementation The source code of TransFun is available at https://github.com/jianlin-cheng/TransFun.

59 BASIC BIOLOGICAL SCIENCES↗

Self-Supervised Cloud Classification

Abstract Low-level marine clouds play a pivotal role in Earth’s weather and climate through their interactions with radiation, heat and moisture transport, and the hydrological cycle. These interactions depend on a range of dynamical and microphysical processes that result in a broad diversity of cloud types and spatial structures, and a comprehensive understanding of cloud morphology is critical for continued improvement of our atmospheric modeling and prediction capabilities moving forward. Deep learning has recently accelerated our ability to study clouds using satellite remote sensing, and machine learning classifiers have enabled detailed studies of cloud morphology. A major limitation of deep learning approaches to this problem, however, is the large number of hand-labeled samples that are required for training. This work applies a recently developed self-supervised learning scheme to train a deep convolutional neural network (CNN) to map marine cloud imagery to vector embeddings that capture information about mesoscale cloud morphology and can be used for satellite image classification. The model is evaluated against existing cloud classification datasets and several use cases are demonstrated, including training cloud classifiers with very few labeled samples, interrogation of the CNN’s learned internal feature representations, cross-instrument application, and resilience against sensor calibration drift and changing scene brightness. The self-supervised approach learns meaningful internal representations of cloud structures and achieves comparable classification accuracy to supervised deep learning methods without the expense of creating large hand-annotated training datasets. Significance Statement Marine clouds heavily influence Earth’s weather and climate, and improved understanding of marine clouds is required to improve our atmospheric modeling capabilities and physical understanding of the atmosphere. Recently, deep learning has emerged as a powerful research tool that can be used to identify and study specific marine cloud types in the vast number of images collected by Earth-observing satellites. While powerful, these approaches require hand-labeling of training data, which is prohibitively time intensive. This study evaluates a recently developed self-supervised deep learning method that does not require human-labeled training data for processing images of clouds. We show that the trained algorithm performs competitively with algorithms trained on hand-labeled data for image classification tasks. We also discuss potential downstream uses and demonstrate some exciting features of the approach including application to multiple satellite instruments, resilience against changing image brightness, and its learned internal representations of cloud types. The self-supervised technique removes one of the major hurdles for applying deep learning to very large atmospheric datasets.

54 ENVIRONMENTAL SCIENCES↗

Learning Distributed Geometric Koopman Operator for Sparse Networked Dynamical Systems

Koopman operator theory provides an alternative to study nonlinear networked dynamical systems by mapping the state space to an abstract higher dimensional space where the system evolution is linear. Recent works show the application of graph neural networks (GNNs) to learn state to object-centric embeddings and achieve centralized block-wise computation of Koopman operator (KO) under additional assumptions on the underlying agents properties and constraints on the KO structure. However, the computational complexity of learning the Koopman increases exponentially for networked systems where the number of possible system states grows in a combinatorial fashion with the number of nodes. The learning challenge is further amplified for sparse networks by two factors: 1) sample sparsity for learning the Koopman operator in the non-linear space, and 2) the divergence in the dynamics of individual nodes or from one subgraph to another. Our work aims to address these challenge by formulating the representation learning of networked dynamical systems into a multi-agent paradigm and learning the Koopman operator in a distributive manner. The computational as well as performance advantages of distributed Koopman is predominant for sparse networks whereas for fully connected networks, it is shown to coincide with the centralized one. The empirical study on rope system, network of oscillators and a synthetic power system show comparable and superior performance along with computational benefits with the state-of-the-art methods.

Mukherjee, Sayak↗

Increasing Mosquito Abundance Under Global Warming

Mosquitoes are a key virus vector that poses significant health threats globally, affecting 700 million individuals and causing 1 million deaths annually. Accurately predicting mosquito abundance and dispersion remains a challenge. Complex interactions between mosquito dynamics and various environmental factors, notably hydrology, contribute to this challenge. Existing models typically focus on precipitation and temperature and often overlook further impacts of hydrological variables within mosquito modeling. In this study, we developed an artificial intelligence‐based model for mosquito dynamics, explicitly accounting for different hydrological variables, such as precipitation, soil moisture and streamflow. Using Toronto, Canada, as a case study, we identified causal relationships between changes in mosquito populations, hydrological factors, vegetation (e.g., leaf area index), and climate variables (e.g., daylight length, precipitation, and temperature). We embedded these relationships into a Long Short‐Term Memory (LSTM) Neural Network Model capable of accurately detecting mosquito dynamics across annual, seasonal, and monthly time scales. The LSTM is able to explain, on average, approximately 40% of the variance in the observed mosquito abundance data. Using the calibrated model, we predicted that the summer season mosquito abundance would increase by ∼16% and ∼19% under an intermediate greenhouse emission scenario, Shared Socioeconomic Pathway (SSP) 2–4.5, and a high greenhouse emission scenario, SSP5‐8.5, respectively. We expect that this model can serve as a valuable tool and inform science‐based decisions affecting mosquito dynamics and public health. It can also build a foundation for future risk analysis at the regional and larger scales.

54 ENVIRONMENTAL SCIENCES↗

Michel Electron Selection with SPINE for DUNE Far Detector Simulation

Michel electrons are a valuable input for particle detector calibration due to their consistent kinetic energy distribution. This report details the evaluation of a Michel electron identification method's application to simulated data from the DUNE (Deep Underground Neutrino Experiment) far detector. This method, which relies on the neural network-based particle classification software SPINE (Scalable Particle Imaging with Neural Embeddings), was developed and calibrated using simulated data for the SBND (Short-Baseline Neutrino Detector) experiment before being applied to simulated DUNE data from a 1x2x6 subset of far detector modules.

Wilson, Dante [Colorado State U.]↗

CoolPINNs: A physics-informed neural network modeling of active cooling in vascular systems

Emerging technologies like hypersonic aircraft, space exploration vehicles, and batteries avail fluid circulation in embedded microvasculatures for efficient thermal regulation. Modeling is vital during the design and operational phases of these engineered systems. However, many challenges exist in developing a modeling framework. What is lacking is an accurate framework that (i) captures sharp jumps in the thermal flux across complex vasculature layouts, (ii) deals with oblique derivatives (involving tangential and normal components), (iii) handles nonlinearity because of radiative heat transfer, (iv) provides a high-speed forecast for real-time monitoring, and (v) facilitates robust inverse modeling. Here, this paper addresses these challenges by availing the power of physics-informed neural networks (PINNs). We develop a fast, reliable, and accurate Scientific Machine Learning (SciML) framework for vascular-based thermal regulation—called CoolPINNs: a PINNs-based modeling framework for active cooling. The proposed mesh-less framework elegantly overcomes all the mentioned challenges. The significance of the reported research is multi-fold. First, the framework is valuable for real-time monitoring of thermal regulatory systems because of rapid forecasting. Second, researchers can address complex thermoregulation designs since the approach is meshless. Finally, the framework facilitates systematic parameter identification and inverse modeling studies, perhaps the most significant utility of the current framework.

97 MATHEMATICS AND COMPUTING↗

Reinforcement Learning via Gaussian Processes with Neural Network Dual Kernels

While deep neural networks (DNNs) and Gaussian Processes (GPs) are both popularly utilized to solve problems in reinforcement learning, both approaches feature undesirable drawbacks for challenging problems. DNNs learn complex non-linear embeddings, but do not naturally quantify uncertainty and are often data-inefficient to train. GPs infer posterior distributions over functions, but popular kernels exhibit limited expressivity on complex and high-dimensional data. Fortunately, recently discovered conjugate and neural tangent kernel functions encode the behavior of overparameterized neural networks in the kernel domain. We demonstrate that these kernels can be efficiently applied to regression and reinforcement learning problems by analyzing a baseline case study.We apply GPs with neural network dual kernels to solve reinforcement learning tasks for the first time. We demonstrate, using the well understood mountain-car problem, that GPs empowered with dual kernels perform at least as well as those using the conventional radial basis function kernel. Finally, we conjecture that by inheriting the probabilistic rigor of GPs and the powerful embedding properties of DNNs, GPs using NN dual kernels will empower future reinforcement learning models on difficult domains.

97 MATHEMATICS AND COMPUTING↗

Augmenting Graph Convolution with Distance Preserving Embedding for Improved Learning

Graph convolution incorporates topological information of a graph into learning. Message passing corresponds to traversal of a local neighborhood in classical graph algorithms. We show that incorporating additional global structures, such as shortest paths, through distance preserving embedding can improve performance. Our approach, Gavotte, significantly improves the performance of a range of popular graph neu-ral networks such as GCN, GA T,Graph SAGE, and GCNII for transductive learning. Gavotte also improves the performance of graph neural networks for full-supervised tasks, albeit to a smaller degree. As high-quality embeddings are generated by Gavotte as a by-product, we leverage clustering algorithms on these embed dings to augment the training set and introduce Gavotte+. Our results of Gavotte+ on datasets with very few labels demonstrate the advantage of augmenting graph convolution with distance preserving embedding.

Cong, Guojing↗

Automatic Generation of High-Performance Convolution Kernels on ARM CPUs for Deep Learning

In this work, we present FastConv, a template-based code auto-generation open source library that can automatically generate high-performance deep learning convolution kernels of arbitrary matrices/tensors shapes. FastConv is based on the Winograd algorithm, which is reportedly the highest performing algorithm for the time-consuming convolution layers of convolutional neural networks. ARM CPUs cover a wide range designs and specifications, from embedded devices to HPC-grade CPUs. The leads to the dilemma of how to consistently optimize Winograd-based convolution solvers for convolution layers of different shapes. FastConv addresses this problem by using templates to auto-generate multiple shapes of tuned kernels variants suitable for skinny tall matrices. As a performance portable library, FastConv transparently searches for the best combination of kernel shapes, cache tiles, scheduling of loop orders, packing strategies, access patterns, and online/offline computations. Auto-tuning is used to search the parameter configuration space for the best performance for a given target architecture and problem size. The experiments with layer-wise evaluation on the VGG--16 model confirms a 1.25x performance gains is got by tuning the Winograd library. Integrated comparison results shows 1.02x to 1.40x, 1.14x to 2.17x, and 1.22x and 2.48x speedup is achieved over NNPACK, Arm NN, and FeatherCNN on the Kunpeng 920 beside few cases. Furthermore, problem size performance portability experiments with various convolution shapes shows that FastConv achieves 1.2x to 1.7x speedup and 2x to 22x speedup over NNPACK and ARM NN inference engine using Winograd on Kunpeng 920 . CPU performance portability evaluation on the VGG--16 show an average speedup over NNPACK of 1.42x, 1.21x, 1.26x, 1.37x, 2.26x, and 11.02x is observed on Kunpeng 920, Snapdragon 835, 855, 888, Apple M1, and AWS Graviton2, respectively.

97 MATHEMATICS AND COMPUTING↗

Transformer Neural Networks with Spatiotemporal Attention for Predictive Control and Optimization of Industrial Processes

In the context of real-time optimization and model predictive control of industrial systems, machine learning, and neural networks represent cutting-edge tools that hold promise for enhancing dynamic modeling. This work presents a novel transformer neural network architecture for real-time optimization and model predictive control. This network design includes a modified attention mechanism inspired by positional embedding attention from vision transformers and task-specific modifications to the input-output structure of the transformer’s decoder stack. Experiments were conducted using data from a 450 MW coal-fired power plant to evaluate this approach's effectiveness. The transformer neural network was compared with conventional recurrent models, including GRU and LSTM. The transformer exhibited a 6% increase in the R-squared (R2) value of predictions and an 83% reduction in mean squared error (MSE). Computation time was also reduced by 84% compared to conventional recurrent models.

Gallup, Ethan R.↗

Deep anomaly detection for industrial systems: a case study

We explore the use of deep neural networks for anomaly detection of industrial systems where the data are multivariate time series measurements. We formulate the problem as a self-supervised learning where data under normal operation are used to train a deep neural network autoregressive model, i.e., use a window of time series data to predict future data values. The aim of such a model is to learn to represent the system dynamic behavior under normal conditions, while expect higher model vs. measurement discrepancies under faulty conditions. In real world applications, many control settings are discrete in nature. In this paper, vector embedding and joint losses are employed to deal with such situations. Both LSTM and CNN based deep neural network backbones are studied on the Secure Water Treatment (SWaT) testbed datasets. Also, Support Vector Data Description (SVDD) method is adapted to such anomaly detection settings with deep neural networks. Evaluation methods and results are discussed based on the SWaT dataset along with potential pitfalls.

anomaly detection, deep neural network↗

FTL: Transfer Learning Nonlinear Plasma Dynamic Transitions in Low Dimensional Embeddings (FTL) v1.0

Fusion Transfer Learning (FTL) model provides a new paradigm to study high-dimensional dynamical behaviors, such as those in fusion plasma systems. The knowledge transfer process leverages a pre-trained neural encoder-decoder network, initially trained on linear simulations, to effectively capture nonlinear dynamics. The low-dimensional embeddings extract the coherent structures of interest, while preserving the inherent dynamics of the complex system. Experimental results highlight FTL's capacity to capture transitional behaviors and dynamical features in plasma dynamics -- a task often challenging for conventional methods. The model developed in this study is generalizable and can be extended broadly through transfer learning to address various magnetohydrodynamics (MHD) modes.

Bai, Zhe↗

Equipping Neural Network Surrogates with Uncertainty for Propagation in Physical Systems

Coarse-grained or filtered models typically rely on closure models to account for unresolved scales. For instance, large eddy simulation for modeling turbulent fluid flows explicitly resolves the largest scales, but requires modeling closure terms to account for the sub-filter scales. With the vast amount of data available from high-fidelity simulations, there are unique opportunities to leverage data-driven modeling techniques to formulate expressive and flexible closure models. Despite their flexibility, data-driven models struggle in domain shift settings, i.e. when deployed in configurations not captured in the training dataset. In particular, the efficacy of neural network surrogates is difficult to assess a priori due to the deterministic, point-estimate nature of predictions. In high-consequence applications, such models require reliable uncertainty estimates in the data-informed and out-of-distribution regimes. To quantify uncertainties in both regimes, we employ Bayesian neural networks which are able to capture both epistemic and aleatoric uncertainties. We will discuss challenges associated with the training and evaluation of these networks. Furthermore, we will discuss uncertainty embedding strategies to enable efficient sampling and propagation of uncertainty through high-fidelity simulations.

Bayesian neural networks↗

Towards POI-based large-scale land use modeling: spatial scale, semantic granularity, and geographic context

The combination of spatial distribution, semantic characteristics, and sometimes temporal dynamics of POIs inside a geographic region can capture its unique land use characteristics. Most previous studies on POI-based land use modeling research focused on one geographic region and select one spatial scale and semantic granularity for land use characterization. There is a lack of understanding on the impact of spatial scale, semantic granularity, and geographic context on POI-based land use modeling, particularly large-scale land use modeling. In this study, we developed a scalable POI-based land use modeling framework and examined the impact of these three factors on POI-based land use characterization using data from three geographic regions. We developed a unified semantic representation framework for POI semantics that can help fuse heterogeneous POI data sources. Then, by combining POIs with a neural network language model, we developed a spatially explicit approach to learn the embedding representation of POIs and AOIs. We trained multiple supervised classifiers using AOI embeddings as input features to predict AOI land use at different semantic granularities. The classification performance of different land use classes was analyzed and compared across three geographic regions to identify the semantic representativeness of POI-based AOI embedding and the impact of geographic context.

58 GEOSCIENCES↗

Clustering of electromagnetic showers and particle interactions with graph neural networks in liquid argon time projection chambers

Liquid argon time projection chambers (LArTPCs) are a class of detectors that produce high resolution images of charged particles within their sensitive volume. In these images, the clustering of distinct particles into superstructures is of central importance to the current and future neutrino physics program. Electromagnetic (EM) activity typically exhibits spatially detached fragments of varying morphology and orientation that are challenging to efficiently assemble using traditional algorithms. Similarly, particles that are spatially removed from each other in the detector may originate from a common interaction. Graph neural networks (GNNs) were developed in recent years to find correlations between objects embedded in an arbitrary space. The graph particle aggregator (GrapPA) first leverages GNNs to predict the adjacency matrix of EM shower fragments and to identify the origin of showers, i.e., primary fragments. On the PILArNet public LArTPC simulation dataset, the algorithm achieves a shower clustering accuracy characterized by a mean purity of 99.4%, a mean efficiency of 99.6% and a primary identification accuracy of 99.8%. It yields a relative shower energy uncertainty of (4.1 + 1.4 / $\sqrt{\text{E(GeV)})}$% and a shower direction uncertainty of (2.1/ $\sqrt{\text{E(GeV)})}$°. Finally, the optimized algorithm is then applied to the related task of clustering particle instances into interactions and yields a mean purity of 99.8% and a mean efficiency of 99.5% for an interaction density of $\mathcal{O}(1)$ m –3 .

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Interactive Visual Study of Multiple Attributes Learning Model of X-Ray Scattering Images

Existing interactive visualization tools for deep learning are mostly applied to the training, debugging, and refinement of neural network models working on natural images. However, visual analytics tools are lacking for the specific application of x-ray image classification with multiple structural attributes. In this paper, we present an interactive system for domain scientists to visually study the multiple attributes learning models applied to x-ray scattering images. It allows domain scientists to interactively explore this important type of scientific images in embedded spaces that are defined on the model prediction output, the actual labels, and the discovered feature space of neural networks. Users are allowed to flexibly select instance images, their clusters, and compare them regarding the specified visual representation of attributes. The exploration is guided by the manifestation of model performance related to mutual relationships among attributes, which often affect the learning accuracy and effectiveness. The system thus supports domain scientists to improve the training dataset and model, find questionable attributes labels, and identify outlier images or spurious data clusters. Case studies and scientists feedback demonstrate its functionalities and usefulness.

97 MATHEMATICS AND COMPUTING↗

Energy efficient photonic memory based on electrically programmable embedded III-V/Si memristors: switches and filters

Abstract Over the past few years, extensive work on optical neural networks has been investigated in hopes of achieving orders of magnitude improvement in energy efficiency and compute density via all-optical matrix-vector multiplication. However, these solutions are limited by a lack of high-speed power power-efficient phase tuners, on-chip non-volatile memory, and a proper material platform that can heterogeneously integrate all the necessary components needed onto a single chip. We address these issues by demonstrating embedded multi-layer HfO 2 /Al 2 O 3 memristors with III-V/Si photonics which facilitate non-volatile optical functionality for a variety of devices such as Mach-Zehnder Interferometers, and (de-)interleaver filters. The Mach-Zehnder optical memristor exhibits non-volatile optical phase shifts > π with ~33 dB signal extinction while consuming 0 electrical power consumption. We demonstrate 6 non-volatile states each capable of 4 Gbps modulation. (De-) interleaver filters were demonstrated to exhibit memristive non-volatile passband transformation with full set/reset states. Time duration tests were performed on all devices and indicated non-volatility up to 24 hours and beyond. We demonstrate non-volatile III-V/Si optical memristors with large electric-field driven phase shifts and reconfigurable filters with true 0 static power consumption. As a result, co-integrated photonic memristors offer a pathway for in-memory optical computing and large-scale non-volatile photonic circuits.

Cheung, Stanley (ORCID:0000000248860013)↗

A reusable neural network pipeline for unidirectional fiber segmentation

Abstract Fiber-reinforced ceramic-matrix composites are advanced, temperature resistant materials with applications in aerospace engineering. Their analysis involves the detection and separation of fibers, embedded in a fiber bed, from an imaged sample. Currently, this is mostly done using semi-supervised techniques. Here, we present an open, automated computational pipeline to detect fibers from a tomographically reconstructed X-ray volume. We apply our pipeline to a non-trivial dataset by Larson et al . To separate the fibers in these samples, we tested four different architectures of convolutional neural networks. When comparing our neural network approach to a semi-supervised one, we obtained Dice and Matthews coefficients reaching up to 98%, showing that these automated approaches can match human-supervised methods, in some cases separating fibers that human-curated algorithms could not find. The software written for this project is open source, released under a permissive license, and can be freely adapted and re-used in other domains.

79 ASTRONOMY AND ASTROPHYSICS↗