Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “neural encoding”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Advancing molecular machine learning representations with stereoelectronics-infused molecular graphs

Molecular representation is a critical element in our understanding of the physical world and the foundation for modern molecular machine learning. Previous molecular machine learning models have used strings, fingerprints, global features and simple molecular graphs that are inherently information-sparse representations. However, as the complexity of prediction tasks increases, the molecular representation needs to encode higher fidelity information. This work introduces a new approach to infusing quantum-chemical-rich information into molecular graphs via stereoelectronic effects, enhancing expressivity and interpretability. Learning to predict the stereoelectronics-infused representation with a tailored double graph neural network workflow enables its application to any downstream molecular machine learning task without expensive quantum-chemical calculations. We show that the explicit addition of stereoelectronic information substantially improves the performance of message-passing two-dimensional machine learning models for molecular property prediction. We show that the learned representations trained on small molecules can accurately extrapolate to much larger molecular structures, yielding chemical insight into orbital interactions for previously intractable systems, such as entire proteins, opening new avenues of molecular design. Finally, we have developed a web application (simg.cheme.cmu.edu) where users can rapidly explore stereoelectronic information for their own molecular systems.

Boiko, Daniil A↗

Convolutional Neural Networks Based Remote Sensing Scene Classification under Clear and Cloudy Environments

Remote sensing (RS) scene classification has wide applications in the environmental monitoring and geological survey. In the real-world applications, the RS scene images taken by the satellite might have two scenarios: clear and cloudy environments. However, most of existing methods did not consider these two environments simultaneously. In this paper, we assume that the global and local features are iscriminative in either clear or cloudy environments. Many existing Convolution Neural Networks (CNN) based models have made excellent achievements in the image classification, however they somewhat ignored the global and local features in their network structure. In this paper, we propose a new CNN based network (named GLNet) with the Global Encoder and Local Encoder to extract the discriminative global and local features for the RS scene classification, where the constraints for inter-class dispersion and intra-class compactness are embedded in the GLNet training. The experimental results on two publicized RS scene classification datasets show that the proposed GLNet could achieve better performance based on many existing CNN backbones under both clear and cloudy environments.

97 MATHEMATICS AND COMPUTING↗

Non‐Linear Dimensionality Reduction With a Variational Encoder Decoder to Understand Convective Processes in Climate Models

Abstract Deep learning can accurately represent sub‐grid‐scale convective processes in climate models, learning from high resolution simulations. However, deep learning methods usually lack interpretability due to large internal dimensionality, resulting in reduced trustworthiness in these methods. Here, we use Variational Encoder Decoder structures (VED), a non‐linear dimensionality reduction technique, to learn and understand convective processes in an aquaplanet superparameterized climate model simulation, where deep convective processes are simulated explicitly. We show that similar to previous deep learning studies based on feed‐forward neural nets, the VED is capable of learning and accurately reproducing convective processes. In contrast to past work, we show this can be achieved by compressing the original information into only five latent nodes. As a result, the VED can be used to understand convective processes and delineate modes of convection through the exploration of its latent dimensions. A close investigation of the latent space enables the identification of different convective regimes: (a) stable conditions are clearly distinguished from deep convection with low outgoing longwave radiation and strong precipitation; (b) high optically thin cirrus‐like clouds are separated from low optically thick cumulus clouds; and (c) shallow convective processes are associated with large‐scale moisture content and surface diabatic heating. Our results demonstrate that VEDs can accurately represent convective processes in climate models, while enabling interpretability and better understanding of sub‐grid‐scale physical processes, paving the way to increasingly interpretable machine learning parameterizations with promising generative properties.

54 ENVIRONMENTAL SCIENCES↗

Accelerating the Inference of the Exa.TrkX Pipeline

Recently, graph neural networks (GNNs) have been successfully used for a variety of particle reconstruction problems in high energy physics, including particle tracking. The Exa.TrkX pipeline based on GNNs demonstrated promising performance in reconstructing particle tracks in dense environments. It includes five discrete steps: data encoding, graph building, edge filtering, GNN, and track labeling. All steps were written in Python and run on both GPUs and CPUs. In this work, we accelerate the Python implementation of the pipeline through customized and commercial GPU-enabled software libraries, and develop a C++ implementation for inferencing the pipeline. The implementation features an improved, CUDA-enabled fixed-radius nearest neighbor search for graph building and a weakly connected component graph algorithm for track labeling. GNNs and other trained deep learning models are converted to ONNX and inferenced via the ONNX Runtime C++ API. The complete C++ implementation of the pipeline allows integration with existing tracking software. We report the memory usage and average event latency tracking performance of our implementation applied to the TrackML benchmark dataset.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Parameterized Neural Ordinary Differential Equations: Applications to Computational Physics Problems

This work proposes an extension of neural ordinary differential equations (NODEs) by introducing an additional set of ODE input parameters to NODEs. This extension allows NODEs to learn multiple dynamics specified by the input parameter instances. Our extension is inspired by the concept of parameterized ordinary differential equations, which are widely investigated in computational science and engineering contexts, where characteristics of the governing equations vary over the input parameters. We apply the proposed parameterized NODEs (PNODEs) for learning latent dynamics of complex dynamical processes that arise in computational physics, which is an essential component for enabling rapid numerical simulations for time-critical physics applications. For this, we propose an encoder-decoder-type framework, which models latent dynamics as PNODEs. We demonstrate the effectiveness of PNODEs with important benchmark problems from computational physics.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Latent-Space Dynamics for Prediction and Fault Detection in Geothermal Power Plant Operations

This paper presents a latent-space dynamic neural network (LSDNN) model for the multi-step-ahead prediction and fault detection of a geothermal power plant’s operation. The model was trained to learn the dynamics of the power generation process from multivariate time-series data and the effects of exogenous variables, such as control adjustment and ambient temperature. In the LSDNN model, an encoder–decoder architecture was designed to capture cross-correlation among different measured variables. In addition, a latent space dynamic structure was proposed to propagate the dynamics in the latent space to enable prediction. The prediction power of the LSDNN was utilized for monitoring a geothermal power plant and detecting abnormal events. The model was integrated with principal component analysis (PCA)-based process monitoring techniques to develop a fault-detection procedure. The performance of the proposed LSDNN model and fault detection approach was demonstrated using field data collected from a geothermal power plant.

15 GEOTHERMAL ENERGY↗

Advancing spatiotemporal forecasts of CO 2 plume migration using deep learning networks with transfer learning and interpretation analysis

Accurate and timely forecasts of CO 2 plume distribution throughout the injection and post-injection phases are crucial for detecting plume migration, assessing leakage risks, and supporting operational decisions in geologic carbon storage (GCS). Current convolutional neural network-based approaches primarily focus on spatial information and overlook temporal dependencies in plume distributions, thus limiting their ability to capture dynamic movement effects and provide accurate predictions of plume migration. In this work, we propose two deep learning models, Auto-Encoder (AE)-LSTM and Encoder-Decoder (ED)-ConvLSTM, each uniquely designed to capture both spatial and temporal features. We apply the proposed methods to forecast the dynamic distribution of CO 2 plumes based on 108 reservoir simulations over a 30-year injection and a 30-year post-injection period. The results indicate that the ED-ConvLSTM model outperforms the AE-LSTM model in accurately predicting the spatiotemporal dynamics of CO 2 plume migration, achieving R 2 values above 0.99. To provide a deeper understanding of these model predictions, we employ a gradient-based explanation method on the trained models. This approach provides insights into the influence of input variables on plume migration forecasts and uncovers the underlying prediction mechanisms of the proposed models. Furthermore, we introduce a transfer learning technique, enabling fast and accurate plume migration forecasting in the post-injection phase by leveraging the trained model during the injection phase. This reduces the necessity for extensive data collection or re-training. In conclusion, the methods proposed in our work enhances the performance and interpretability of CO 2 plume migration forecasts, thereby facilitating informed decision-making throughout the entire lifecycle of GCS applications.

58 GEOSCIENCES↗

NSGA-PINN: A Multi-Objective Optimization Method for Physics-Informed Neural Network Training

This paper presents NSGA-PINN, a multi-objective optimization framework for the effective training of physics-informed neural networks (PINNs). The proposed framework uses the non-dominated sorting genetic algorithm (NSGA-II) to enable traditional stochastic gradient optimization algorithms (e.g., ADAM) to escape local minima effectively. Additionally, the NSGA-II algorithm enables satisfying the initial and boundary conditions encoded into the loss function during physics-informed training precisely. We demonstrate the effectiveness of our framework by applying NSGA-PINN to several ordinary and partial differential equation problems. In particular, we show that the proposed framework can handle challenging inverse problems with noisy data.

Lu, Binghang (ORCID:0009000160016632)↗

luoyunan/ECNet: First release

ECNet (evolutionary context-integrated neural network) is a deep learning model that guides protein engineering by predicting protein fitness from the sequence. It integrates local evolutionary context from homologous sequences that explicitly model residue-residue epistasis for the protein of interest with the global evolutionary context that encodes rich semantic and structural features from the enormous protein sequence universe.

Luo, Yunan↗

ECNet is an evolutionary context-integrated deep learning framework for protein engineering

Abstract Machine learning has been increasingly used for protein engineering. However, because the general sequence contexts they capture are not specific to the protein being engineered, the accuracy of existing machine learning algorithms is rather limited. Here, we report ECNet (evolutionary context-integrated neural network), a deep-learning algorithm that exploits evolutionary contexts to predict functional fitness for protein engineering. This algorithm integrates local evolutionary context from homologous sequences that explicitly model residue-residue epistasis for the protein of interest with the global evolutionary context that encodes rich semantic and structural features from the enormous protein sequence universe. As such, it enables accurate mapping from sequence to function and provides generalization from low-order mutants to higher-order mutants. We show that ECNet predicts the sequence-function relationship more accurately as compared to existing machine learning algorithms by using ~50 deep mutational scanning and random mutagenesis datasets. Moreover, we used ECNet to guide the engineering of TEM-1 β-lactamase and identified variants with improved ampicillin resistance with high success rates.

59 BASIC BIOLOGICAL SCIENCES↗

Learning nonlinear operators via DeepONet based on the universal approximation theorem of operators

It is widely known that neural networks (NNs) are universal approximators of continuous functions. However, a less known but powerful result is that a NN with a single hidden layer can accurately approximate any nonlinear continuous operator. This universal approximation theorem of operators is suggestive of the structure and potential of deep neural networks (DNNs) in learning continuous operators or complex systems from streams of scattered data. Here, in this work, we thus extend this theorem to DNNs. We design a new network with small generalization error, the deep operator network (DeepONet), which consists of a DNN for encoding the discrete input function space (branch net) and another DNN for encoding the domain of the output functions (trunk net). We demonstrate that DeepONet can learn various explicit operators, such as integrals and fractional Laplacians, as well as implicit operators that represent deterministic and stochastic differential equations. We study different formulations of the input function space and its effect on the generalization error for 16 different diverse applications.

97 MATHEMATICS AND COMPUTING↗

Use of Graph Theory and Neural Networks for Microstructural Classification

Recent advances in materials data analytics have provided new avenues for determining process-structure-property (PSP) linkages in a variety of materials. Machine learning techniques including few-shot learning have increased the efficiency of classifying microscopy images for the purposes of material characterization. Modifications in segmentation also show potential in improving the accuracy of our current pyCHIP classifier. Replacing previous encoders trained on ImageNet with those trained on microscopy images like MicroNet has initially shown better performance at classifying images of irradiated samples. Additionally, different normalization approaches were tested to show no discernable effect on classification. The Louvain method for community detection is analyzed on a set of irradiated samples with different parameters to determine which proved beneficial under what circumstances. We suggest that microscopy experiments be automated in the future using a combination of these techniques to enable high-throughput analyses.

36 MATERIALS SCIENCE↗

Mitigating Catastrophic Forgetting in Deep Learning in a Streaming Setting Using Historical Summary

Recent advancements in scientific equipment and the adaptation of electronics and the Internet of Things (IoT) in our everyday lives resulted in large and complex data production at a high rate. Making meaningful and timely knowledge discovery at a modest cost from this big data is difficult for computing power and storage limitations. Training deep learning models incrementally in a streaming setting can help us with overcoming these limitations. However, in a well-known phenomenon named catastrophic forgetting, incrementally trained models increasingly perform poorly on the past data. To mitigate catastrophic forgetting in training in a streaming setting, we propose constructing a historical summary over time and use the summary with newly arrived data during incremental training. We propose various data summarization techniques such as random sampling, micro clustering, coreset computation, and Auto Encoders to counteract catastrophic forgetting. We built a pipeline for incremental training with a historical summary for training deep learning models for streaming data. We demonstrate the effectiveness of historical summary in mitigating catastrophic forgetting using three case studies involving three different deep learning applications: an Artificial Neural Network (ANN) for classification task on MNIST dataset, a language model (RNN-LM) on the WikiText2 dataset, and a Convolutional Neural Network (CNN), ResNet50 to classify the ImageNet dataset. Through the training of the models, we observe that catastrophic forgetting is evident in ANN and CNN but not in an RNN. For the first task, our method recovers up to 47.9% lost accuracy due to catastrophic forgetting. For the third task, the historical summary recovers classification accuracy by up to 25%. For the second task, though there is not proof of catastrophic forgetting, the training performance (PPL) improves by up to 26% with historical summary.

Dash, Sajal↗

Fast and efficient identification of anomalous galaxy spectra with neural density estimation

ABSTRACT Current large-scale astrophysical experiments produce unprecedented amounts of rich and diverse data. This creates a growing need for fast and flexible automated data inspection methods. Deep learning algorithms can capture and pick up subtle variations in rich data sets and are fast to apply once trained. Here, we study the applicability of an unsupervised and probabilistic deep learning framework, the probabilistic auto-encoder, to the detection of peculiar objects in galaxy spectra from the SDSS survey. Different to supervised algorithms, this algorithm is not trained to detect a specific feature or type of anomaly, instead it learns the complex and diverse distribution of galaxy spectra from training data and identifies outliers with respect to the learned distribution. We find that the algorithm assigns consistently lower probabilities (higher anomaly score) to spectra that exhibit unusual features. For example, the majority of outliers among quiescent galaxies are E+A galaxies, whose spectra combine features from old and young stellar population. Other identified outliers include LINERs, supernovae, and overlapping objects. Conditional modelling further allows us to incorporate additional information. Namely, we evaluate the probability of an object being anomalous given a certain spectral class, but other information such as metrics of data quality or estimated redshift could be incorporated as well. We make our code publicly available.

Böhm, Vanessa↗

A General Framework to Learn Tertiary Structure for Protein Sequence Characterization

During the past five years, deep-learning algorithms have enabled ground-breaking progress towards the prediction of tertiary structure from a protein sequence. Very recently, we developed SAdLSA, a new computational algorithm for protein sequence comparison via deep-learning of protein structural alignments. SAdLSA shows significant improvement over established sequence alignment methods. In this contribution, we show that SAdLSA provides a general machine-learning framework for structurally characterizing protein sequences. By aligning a protein sequence against itself, SAdLSA generates a fold distogram for the input sequence, including challenging cases whose structural folds were not present in the training set. About 70% of the predicted distograms are statistically significant. Although at present the accuracy of the intra-sequence distogram predicted by SAdLSA self-alignment is not as good as deep-learning algorithms specifically trained for distogram prediction, it is remarkable that the prediction of single protein structures is encoded by an algorithm that learns ensembles of pairwise structural comparisons, without being explicitly trained to recognize individual structural folds. As such, SAdLSA can not only predict protein folds for individual sequences, but also detects subtle, yet significant, structural relationships between multiple protein sequences using the same deep-learning neural network. The former reduces to a special case in this general framework for protein sequence annotation.

59 BASIC BIOLOGICAL SCIENCES↗

Theoretical guarantees for permutation-equivariant quantum neural networks

Despite the great promise of quantum machine learning models, there are several challenges one must overcome before unlocking their full potential. For instance, models based on quantum neural networks (QNNs) can suffer from excessive local minima and barren plateaus in their training landscapes. Recently, the nascent field of geometric quantum machine learning (GQML) has emerged as a potential solution to some of those issues. The key insight of GQML is that one should design architectures, such as equivariant QNNs, encoding the symmetries of the problem at hand. Here, we focus on problems with permutation symmetry (i.e., symmetry group $S_n$), and show how to build $S_n$-equivariant QNNs We provide an analytical study of their performance, proving that they do not suffer from barren plateaus, quickly reach overparametrization, and generalize well from small amounts of data. To verify our results, we perform numerical simulations for a graph state classification task. Our work provides theoretical guarantees for equivariant QNNs, thus indicating the power and potential of GQML.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Group-equivariant autoencoder for identifying spontaneously broken symmetries

We introduce the group-equivariant autoencoder (GE autoencoder), a deep neural network (DNN) method that locates phase boundaries by determining which symmetries of the Hamiltonian have spontaneously broken at each temperature. We use group theory to deduce which symmetries of the system remain intact in all phases, and then use this information to constrain the parameters of the GE autoencoder such that the encoder learns an order parameter invariant to these “never-broken” symmetries. This procedure produces a dramatic reduction in the number of free parameters such that the GE-autoencoder size is independent of the system size. We include symmetry regularization terms in the loss function of the GE autoencoder so that the learned order parameter is also equivariant to the remaining symmetries of the system. By examining the group representation by which the learned order parameter transforms, we are then able to extract information about the associated spontaneous symmetry breaking. We test the GE autoencoder on the 2D classical ferromagnetic and antiferromagnetic Ising models, finding that the GE autoencoder (1) accurately determines which symmetries have spontaneously broken at each temperature; (2) estimates the critical temperature in the thermodynamic limit with greater accuracy, robustness, and time efficiency than a symmetry-agnostic baseline autoencoder; and (3) detects the presence of an external symmetry-breaking magnetic field with greater sensitivity than the baseline method. Lastly, we describe various key implementation details, including a quadratic-programming-based method for extracting the critical temperature estimate from trained autoencoders and calculations of the DNN initialization and learning rate settings required for fair model comparisons.

42 ENGINEERING↗

ARENA: Adversary-Resistant Evolving Neural Architectures

Neural networks are becoming the cornerstone for national security prediction tasks. However, designing them requires significant research and trial/error, as they have many hyperparameters, including their computation graph (“architecture”). Neural architecture search (NAS) employs secondary optimizers to search for architectures maximizing objectives like accuracy. Evolutionary algorithms (EAs) are the most used class of optimizer for NAS. However, existing Python libraries for writing EAs limit the complexity of experiments a user can design. In this project, we built ARENA, a Python framework that encodes complex, hyper-realistic EAs. ARENA collects detailed information as it runs and is flexible enough to encode non-EA search algorithms. We tested ARENA on 4 toy optimization problems by encoding 3 search algorithms for each—random search, an EA, and simulated annealing. We also designed an EA that performs NAS on the MNIST dataset. Our experiments suggest the potential for immediate mission impact through solving lab-wide optimization problems.

97 MATHEMATICS AND COMPUTING↗