Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “convolutional neural network model”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Recurrent Convolutional Deep Neural Networks for Modeling Time-Resolved Wildfire Spread Behavior

The increasing incidence and severity of wildfires underscores the necessity of accurately predicting their behavior. While high-fidelity models derived from first principles offer physical accuracy, they are too computationally expensive for use in real-time fire response. Low-fidelity models sacrifice some physical accuracy and generalizability via the integration of empirical measurements, but enable real-time simulations for operational use in fire response. Machine learning techniques have demonstrated the ability to bridge these objectives by learning first-principles physics while achieving computational speedups. While deep learning approaches have demonstrated the ability to predict wildfire propagation over large time periods, time-resolved fire-spread predictions are needed for active fire management. Here, in this work, we evaluate the ability of deep learning approaches in accurately modeling the time-resolved dynamics of wildfires. We use an autoregressive process in which a convolutional recurrent deep learning model makes predictions that propagate a wildfire over 15 min increments. We apply the model to four simulated datasets of increasing complexity, containing both field fires with homogeneous fuel distribution as well as real-world topologies sampled from the California region of the United States. We show that even after 100 autoregressive predictions representing more than 24 h of simulated fire spread, the resulting models generate stable and realistic propagation dynamics, achieving a Jaccard score between 0.89 and 0.94 when predicting the resulting fire scar. The inference time of the deep learning models are examined and compared, and directions for future work are discussed.

54 ENVIRONMENTAL SCIENCES↗

Learning model combining convolutional deep neural network with a self-attention mechanism for AC optimal power flow

Alternating current optimal power flow (OPF) analysis is critical for efficient and reliable operation of power systems. For large systems or repetitive computations, the traditional methods such as the direct and gradient methods, or non-traditional methods, such as the genetic algorithm and simulating annealing, are time-consuming and unsuitable for real-time computing. The work in this paper proposes a novel framework to obtain the optimal solution of power flow in real-time using a combination of convolutional neural networks and a self-attention mechanism. All parameters of the power networks are rearranged in an image-like shape of a multi-channel image where each channel is a two-dimensional matrix. The proposed approach is adaptive with every input size of power systems as well as frequent variations of network topologies without intervention to the framework core. The encompassment of all power system contexts in which all parameters of internal elements, generation costs, and topology information are included, contributes to the higher accuracy of inference compared to other current machine-learning-based OPF-solving methods. Besides, the proposed framework established on ubiquitous platforms is effortlessly integrated into current infrastructures of power systems, and the great efficiency along with the computation speed may serve as a critical point for practical implications, such as enabling faster decision-making during real-time operations, predicting system contingencies, and remedial actions based on an offline pre-trained model. Furthermore, this supervised learning process is applied to the dataset of four case studies of meshed power systems: the IEEE 5-bus system (IEEE-5), the IEEE 30-bus system (IEEE-30), the IEEE 39-bus system (IEEE-39), and the IEEE 57-bus system (IEEE-57) to prove the efficacy of the proposed method.

42 ENGINEERING↗

Neural network based fast prediction of β N limits in HL-2M

Artificial neural networks (NNs) are trained, based on the numerical database, to predict the no-wall and ideal-wall β N limits, due to onset of the n = 1 ( n is the toroidal mode number) ideal external kink instability, for the HL-2M tokamak. The database is constructed by toroidal computations utilizing both the equilibrium code CHEASE (Lütjens et al 1992 Comput. Phys. Commun. 69 287) and the stability code MARS-F (Liu et al 2000 Phys. Plasmas 7 3681). The stability results show that (1) the plasma elongation generally enhances both β N limits, for either positive or negative triangularity plasmas; (2) the effect is more pronounced for positive triangularity plasmas; (3) the computed no-wall β N limit linearly scales with the plasma internal inductance, with the proportionality coefficient ranging between 1 and 5 for HL-2M; (4) the no-wall limit substantially decreases with increasing pressure peaking factor. Furthermore, both the NN model and the convolutional neural network (CNN) model are trained and tested, producing consistent results. The trained NNs predict both the no-wall and ideal-wall limits with as high as 95% accuracy, compared to those directly computed by the stability code. Additional test cases, produced by the Tokamak Simulation Code (Jardin et al 1993 Nucl. Fusion 33 371), also show reasonable performance of the trained NNs, with the relative error being within 10%. The constructed database provides effective references for the future HL-2M operations. The trained NNs can be used as a real-time monitor for disruption prevention in the HL-2M experiments, or serve as part of the integrated modeling tools for ideal kink stability analysis.

Physics↗

Surrogate modeling of Monte Carlo radiation transport with convolutional neural networks for shielding optimization

Here, we present a machine learning (ML)-based surrogate model using convolutional neural networks (CNN) designed to emulate the attenuation of neutron fields as they pass through various shielding materials. This model can compute the outgoing neutron flux almost instantaneously and achieves reasonable accuracy compared to traditional Monte Carlo (MC)-based codes, which are computationally intensive. This emulator alleviates the complexity of neutron radiation transport through shielding materials by reducing the dimensionality and enables shielding optimization for a known radiation environment. This optimization process, which would have taken an unrealistic timeline due to several complex radiation transport simulations, can now be achieved in minutes, thus increasing computational capabilities in radiation shielding assessment. We demonstrate the applications of this emulator in computing effective dose rates and optimizing shielding solutions for a heavy-ion accelerator facility, such as the Facility for Rare Isotope Beams, where secondary neutrons produced via beam interactions dominate the radiation environment.

accelerator shielding↗

DISTEMA: distance map-based estimation of single protein model accuracy with attentive 2D convolutional neural network

Abstract Background Estimation of the accuracy (quality) of protein structural models is important for both prediction and use of protein structural models. Deep learning methods have been used to integrate protein structure features to predict the quality of protein models. Inter-residue distances are key information for predicting protein’s tertiary structures and therefore have good potentials to predict the quality of protein structural models. However, few methods have been developed to fully take advantage of predicted inter-residue distance maps to estimate the accuracy of a single protein structural model. Result We developed an attentive 2D convolutional neural network (CNN) with channel-wise attention to take only a raw difference map between the inter-residue distance map calculated from a single protein model and the distance map predicted from the protein sequence as input to predict the quality of the model. The network comprises multiple convolutional layers, batch normalization layers, dense layers, and Squeeze-and-Excitation blocks with attention to automatically extract features relevant to protein model quality from the raw input without using any expert-curated features. We evaluated DISTEMA’s capability of selecting the best models for CASP13 targets in terms of ranking loss of GDT-TS score. The ranking loss of DISTEMA is 0.079, lower than several state-of-the-art single-model quality assessment methods. Conclusion This work demonstrates that using raw inter-residue distance information with deep learning can predict the quality of protein structural models reasonably well. DISTEMA is freely at https://github.com/jianlin-cheng/DISTEMA

59 BASIC BIOLOGICAL SCIENCES↗

Array-Based Machine Learning for Functional Group Detection in Electron Ionization Mass Spectrometry

Mass spectrometry is a ubiquitous technique capable of complex chemical analysis. The fragmentation patterns that appear in mass spectrometry are an excellent target for artificial intelligence methods to automate and expedite the analysis of data to identify targets such as functional groups. To develop this approach, we trained models on electron ionization (a reproducible hard fragmentation technique) mass spectra so that not only the final model accuracies but also the reasoning behind model assignments could be evaluated. The convolutional neural network (CNN) models were trained on 2D images of the spectra using transfer learning of Inception V3, and the logistic regression models were trained using array-based data and Scikit Learn implementation in Python. Our training dataset consisted of 21,166 mass spectra from the United States’ National Institute of Standards and Technology (NIST) Webbook. The data was used to train models to identify functional groups, both specific (e.g., amines, esters) and generalized classifications (aromatics, oxygen-containing functional groups, and nitrogen-containing functional groups). We found that the highest final accuracies on identifying new data were observed using logistic regression rather than transfer learning on CNN models. It was also determined that the mass range most beneficial for functional group analysis is 0–100 m/z. We also found success in correctly identifying functional groups of example molecules selected from both the NIST database and experimental data. Beyond functional group analysis, we also have developed a methodology to identify impactful fragments for the accurate detection of the models’ targets. The results demonstrate a potential pathway for analyzing and screening substantial amounts of mass spectral data.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Data-Driven State of Health Estimation for Second-Life Batteries Using Interpolated Synthetic Data and Feature Selection

Accurate estimation of the State of Health (SOH) for second-life batteries (SLBs) is crucial given their increasing use in energy storage applications. Precise SOH prediction is essential for safe operation and robust battery management systems. A major challenge is the limited availability of datasets for building reliable degradation models. To address this, synthetic data generation through linear interpolation is performed to extend the available data, making it more representative of real-world battery operating conditions. By analyzing feature correlation with SOH, the most relevant features are selected for the model. The proposed approach employs a convolutional neural network (CNN) model trained on this interpolated, feature-selected dataset, using time series data of voltage, temperature, and current over a cycle. By focusing on highly correlated features, the model achieves over 95% accuracy, with mean absolute error and root mean squared error up to 2.27% and 2.64%, respectively, in SOH estimation for two battery datasets tested. These results highlight the potential of combining synthetic data generation and feature selection to enhance SOH predictions, showcasing the superior performance of the proposed CNN model for both new batteries and SLBs.

feature selection↗

A Modified Sequence-to-point HVAC Load Disaggregation Algorithm

This paper presents a modified sequence-to-point (S2P) algorithm for disaggregating the heat, ventilation, and air conditioning (HVAC) load from the total building electricity consumption. The original S2P model is convolutional neural network (CNN) based, which uses load profiles as inputs. We propose three modifications. First, the input convolution layer is changed from 1D to 2D so that normalized temperature profiles are also used inputs to the S2P model. Second, a drop-out layer is added to improve adaptability and generalizability so that the model trained in one area can be transferred to other geographical areas without labelled HVAC data. Third, a fine-tuning process is proposed for areas with a small amount of labelled HVAC data so that the pre-trained S2P model can be fine-tuned to achieve higher disaggregation accuracy (i.e., better transferability) in other areas. The model is first trained and tested using smart meter and sub-metered HVAC data collected in Austin, Texas. Then, the trained model is tested on two other areas: Boulder, Colorado and San Diego, California. Simulation results show that the proposed modified S2P algorithm outperforms the original S2P model and the support-vector machine based approach in accuracy, adaptability, and transferability.

Ye, Kai↗

Predictive Skill of Deep Learning Models Trained on Limited Sequence Data

In this report we investigate the utility of one-dimensional convolutional neural network (CNN) models in epidemiological forecasting. Deep learning models, especially variants of recurrent neural networks (RNNs) have been studied for influenza forecasting, and have achieved higher forecasting skill compared to conventional models such as ARIMA models. In this study, we adapt two neural networks that employ one-dimensional temporal convolutional layers as a primary building block temporal convolutional networks and simple neural attentive meta-learner for epidemiological forecasting and test them with influenza data from the US collected over 2010-2019. We find that epidemiological forecasting with CNNs is feasible, and their forecasting skill is comparable to, and at times, superior to, RNNs. Thus CNNs and RNNs bring the power of nonlinear transformations to purely data-driven epidemiological models, a capability that heretofore has been limited to more elaborate mechanistic/compartmental disease models.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

PyTorch Implementation of Log-Additive Convolutional Neural Networks

This code is a collection of python code that defines, trains, and tests Log-Additive Convolutional Neural Networks. The model components and training routine are based on the PyTorch python library. The code implements the Log-Additive Convolutional Neural Networks as described in Pagendam et al. 2023. In addition to the Log-Additive Convolutional Neural Networks, this library also defines the Log-Normal Density loss function as described in Pagendam et al. 2023. Code from this paper is not publicly available, so the Pytorch implementation of this type of model is unique to this library.

Callis, Skylar↗

Data-driven cyber-attack detection for photovoltaic systems: A transfer learning approach

With increasing exposure to software-based sensing and control, power systems are facing higher risks of cyber/physical attacks. Here, to ensure system stability and minimize the potential economic losses, it is imperative to monitor the operating states and detect those attacks at the early stage. In this paper, a transfer learning method is proposed to detect cyber-attacks in photovoltaic (PV) systems with much less training data. First of all, two PV systems with a different number of PV inverters and power ratings are analyzed and their attack models are studied. Next, an attack detection Convolutional Neural Network (CNN) model was trained with rich amount of data from PV #1. Then, transfer learning was proposed to transfer the well-trained features from PV #1 to PV #2. Lastly, the attack detection model on PV #2 was trained based on the transferred CNN model. The experiment results show that the proposed transfer learning method achieves better accuracy and a faster convergence rate with a much less training dataset than conventional deep learning.

14 SOLAR ENERGY↗

Evaluating the Trustworthiness of Explainable Artificial Intelligence (XAI) Methods Applied to Regression Predictions of Arctic Sea Ice Motion

Abstract Recent advances in explainable artificial intelligence (XAI) methods show promise for understanding predictions made by machine learning (ML) models. XAI explains how the input features are relevant or important for the model predictions. We train linear regression (LR) and convolutional neural network (CNN) models to make 1-day predictions of sea ice velocity in the Arctic from inputs of present-day wind velocity and previous-day ice velocity and concentration. We apply XAI methods to the CNN and compare explanations to variance explained by LR. We confirm the feasibility of using a novel XAI method [i.e., global layerwise relevance propagation (LRP)] to understand ML model predictions of sea ice motion by comparing it to established techniques. We investigate a suite of linear, perturbation-based, and propagation-based XAI methods in both local and global forms. Outputs from different explainability methods are generally consistent in showing that wind speed is the input feature with the highest contribution to ML predictions of ice motion, and we discuss inconsistencies in the spatial variability of the explanations. Additionally, we show that the CNN relies on both linear and nonlinear relationships between the inputs and uses nonlocal information to make predictions. LRP shows that wind speed over land is highly relevant for predicting ice motion offshore. This provides a framework to show how knowledge of environmental variables (i.e., wind) on land could be useful for predicting other properties (i.e., sea ice velocity) elsewhere. Significance Statement Explainable artificial intelligence (XAI) is useful for understanding predictions made by machine learning models. Our research establishes trustability in a novel implementation of an explainable AI method known as layerwise relevance propagation for Earth science applications. To do this, we provide a comparative evaluation of a suite of explainable AI methods applied to machine learning models that make 1-day predictions of Arctic sea ice velocity. We use explainable AI outputs to understand how the input features are used by the machine learning to predict ice motion. Additionally, we show that a convolutional neural network uses nonlinear and nonlocal information in making its predictions. We take advantage of the nonlocality to investigate the extent to which knowledge of wind on land is useful for predicting sea ice velocity elsewhere.

Hoffman, Lauren [Scripps Institution of Oceanograp↗

A Deep Learning Filter for the Intraseasonal Variability of the Tropics

Abstract This paper presents a novel application of convolutional neural network (CNN) models for filtering the intraseasonal variability of the tropical atmosphere. In this deep learning filter, two convolutional layers are applied sequentially in a supervised machine learning framework to extract the intraseasonal signal from the total daily anomalies. The CNN-based filter can be tailored for each field similarly to fast Fourier transform filtering methods. When applied to two different fields (zonal wind stress and outgoing longwave radiation), the index of agreement between the filtered signal obtained using the CNN-based filter and a conventional weight-based filter is between 95% and 99%. The advantage of the CNN-based filter over the conventional filters is its applicability to time series with the length comparable to the period of the signal being extracted. Significance Statement This study proposes a new method for discovering hidden connections in data representative of tropical atmosphere variability. The method makes use of an artificial intelligence (AI) algorithm that combines a mathematical operation known as convolution with a mathematical model built to reflect the behavior of the human brain known as artificial neural network. Our results show that the filtered data produced by the AI-based method are consistent with the results obtained using conventional mathematical algorithms. The advantage of the AI-based method is that it can be applied to cases for which the conventional methods have limitations, such as forecast (hindcast) data or real-time monitoring of tropical variability in the 20–100-day range.

Stan, Cristiana↗

Interpretable Convolutional Learning Classifier System (C-LCS) for Higher Dimensional Datasets

The purpose of this paper is to devise an interpretable hybrid classification model for Convolutional Neural Networks (CNN) and a Learning Classifier System (LCS). The presented hybrid system integrates the fundamental attributes from both types of these classifiers. In the proposed hybrid model CNN works as an automatic feature extractor, and LCS works to provide interpretable rule-based classification results. Although LCS has limitations working on higher dimensional datasets, we resolve this limitation by using CNN as a feature extractor. The other concept of the non-interpretability of CNN is addressed by using the LCS rule. Furthermore, our experiment with higher dimensional datasets like CIFAR-10 and Fashion-MNIST shows that extended LCS provides comparable performance to the standard neural network model while also providing interpretable results. We named this extended LCS method Convolutional Learning Classifier Cystem (C-LCS).

Jelani Owens↗

A data-driven method for modelling dissipation rates in stratified turbulence

We present a deep probabilistic convolutional neural network (PCNN) model for predicting local values of small-scale mixing properties in stratified turbulent flows, namely the dissipation rates of turbulent kinetic energy and density variance, $\varepsilon$ and $\chi$ . Inputs to the PCNN are vertical columns of velocity and density gradients, motivated by data typically available from microstructure profilers in the ocean. The architecture is designed to enable the model to capture several characteristic features of stratified turbulence, in particular the dependence of small-scale isotropy on the buoyancy Reynolds number $Re_b:=\varepsilon /(\nu N^2)$ , where $\nu$ is the kinematic viscosity and $N$ is the background buoyancy frequency, the correlation between suitably locally averaged density gradients and turbulence intensity and the importance of capturing the tails of the probability distribution functions of values of dissipation. Empirically modified versions of commonly used isotropic models for $\varepsilon$ and $\chi$ that depend only on vertical derivatives of density and velocity are proposed based on the asymptotic regimes $Re_b\ll 1$ and $Re_b\gg 1$ , and serve as an instructive benchmark for comparison with the data-driven approach. When trained and tested on a simulation of stratified decaying turbulence which accesses a range of turbulent regimes (associated with differing values of $Re_b$ ), the PCNN outperforms assumptions of isotropy significantly as $Re_b$ decreases, and additionally demonstrates improvements over the fitted empirical models. A differential sensitivity analysis of the PCNN facilitates a comparison with the theoretical models and provides a physical interpretation of the features enabling it to make improved predictions.

42 ENGINEERING↗