Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Convolution”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

Identification of Grand-design and Flocculent spirals from SDSS using deep convolutional neural network

ABSTRACT Spiral galaxies can be classified into the Grand-designs and Flocculents based on the nature of their spiral arms. The Grand-designs exhibit almost continuous and high contrast spiral arms and are believed to be driven by stationary density waves, while the Flocculents have patchy and low-contrast spiral features and are primarily stochastic in origin. We train a deep convolutional neural network model to classify spirals into Grand-designs and Flocculents, with a testing accuracy of $\mathrm{97.2{{\ \rm per\ cent}}}$. We then use the above model for classifying 1354 spirals from the SDSS. Out of these, 721 were identified as Flocculents, and the rest as Grand-designs. Interestingly, we find the mean asymptotic rotational velocities of our newly classified Grand-designs and Flocculents are 218 ± 86 and 146 ± 67 km s−1, respectively, indicating that the Grand-designs are mostly the high-mass and the Flocculents the intermediate-mass spirals. This is further corroborated by the observation that the mean morphological indices of the Grand-designs and Flocculents are 2.6 ± 1.8 and 4.7 ± 1.9, respectively, implying that the Flocculents primarily consist of a late-type galaxy population in contrast to the Grand-designs. Finally, an almost equal fraction of bars ∼0.3 in both the classes of spiral galaxies reveals that the presence of a bar component does not regulate the type of spiral arm hosted by a galaxy. Our results may have important implications for formation and evolution of spiral arms in galaxies.

Astronomy & Astrophysics↗

Lessons learned from the two largest Galaxy morphological classification catalogues built by convolutional neural networks

ABSTRACT We compare the two largest galaxy morphology catalogues, which separate early- and late-type galaxies at intermediate redshift. The two catalogues were built by applying supervised deep learning (convolutional neural networks, CNNs) to the Dark Energy Survey data down to a magnitude limit of ∼21 mag. The methodologies used for the construction of the catalogues include differences such as the cutout sizes, the labels used for training, and the input to the CNN – monochromatic images versus gri-band normalized images. In addition, one catalogue is trained using bright galaxies observed with DES (i < 18), while the other is trained with bright galaxies (r < 17.5) and ‘emulated’ galaxies up to r-band magnitude 22.5. Despite the different approaches, the agreement between the two catalogues is excellent up to i < 19, demonstrating that CNN predictions are reliable for samples at least one magnitude fainter than the training sample limit. It also shows that morphological classifications based on monochromatic images are comparable to those based on gri-band images, at least in the bright regime. At fainter magnitudes, i > 19, the overall agreement is good (∼95 per cent), but is mostly driven by the large spiral fraction in the two catalogues. In contrast, the agreement within the elliptical population is not as good, especially at faint magnitudes. By studying the mismatched cases, we are able to identify lenticular galaxies (at least up to i < 19), which are difficult to distinguish using standard classification approaches. The synergy of both catalogues provides an unique opportunity to select a population of unusual galaxies.

79 ASTRONOMY AND ASTROPHYSICS↗

Effective cosmic density field reconstruction with convolutional neural network

ABSTRACT We present a cosmic density field reconstruction method that augments the traditional reconstruction algorithms with a convolutional neural network (CNN). Following previous work, the key component of our method is to use the reconstructed density field as the input to the neural network. We extend this previous work by exploring how the performance of these reconstruction ideas depends on the input reconstruction algorithm, the reconstruction parameters, and the shot noise of the density field, as well as the robustness of the method. We build an eight-layer CNN and train the network with reconstructed density fields computed from the Quijote suite of simulations. The reconstructed density fields are generated by both the standard algorithm and a new iterative algorithm. In real space at z = 0, we find that the reconstructed field is 90 per cent correlated with the true initial density out to $k\sim 0.5 \, \mathrm{ h}\, \rm {Mpc}^{-1}$, a significant improvement over $k\sim 0.2 \, \mathrm{ h}\, \rm {Mpc}^{-1}$ achieved by the input reconstruction algorithms. We find similar improvements in redshift space, including an improved removal of redshift space distortions at small scales. We also find that the method is robust across changes in cosmology. Additionally, the CNN removes much of the variance from the choice of different reconstruction algorithms and reconstruction parameters. However, the effectiveness decreases with increasing shot noise, suggesting that such an approach is best suited to high density samples. This work highlights the additional information in the density field beyond linear scales as well as the power of complementing traditional analysis approaches with machine learning techniques.

Astronomy & Astrophysics↗

Reconstructing cosmological initial conditions from late-time structure with convolutional neural networks

ABSTRACT We present a method to reconstruct the initial linear-regime matter density field from the late-time non-linearly evolved density field in which we channel the output of standard first-order reconstruction to a convolutional neural network (CNN). Our method shows dramatic improvement over the reconstruction of either component alone. We show why CNNs are not well-suited for reconstructing the initial density directly from the late-time density: CNNs are local models, but the relationship between initial and late-time density is not local. Our method leverages standard reconstruction as a preprocessing step, which inverts bulk gravitational flows sourced over very large scales, transforming the residual reconstruction problem from long-range to local and making it ideally suited for a CNN. We develop additional techniques to account for redshift distortions, which warp the density fields measured by galaxy surveys. Our method improves the range of scales of high-fidelity reconstruction by a factor of 2 in wavenumber above standard reconstruction, corresponding to a factor of 8 increase in the number of well-reconstructed modes. In addition, our method almost completely eliminates the anisotropy caused by redshift distortions. As galaxy surveys continue to map the Universe in increasingly greater detail, our results demonstrate the opportunity offered by CNNs to untangle the non-linear clustering at intermediate scales more accurately than ever before.

Astronomy & Astrophysics↗

Survey of gravitationally lensed objects in HSC imaging (SuGOHI) – X. Strong lens finding in the HSC-SSP using convolutional neural networks

ABSTRACT We apply a novel model based on convolutional neural networks (CNN) to identify gravitationally lensed galaxies in multiband imaging of the Hyper Suprime-Cam Subaru Strategic Program (HSC-SSP) Survey. The trained model is applied to a parent sample of 2350 061 galaxies selected from the $\sim$ 800 deg$^2$ Wide area of the HSC-SSP Public Data Release 2. The galaxies in HSC Wide are selected based on stringent pre-selection criteria, such as multiband magnitudes, stellar mass, star formation rate, extendedness limit, photometric redshift range, etc. The trained CNN assigns a score from 0 to 1, with 1 representing lenses and 0 representing non-lenses. Initially, the CNN selects a total of 20 241 cutouts with a score greater than 0.9, but this number is subsequently reduced to 1522 cutouts after removing definite non-lenses for further visual inspection. We discover 43 grade A (definite) and 269 grade B (probable) strong lens candidates, of which 97 are completely new. In addition, we also discover 880 grade C (possible) lens candidates, 289 of which are known systems in the literature. We identify 143 candidates from the known systems of grade C that had higher confidence in previous searches. Our model can also recover 285 candidate galaxy-scale lenses from the Survey of Gravitationally lensed Objects in HSC Imaging (SuGOHI), where a single foreground galaxy acts as the deflector. Even though group-scale and cluster-scale lens systems are not included in the training, a sample of 32 SuGOHI-c (i.e. group/cluster-scale systems) lens candidates is retrieved. Our discoveries will be useful for ongoing and planned spectroscopic surveys, such as the Subaru Prime Focus Spectrograph project, to measure lens and source redshifts in order to enable detailed lens modelling.

Jaelani, Anton T. (ORCID:0000000162825778)↗

Deciphering enhancer sequence using thermodynamics-based models and convolutional neural networks

Abstract Deciphering the sequence-function relationship encoded in enhancers holds the key to interpreting non-coding variants and understanding mechanisms of transcriptomic variation. Several quantitative models exist for predicting enhancer function and underlying mechanisms; however, there has been no systematic comparison of these models characterizing their relative strengths and shortcomings. Here, we interrogated a rich data set of neuroectodermal enhancers in Drosophila, representing cis- and trans- sources of expression variation, with a suite of biophysical and machine learning models. We performed rigorous comparisons of thermodynamics-based models implementing different mechanisms of activation, repression and cooperativity. Moreover, we developed a convolutional neural network (CNN) model, called CoNSEPT, that learns enhancer ‘grammar’ in an unbiased manner. CoNSEPT is the first general-purpose CNN tool for predicting enhancer function in varying conditions, such as different cell types and experimental conditions, and we show that such complex models can suggest interpretable mechanisms. We found model-based evidence for mechanisms previously established for the studied system, including cooperative activation and short-range repression. The data also favored one hypothesized activation mechanism over another and suggested an intriguing role for a direct, distance-independent repression mechanism. Our modeling shows that while fundamentally different models can yield similar fits to data, they vary in their utility for mechanistic inference. CoNSEPT is freely available at: https://github.com/PayamDiba/CoNSEPT.

59 BASIC BIOLOGICAL SCIENCES↗

Using convolutional neural networks to accelerate three-dimensional coherent synchrotron radiation computations

Calculating the effects of coherent synchrotron radiation (CSR) is one of the most computationally expensive tasks in accelerator physics. Here, we use convolutional neural networks (CNNs), along with a latent conditional diffusion (LCD) model, trained on physics-based simulations to speed up calculations. Specifically, we produce the 3D CSR wakefields generated by electron bunches in circular orbit in the steady-state condition. Two datasets are used for training and testing the models: wakefields generated by three-dimensional Gaussian electron distributions and wakefields from a sum of up to 25 three-dimensional Gaussian distributions. The CNNs are able to accurately produce the 3D wakefields ∼250–1000 times faster than the numerical calculations, while the LCD achieves a gain of a factor of ∼34. We also test the extrapolation and out-of-distribution generalization ability of the models. They generalize well on distributions with larger spreads than what they were trained on but struggle with smaller spreads.

43 PARTICLE ACCELERATORS↗

Neutrino interaction classification with a convolutional neural network in the DUNE far detector

The Deep Underground Neutrino Experiment is a next-generation neutrino oscillation experiment that aims to measure $CP$-violation in the neutrino sector as part of a wider physics program. A deep learning approach based on a convolutional neural network has been developed to provide highly efficient and pure selections of electron neutrino and muon neutrino charged-current interactions. The electron neutrino (antineutrino) selection efficiency peaks at 90% (94%) and exceeds 85% (90%) for reconstructed neutrino energies between 2-5 GeV. The muon neutrino (antineutrino) event selection is found to have a maximum efficiency of 96% (97%) and exceeds 90% (95%) efficiency for reconstructed neutrino energies above 2 GeV. When considering all electron neutrino and antineutrino interactions as signal, a selection purity of 90% is achieved. These event selections are critical to maximize the sensitivity of the experiment to $CP$-violating effects.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Measurement of Atmospheric Neutrino Oscillation Parameters Using Convolutional Neural Networks with 9.3 Years of Data in IceCube DeepCore

The DeepCore subdetector of the IceCube Neutrino Observatory provides access to neutrinos with energies above approximately 5 GeV. Data taken between 2012 and 2021 (3387 days) are utilized for an atmospheric ν μ disappearance analysis that studied 150 257 neutrino-candidate events with reconstructed energies between 5 and 100 GeV. An advanced reconstruction based on a convolutional neural network is applied, providing increased signal efficiency and background suppression, resulting in a measurement with both significantly increased statistics compared to previous DeepCore oscillation results and high neutrino purity. For the normal neutrino mass ordering, the atmospheric neutrino oscillation parameters and their 1 σ errors are measured to be Δ m 32 2 = 2.40 − 0.04 + 0.05 × 10 − 3 eV 2 and sin 2 θ 23 = 0.54 − 0.03 + 0.04 . The results are the most precise to date using atmospheric neutrinos, and are compatible with measurements from other neutrino detectors including long-baseline accelerator experiments. Published by the American Physical Society 2025

Abbasi, R.↗

Identification and denoising of radio signals from cosmic-ray air showers using convolutional neural networks

Radio pulses generated by cosmic-ray air showers can be used to reconstruct key properties like the energy and depth of the electromagnetic component of cosmic-ray air showers. Radio detection threshold, influenced by natural and anthropogenic radio background, can be reduced through various techniques. In this work, we demonstrate that convolutional neural networks (CNNs) are an effective way to lower the threshold. We developed two CNNs: a classifier to distinguish radio signal waveforms from background noise and a denoiser to clean contaminated radio signals. Following the training and testing phases, we applied the networks to air-shower data triggered by scintillation detectors of the prototype station for the enhancement of IceTop, IceCube’s surface array at the South Pole. Over a four-month period, we identified 554 cosmic-ray events in coincidence with IceTop, approximately five times more compared to a reference method based on a cut on the signal-to-noise ratio. Comparisons with IceTop measurements of the same air showers confirmed that the CNNs reliably identified cosmic-ray radio pulses and outperformed the reference method. Additionally, we find that CNNs reduce the false-positive rate of air-shower candidates and effectively denoise radio waveforms, thereby improving the accuracy of the power and arrival time reconstruction of radio pulses.

Abbasi, R↗

Fast Gaussian Process Estimation for Large-Scale In Situ Inference using Convolutional Neural Networks

Exascale computing will bring with it significant I/O limitations. One foreseeable consequence of such restrictions is that the user can save only a small fraction of complex simulation data to disk for subsequent analysis. An alternative is to fit statistical models to data in situ, that is, inside the simulation as it runs. This option requires extremely fast statistical estimation to avoid slowing down the simulation. Gaussian processes (GPs) have state-of-the-art predictive performance for modeling spatial data. However, standard estimation methods for GPs scale quite poorly to large data sets as parameter estimation requires inverting a covariance matrix to the size of the data set. In the presented work, we use a convolutional neural network (CNN) to predict the GP parameters for a spatial data set, from a simulation or otherwise, rather than optimize the parameters directly. Here, our presented case study models spatial data from E3SM, the Department of Energy’s Exascale climate model. The CNN is trained on synthetic data simulated from GP models with known parameters and then applied to data from the climate simulation. In the presented examples, the neural network scheme produces parameter estimates that compare well with standard methods such as maximum likelihood estimation in predictive performance but is obtained four orders of magnitude faster.

big data↗

Hybrid Attack Graph Generation with Graph Convolutional Deep-Q Learning

Critical infrastructures such as power grids have become increasingly complex, connected, and vulnerable to adverse scenarios, including cyber and physical attacks and faults. Effective risk mitigation for such cyber-physical energy systems (CPES), requires preemptive knowledge of likely adversarial attack scenarios. Hybrid Attack Graph (HAG) is a structured way to represent an adversarial scenario as an attack sequence using a threat model. However, the scarcity of documented attack sequences hinders analysts and CPES planners’ ability to identify credible attack scenarios for a given CPES. We propose a data-driven Graph Convolutional Deep-Q Network (GCDQ) to address this data challenge through generating HAGs. By leveraging limited real-world observations from the MITRE ATT&CK knowledge base, our GCDQ model synthesizes realistic graphs with the targeted attribute of minimum detectability via reinforcement learning. This generative model is the first step in creating a tool to substantially boost the attack sequence dataset and enhance the performance of CPS defense-related tasks by providing insights into likely attack sequences with given attributes.

deep learning, artificial intelligence↗

GCoD: Graph Convolutional Network Acceleration via Dedicated Algorithm and Accelerator Co-Design

Graph Convolutional Networks (GCNs) have emerged as the state-of-the-art graph learning model. However, it remains notoriously challenging to inference GCNs over large graph datasets, limiting their application to large real-world graphs and hindering the exploration of deeper and more sophisticated GCN graphs. This is because real-world graphs can be extremely large and sparse. Furthermore, the node degree of GCNs tends to follow the power-law distribution and therefore have highly irregular adjacency matrices, resulting in prohibitive inefficiencies in both data processing and movement and thus substantially limiting the achievable GCN acceleration efficiency. To this end, this paper proposes the first GCN algorithm and accelerator Co-Design framework dubbed GCoD which can largely alleviate the aforementioned GCN irregularity and boost GCNs' inference efficiency. Specifically, on the algorithm level, GCoD integrates a divide and conquer GCN training strategy that polarizes the graphs to be either denser or sparser in local neighborhoods without compromising the model accuracy, resulting in graph adjacency matrices that (mostly) have merely two levels of workload and enjoys largely enhanced regularity and thus ease of acceleration. On the hardware level, we further develop a dedicated two-pronged accelerator with a separated engine to process each of the aforementioned workloads, further boosting the overall utilization and acceleration efficiency. Extensive experiments and ablation studies validate that our GCoD consistently outperforms state-of-the-art designs in terms of accelerator efficiency while maintaining or even improving the task accuracy. Additionally, we visualize GCoD trained graph adjacency matrices to better understand its advantages. All codes and pre-trained models will be released upon acceptance.

You, Haoran↗

Efficient Data Compression for 3D Sparse TPC via Bicephalous Convolutional Autoencoder

Real-time data collection and analysis in large experimental facilities present a great challenge across multiple domains, including high energy physics, nuclear physics, and cosmology. To address this, machine learning (ML)-based methods for real-time data compression have drawn significant attention. However, unlike natural image data, such as CIFAR and ImageNet that are relatively small-sized and continuous, scientific data often come in as three-dimensional 3D data volumes at high rates with high sparsity (many zeros) and non-Gaussian value distribution. This makes direct application of popular ML compression methods, as well as conventional data compression methods, suboptimal. To address these obstacles, this work introduces a dual-head autoencoder to resolve sparsity and regression simultaneously, called Bicephalous Convolutional AutoEncoder (BCAE). This method shows advantages both in compression fidelity and ratio compared to traditional data compression methods, such as MGARD, SZ, and ZFP. To achieve similar fidelity, the best performer among the traditional methods can reach only half the compression ratio of BCAE. Moreover, a thorough ablation study of the BCAE method shows that a dedicated segmentation decoder improves the reconstruction.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Convolutional Variational Autoencoder-based Unsupervised Learning for Power Systems Faults

Classification of power system event data is a growing need, particularly where non-protective relaying-based sensors are used to monitor grid performance. Given the high burden of obtaining event data with appropriate labeling, an unsupervised approach is highly valuable. This approach enables using event data without labeling, which is far easier to obtain. This paper presents an unsupervised learning method to classify and label transients observed in the distribution grid. A Convolutional Variational Autoencoder (CVAE) was developed for this purpose. We demonstrate the efficacy of our approach using the transient data generated from the simulations. The simulation data is used to train the CVAE that identifies different faults as different clusters in the latent space. The clusters are then used as the foundation model to categorize the real-world data.

Alam, Maksudul↗

A 3D Implementation of Convolutional Neural Network for Fast Inference

Low latency inference has many applications in edge machine learning. In this paper, we present a run-time configurable convolutional neural network (CNN) inference ASIC design for low-latency edge machine learning. By implementing a 5-stage pipelined CNN inference model in a 3D ASIC technology, we demonstrate that the model distributed on two dies utilizing face-to-face (F2F) 3D integration achieves superior performance. Our experimental results show that the design based on 3D integration achieves 43% better energy-delay product when compared to the traditional 2D technology.

Miniskar, Narasinga Rao↗

Temporal Convolutional Network Using Empirical Mode Decomposition to Detect Faults in Grid Connected Systems

Grid-connected power electronic systems require timely and reliable fault detection to prevent equipment damage and reduce downtime. This paper presents a forecasting-based anomaly detection pipeline that decomposes voltage and current measurements into intrinsic mode functions (IMFs) using empirical mode decomposition (EMD), then trains a causal temporal convolutional network (TCN) on normal-operation IMF data to predict short-horizon future dynamics. Deviations between forecasts and observations are summarized as reliability-weighted residual scores and thresholded per sensor using robust statistics with temporal persistence constraints to suppress false positives. To reduce runtime, EMD is performed on downsampled signals for detection, while raw-rate EMD is applied only within a short region of interest for high-frequency interpretability near detected events. Results on a simulated grid-connected converter system demonstrate that IMF-domain forecasting improves anomaly separability relative to raw-signal forecasting and provides interpretable evidence of faults across decomposition channels.

Sutton, Elizabeth [ORNL] (ORCID:0009000078885935)↗

Defect Recognition for Eddy Current Testing of Spent Nuclear Fuel Canister using Convolutional Neural Network

This paper proposes an accurate and robust defect detection solution for 304L and 306L stainless steel (SS) weld. In the proposed solution, Eddy current testing (ECT) is employed to generate 2-dimensional (2D) data for samples under test with defects. The 2D data can be treated as images for deep learning-based defect detection. Since convolutional neural networks (CNNs) are powerful in processing images, CNN is employed in this study for defect detection. Experiments are conducted on a submerged arc welding (SAW) 304L SS weld sample with an artificial crack generated by waterjet cutting. The ECT data on this seeded fault sample is utilized to verify the proposed solution. For this purpose, the ECT measurement are separated as from Fault area and Normal area, which are used for CNN training. After training, the testing data is used for verification. Experimental results demonstrate the feasibility and effectiveness of the proposed solution.

Niu, Guangxing↗