Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Convolution”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 649 records · Page 36

Protograph-Based Raptor-Like Codes

Theoretical analysis has long indicated that feedback improves the error exponent but not the capacity of pointto- point memoryless channels. The analytic and empirical results indicate that at short blocklength regime, practical rate-compatible punctured convolutional (RCPC) codes achieve low latency with the use of noiseless feedback. In 3GPP, standard rate-compatible turbo codes (RCPT) did not outperform the convolutional codes in the short blocklength regime. The reason is the convolutional codes for low number of states can be decoded optimally using Viterbi decoder. Despite excellent performance of convolutional codes at very short blocklengths, the strength of convolutional codes does not scale with the blocklength for a fixed number of states in its trellis.

Divsalar, Dariush↗

A Bayesian Deep Learning Approach to Near-Term Climate Prediction

Since model bias and associated initialization shock are serious shortcomings that reduce prediction skills in state-of-the-art decadal climate prediction efforts, we pursue a complementary machine-learning-based approach to climate prediction. The example problem setting we consider consists of predicting natural variability of the North Atlantic sea surface temperature on the interannual timescale in the pre-industrial control simulation of the Community Earth System Model. While previous works have considered the use of recurrent networks such as convolutional LSTMs and reservoir computing networks in this and other similar problem settings, we currently focus on the use of feedforward convolutional networks. In particular, we find that a feedforward convolutional network with a Densenet architecture is able to outperform a convolutional LSTM in terms of predictive skill. Next, we go on to consider a probabilistic formulation of the same network based on Stein variational gradient descent and find that in addition to providing useful measures of predictive uncertainty, the probabilistic (Bayesian) version improves on its deterministic counterpart in terms of predictive skill. Finally, we characterize the reliability of the ensemble of machine learning models obtained in the probabilistic setting by using analysis tools developed in the context of ensemble numerical weather prediction.

54 ENVIRONMENTAL SCIENCES↗

A Modified Sequence-to-point HVAC Load Disaggregation Algorithm

This paper presents a modified sequence-to-point (S2P) algorithm for disaggregating the heat, ventilation, and air conditioning (HVAC) load from the total building electricity consumption. The original S2P model is convolutional neural network (CNN) based, which uses load profiles as inputs. We propose three modifications. First, the input convolution layer is changed from 1D to 2D so that normalized temperature profiles are also used inputs to the S2P model. Second, a drop-out layer is added to improve adaptability and generalizability so that the model trained in one area can be transferred to other geographical areas without labelled HVAC data. Third, a fine-tuning process is proposed for areas with a small amount of labelled HVAC data so that the pre-trained S2P model can be fine-tuned to achieve higher disaggregation accuracy (i.e., better transferability) in other areas. The model is first trained and tested using smart meter and sub-metered HVAC data collected in Austin, Texas. Then, the trained model is tested on two other areas: Boulder, Colorado and San Diego, California. Simulation results show that the proposed modified S2P algorithm outperforms the original S2P model and the support-vector machine based approach in accuracy, adaptability, and transferability.

Ye, Kai↗

Adaptive methods of generating complex light arrays

Structured light arrays of various shapes have been a cornerstone in optical science, driven by the complexities of precise and adaptable generation. This study introduces an approach using a spatial light modulator (SLM) as a generator for these arrays. By projecting a holographic mask onto the SLM, it functions simultaneously as an optical convolution device, focusing mechanism, and structured light beam mask. Our approach offers unmatched versatility, allowing for the experimental fabrication of traditional beam arrays like azimuthal Laguerre–Gaussian (LG), Bessel–Gaussian (BG), and Hermite–Gauss (HG) in the far-field. Notably, it has enabled a method of generating Ince–Gauss (IG) and LG radial mode beam arrays using a convolution solution. Our system provides exceptional control over array periodicity and intensity distribution, bypassing the Talbot self-imaging phenomenon seen in traditional setups. We provide an in-depth theoretical discussion, supported by empirical evidence, of our far-field results. This method has vast potential for applications in optical communication, data processing, and multi-particle manipulation. It paves the way for rapid generation of structured light with high spatial frequencies and complex shapes, promising transformative advances in these domains.

Optics↗

Resampling study

The author has identified the following significant results. The nearest neighbor and cubic convolution resampling algorithms were applied to a variety of images extracted from LANDSAT MSS data. A comparison of the results demonstrated that (1) cubic convolution can cause spreading of small features and can introduce noticeable overshoot (ringing) into the data; (2) cubic convolution attenuates the high spatial frequencies compared to the original and nearest neighbor resampled data; and (3) cubic convolution generally produces photographic products of superior visual quality. The effects of the resampling algorithms on multispectral classification were not conclusively determined due to the small number of images tested.

Ferneyhough, D. G.↗

Fast-Polynomial-Transform Program

Computer program uses fast-polynomial-transformation (FPT) algorithm applicable to two-dimensional mathematical convolutions. Two-dimensional cyclic convolutions converted to one-dimensional convolutions in polynomial rings. Program decomposes cyclic polynomials into polynomial convolutions of same length. Only FPT's and fast Fourier transforms of same length required. Modular approach saves computional resources. Program written in C.

Truong, T. K.↗

Error control techniques for satellite and space communications

Worked performed during the reporting period is summarized. Construction of robustly good trellis codes for use with sequential decoding was developed. The robustly good trellis codes provide a much better trade off between free distance and distance profile. The unequal error protection capabilities of convolutional codes was studied. The problem of finding good large constraint length, low rate convolutional codes for deep space applications is investigated. A formula for computing the free distance of 1/n convolutional codes was discovered. Double memory (DM) codes, codes with two memory units per unit bit position, were studied; a search for optimal DM codes is being conducted. An algorithm for constructing convolutional codes from a given quasi-cyclic code was developed. Papers based on the above work are included in the appendix.

Costello, Daniel J., Jr.↗

Trellises and Trellis-Based Decoding Algorithms for Linear Block Codes

Decoding algorithms based on the trellis representation of a code (block or convolutional) drastically reduce decoding complexity. The best known and most commonly used trellis-based decoding algorithm is the Viterbi algorithm. It is a maximum likelihood decoding algorithm. Convolutional codes with the Viterbi decoding have been widely used for error control in digital communications over the last two decades. This chapter is concerned with the application of the Viterbi decoding algorithm to linear block codes. First, the Viterbi algorithm is presented. Then, optimum sectionalization of a trellis to minimize the computational complexity of a Viterbi decoder is discussed and an algorithm is presented. Some design issues for IC (integrated circuit) implementation of a Viterbi decoder are considered and discussed. Finally, a new decoding algorithm based on the principle of compare-select-add is presented. This new algorithm can be applied to both block and convolutional codes and is more efficient than the conventional Viterbi algorithm based on the add-compare-select principle. This algorithm is particularly efficient for rate 1/n antipodal convolutional codes and their high-rate punctured codes. It reduces computational complexity by one-third compared with the Viterbi algorithm.

Lin, Shu↗

ExtremeMETA: High-speed Lightweight Image Segmentation Model by Remodeling Multi-channel Metamaterial Imagers

Deep neural networks (DNNs) have heavily relied on traditional computational units, such as CPUs and GPUs. However, this conventional approach brings significant computational burden, latency issues, and high power consumption, limiting their effectiveness. This has sparked the need for lightweight networks such as ExtremeC3Net. Meanwhile, there have been notable advancements in optical computational units, particularly with metamaterials, offering the exciting prospect of energy-efficient neural networks operating at the speed of light. Yet, the digital design of metamaterial neural networks (MNNs) faces precision, noise, and bandwidth challenges, limiting their application to intuitive tasks and low-resolution images. In this study, we proposed a large kernel lightweight segmentation model, ExtremeMETA. Based on ExtremeC3Net, our proposed model, ExtremeMETA maximized the ability of the first convolution layer by exploring a larger convolution kernel and multiple processing paths. With the large kernel convolution model, we extended the optic neural network application boundary to the segmentation task. To further lighten the computation burden of the digital processing part, a set of model compression methods was applied to improve model efficiency in the inference stage. The experimental results on three publicly available datasets demonstrated that the optimized efficient design improved segmentation performance from 92.45 to 95.97 on mIoU while reducing computational FLOPs from 461.07 MMacs to 166.03 MMacs. The large kernel lightweight model ExtremeMETA showcased the hybrid design’s ability on complex tasks.

large convolution kernel↗

Performance of A Real-Time Photon Counting Optical Receiver in the Presence of Emulated Channel Fading

Free-space optical communication links with terrestrial ground stations experience fading due to atmospheric scintillation and beam pointing. Fiber-coupled receiver systems experience additional fading at the interface between the fiber and free-space optics of the telescope. The National Aeronautics and Space Administration (NASA) Glenn Research Center (GRC) has characterized a real-time photon-counting optical ground receiver system with an atmospheric fade emulation system. The receiver system is comprised of a fiber interconnect, an array of superconducting nanowire single photon detectors (SNSPDs), and a field programmable gate array (FPGA) based receive modem. Two fiber interconnect/detector architectures have been studied. One architecture uses a 70-mode photonic lantern coupled to seven single pixel SNSPDs. The other architecture uses a 10-mode few-mode fiber (FMF) coupled to a 15-pixel SNSPD array. The receiver system complies with the Consultative Committee for Space Data Systems (CCSDS) Optical Communications High Photon Efficiency Coding and Synchronization Standard, which uses serially concatenated convolutionally coded pulse-position modulation (SCPPM). The CCSDS standard is designed for use in low photon flux missions, including the Orion Artemis-II Optical (O2O) communications demonstration. The standard utilizes a convolutional symbol interleaver which can be resized to mitigate different fades. The fade emulation system employed in this work emulates scintillation-induced, pointing-induced, and coupling-induced fading. This paper gives an overview of the real-time optical receiver system and the fade emulation system. It presents tests results which show the impact of fading on the performance on the receiver. The test results show that in the presence of channel fading, the 70-mode photonic lantern outperforms the 10-mode FMF under higher (D/r_0=9) turbulence conditions due to high fiber-coupling-induced fading and fiber coupling loss on the 10-mode FMF. When operating in lower turbulence (D/r_0=4), the 10-mode FMF outperforms the 70-mode photonic lantern. The paper also shows a larger convolutional interleaver improves the system performance as long as the receiver does not lose acquisition.

optical communications↗

Performance of a real-time photon counting optical receiver in the presence of emulated channel fading

Free-space optical communication links with terrestrial ground stations experience fading due to atmospheric scintillation and beam pointing. Fiber-coupled receiver systems experience additional fading at the interface between the fiber and free-space optics of the telescope. The National Aeronautics and Space Administration (NASA) Glenn Research Center (GRC) has characterized a real-time photon-counting optical ground receiver system with an atmospheric fade emulation system. The receiver system is comprised of a fiber interconnect, an array of superconducting nanowire single photon detectors (SNSPDs), and a field programmable gate array (FPGA) based receive modem. Two fiber interconnect/detector architectures have been studied. One architecture uses a 70-mode photonic lantern coupled to seven single pixel SNSPDs. The other architecture uses a 10-mode few-mode fiber (FMF) coupled to a 15-pixel SNSPD array. The receiver system complies with the Consultative Committee for Space Data Systems (CCSDS) Optical Communications High Photon Efficiency Coding and Synchronization Standard, which uses serially concatenated convolutionally coded pulse-position modulation (SCPPM). The CCSDS standard is designed for use in low photon flux missions, including the Orion Artemis-II Optical (O2O) communications demonstration. The standard utilizes a convolutional symbol interleaver which can be resized to mitigate different fades. The fade emulation system employed in this work emulates scintillation-induced, pointing-induced, and coupling-induced fading. This paper gives an overview of the real-time optical receiver system and the fade emulation system. It presents tests results which show the impact of fading on the performance on the receiver. The test results show that in the presence of channel fading, the 70-mode photonic lantern outperforms the 10-mode FMF under higher (D/r_0=9) turbulence conditions due to high fiber-coupling-induced fading and fiber coupling loss on the 10-mode FMF. When operating in lower turbulence (D/r_0=4), the 10-mode FMF outperforms the 70-mode photonic lantern. The paper also shows a larger convolutional interleaver improves the system performance as long as the receiver does not lose acquisition.

optical communications↗

ESM data downscaling: a comparison of super-resolution deep learning models

Abstract Climate projections at fine spatial resolutions are required to conduct accurate risk assessment for critical infrastructure and design adaptation planning. Generating these projections using advanced Earth system models (ESM) requires significant computational resources. To address this issue, various statistical downscaling techniques have been introduced to generate fine-resolution data from coarse-resolution simulations. In this study, we evaluate and compare five deep learning-based downscaling techniques, namely, super-resolution convolutional neural networks, fast super-resolution convolutional neural network ESM, efficient sub-pixel convolutional neural network, enhanced deep residual network (EDRN), and super-resolution generative adversarial network (SRGAN). These techniques are applied to a dataset generated by the Energy Exascale Earth System Model (E3SM), focusing on key surface variables such as surface temperature, shortwave heat flux, and longwave heat flux. Models are trained and validated using paired fine-resolution (0.25 $$^{\circ }$$ ∘ ) and coarse-resolution (1 $$^{\circ }$$ ∘ ) monthly data obtained from a 9-year simulation. Next, blind testing is performed using monthly data obtained from two different years outside of the training and validation set. To evaluate the efficiency of each technique, different statistical metrics are used, including mean squared error (MSE), peak signal-to-noise ratio (PSNR), structural similarity index measure (SSIM), and learned perceptual image patch similarity (LPIPS). The results show that EDRN outperforms other algorithms in terms of PSNR, SSIM, and MSE, but struggles to capture fine-scale features in the data. In contrast, SRGAN, a generative model that uses perceptual loss, excels in capturing fine details at boundaries and internal structures, resulting in lower LPIPS than other methods.

Pawar, Nikhil M. (ORCID:0000000211613289)↗

Variable rate neural compression for sparse detector data

Particle colliders produce data at extraordinary rates, posing major challenges for transmission and storage. High-throughput compression algorithms are therefore essential. In the sPHENIX experiment taking data at the Relativistic Heavy Ion Collider, a time projection chamber records three-dimensional (3D) particle trajectories that are highly sparse, making conventional learning-free lossy compression ineffective. Convolutional neural networks have surpassed traditional methods in compression ratio and accuracy. However, they fail to exploit sparsity for efficiency. To address these gaps, we present BCAE-VS, a bicephalous convolutional autoencoder with variable compression ratio for sparse data, which adapts compression to input complexity through key-point identification and sparse convolution. BCAE-VS achieves higher accuracy and compression ratios than prior neural approaches while being orders of magnitude smaller. Moreover, its throughput increases with sparsity—a property not observed in other methods. Although it was developed for collider experiments, BCAE-VS readily extends to other sparse data domains, such as light detection and ranging (LiDAR) sensing and 3D microscopy.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Accelerating magnonic simulations with the pseudospectral Landau-Lifshitz equation

The pseudospectral Landau-Lifshitz (PS-LL) model can describe atomic-scale magnetic exchange interactions within a continuum framework. This is achieved by employing a convolution kernel that models the nonlocal interaction in a grid-independent manner. Even though the PS-LL was originally introduced to address atomic exchange, any nonlocal kernel can be modeled. In the field of magnonics, the dipole field is fundamental to describe the dispersion relation of magnons, the quasiparticle representation of angular momentum. Because dipole-dipole interactions are long-range, numerical approaches typically rely on convolutions. Here, we demonstrate that the PS-LL model can be used to perform magnonic simulations with a single convolution kernel derived from analytical solutions. We demonstrate a twofold increase in computational speed compared with the full dipole calculation. This approach is valid insofar as the excitations are linear, which is typically the case for magnons. Our results have the potential to accelerate magnonic research, particularly for the inverse design method, where several simulations must be performed to achieve the desired outcome.

Mathematics and computing↗

Modeled sensitivity of multi-MA accelerator performance to electrode contaminant inventory

Significant particle-in-cell code development has enabled simulations of power flow in multi-MA accelerators to include the desorption of surface contaminants, their ionization into surface plasmas, and the impact of these plasmas on efficiency. The simulations base desorption on an Arrhenius equation, whose most significant unknown is the surface contaminant inventory. The sensitivity of power-flow simulations to this inventory is studied here using Sandia National Laboratories' Z accelerator with a 7-nH MagLIF load [Phys. Plasmas 17, 056303 (2010)]. Simulations are conducted in 3D cylindrical coordinates for the current-adder, or “convolute,” region of Z and in 2D for the final feed only. Simulated contaminant inventories are varied from 1 to 32 monolayers (MLs) in 2D, and 2 to 4 ML in 3D. The results reveal sensitivities to the local ratio of E/B⁠. The high B-field, low E-field region near the short-circuit load is insensitive to the contaminant inventory, where assumed values of 4–32 ML change the load current by ≤ 2%, and agree with experiment to within 2% at peak current. A 1-ML value is the outlier, increasing the load current by 5%, but still within measurement uncertainty. In contrast, the relatively higher E-field, lower B-field convolute region has slower contaminant desorption and higher-magnitude E-field penetration of the surface plasmas. The current loss in the convolute region does increase with contaminant inventory. The loss assuming 4 ML is 12% larger than for 2 ML, with 4 ML being the better match to experiment.

Arrhenius equation↗

Hierarchical-embedding autoencoder with a predictor as efficient architecture for learning time-evolution in multi-scale turbulent flows

We introduce a scale-aware, data-driven deep learning modeling framework for accurately predicting the time evolution of multi-scale turbulent plasma and liquid flows. The approach is motivated by the idea of scale separation. Structures of vastly different length scales emerge in these systems, and interactions between these structures occur only locally. To exploit this structure, the flow state is transformed by a hierarchical, fully convolutional autoencoder, not into a single embedding layer as in conventional convolutional surrogate models, but into a series of embedding layers. A stepwise training strategy ensures that fine-scale features are encoded on a high-resolution grid, while larger structures are represented on progressively coarser layers. The time evolution predictor advances all embedding layers in sync, capturing local interactions between features at the same scale as well as between all scales. This approach enables efficient modeling of multi-scale systems since negligible interactions between distant, small-scale structures do not need to be directly modeled. Our hierarchical-embedding autoencoder with a predictor framework is evaluated on canonical examples of multi-scale turbulence: two-dimensional Kolmogorov flow and Hasegawa–Wakatani plasma turbulence. In both cases, the proposed framework significantly improves predictive accuracy relative to conventional convolutional network architectures. A significant improvement in prediction accuracy was observed for crucial statistical characteristics of the Hasegawa–Wakatani plasma as well as for individual trajectories of the Kolmogorov flow turbulence. Importantly, the model's rollout for the Hasegawa–Wakatani problem demonstrates a four-order-of-magnitude speedup compared to traditional numerical solvers.

Khrabry, Alexander I. [Princeton Univ., NJ (United↗

Orthogonal Polynomials Defined by Self-Similar Measures with Overlaps

Here, we study orthogonal polynomials with respect to self-similar measures, focusing on the class of infinite Bernoulli convolutions, which are defined by iterated function systems with overlaps, especially those defined by the Pisot, Garsia, and Salem numbers. By using an algorithm of Mantica, we obtain graphs of the coefficients of the 3-term recursion relation defining the orthogonal polynomials. We use these graphs to predict whether the singular infinite Bernoulli convolutions belong to the Nevai class. Based on our numerical results, we conjecture that all infinite Bernoulli convolutions with contraction ratios greater than or equal to 1/2 belong to Nevai’s class, regardless of the probability weights assigned to the self-similar measures.

97 MATHEMATICS AND COMPUTING↗

Direct estimation of the density of states for fermionic systems

Simulating time evolution is one of the most natural applications of quantum computers and is thus one of the most promising prospects for achieving practical quantum advantage. Here, we develop quantum algorithms to extract thermodynamic properties by estimating the density of states (DOS), which is a central object in quantum statistical mechanics. We introduce several key innovations that significantly improve the practicality and extend the generality of previous techniques. First, our approach allows one to estimate the DOS only for a specific subspace of the full Hilbert space. This is crucial for fermionic systems, since both canonical and grand canonical ensemble thermal equilibrium properties depend on subspaces of fixed number. Second, in our approach, by time evolving very simple, random initial states, such as randomly chosen computational basis states, we can exactly recover the DOS on average. Third, due to circuit-depth limitations, we only reconstruct the DOS up to a convolution with a Gaussian window—thus all imperfections that shift the energy levels by less than the width of the convolution window will not significantly affect the estimated DOS. For these reasons, we find the approach is a promising candidate for early quantum advantage as even short-time, noisy dynamics can yield a semiquantitative reconstruction of the DOS (convolution with a broad Gaussian window), while early fault-tolerant devices will likely enable higher-resolution DOS reconstruction through longer time evolutions. We demonstrate the practicality of our approach in representative Fermi-Hubbard and spin models and indeed find that our approach is highly robust against algorithmic errors in the time evolution and against gate noise. We further demonstrate that our approach is compatible with noisy intermediate-scale quantum (NISQ) computing NISQ-friendly variational techniques, introducing and leveraging a technique for variational time evolution.

97 MATHEMATICS AND COMPUTING↗