Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “fast algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Automatic Generation of Algorithms for High-Speed Reliable Lossy Data Compression (Final Report)

Fast reliable data compression is urgently needed for many leading-edge scientific instruments and for exascale high-performance computing applications because they produce vast amounts of data at extremely high rates. The goal of this project has been to develop a framework named LC that is able to automatically generate high-speed lossless and reliable lossy compression and decompression algorithms that can be customized for different kinds of data. The resulting LC framework is freely available on GitHub. To achieve high-speed operation, LC outputs optimized and parallelized CPU and GPU implementations of the generated algorithms. To ensure the quality of lossily compressed data, LC guarantees the user-provided error bound. To be able to customize the compression algorithm to various use cases, LC can synthesize millions of different algorithms and automatically search for the one that works best for the given data. We have already employed LC to create state-of-the-art lossless and lossy compressors for scientific data as well as leading lossless compressors for images. We hope that LC and the customized, fast, reliable, and CPU/GPU-compatible compression algorithms that it can generate will greatly benefit the many scientific applications that need not only high trustworthiness but also high performance.

97 MATHEMATICS AND COMPUTING↗

SREMI: Super-resolution electromagnetic imaging with single-channel ground-penetrating radar

High-resolution near-surface imaging has important applications in civil engineering, infrastructure inspection, military threat detection, geological characterization, and lunar and planetary exploration. Zero-offset, singlechannel ground penetrating radar (GPR) imaging is an established technique for near-surface target imaging and sensing but often suffers from low spatial resolution and imaging artifacts, especially of deep structures. In response, we formulate the GPR imaging as a dual-sparsity optimization problem, and develop a super-resolution electromagnetic imaging method based on a fast iterative shrinkage-thresholding algorithm. We develop our GPR imaging method in the framework of electromagnetic exploding-reflectors simulation theory, therefore the imaging method is computationally efficient. In this work, we demonstrate through synthetic and field data examples that our method can produce sharper, more reliable images with fewer artifacts compared with single-pass reverse-time migration GPR method, thus leading to improved near-surface interpretation and object identification.

58 GEOSCIENCES↗

Explainability and human intervention in autonomous scanning probe microscopy

The broad adoption of machine learning (ML)-based autonomous experiments (AEs) in material characterization and synthesis requires strategies development for understanding and intervention in the experimental workflow. Here, we introduce and realize a post-experimental analysis strategy for deep kernel learning-based autonomous scanning probe microscopy. This approach yields real-time and post-experimental indicators for the progression of an active learning process interacting with an experimental system. We further illustrate how this approach can be applied to human-in-the-loop AEs, where human operators make high-level decisions at high latencies setting the policies for AEs, and the ML algorithm performs low-level, fast decisions. The proposed approach is universal and can be extended to other techniques and applications such as combinatorial library analysis.

47 OTHER INSTRUMENTATION↗

Quantitative imaging and automated fuel pin identification for passive gamma emission tomography

Compliance of member States to the Treaty on the Non-Proliferation of Nuclear Weapons is monitored through nuclear safeguards. The Passive Gamma Emission Tomography (PGET) system is a novel instrument developed within the framework of the International Atomic Energy Agency (IAEA) project JNT 1510, which included the European Commission, Finland, Hungary and Sweden. The PGET is used for the verification of spent nuclear fuel stored in water pools. Advanced image reconstruction techniques are crucial for obtaining high-quality cross-sectional images of the spent-fuel bundle to allow inspectors of the IAEA to monitor nuclear material and promptly identify its diversion. In this work, we have developed a software suite to accurately reconstruct the spent-fuel cross sectional image, automatically identify present fuel rods, and estimate their activity. Unique image reconstruction challenges are posed by the measurement of spent fuel, due to its high activity and the self-attenuation. While the former is mitigated by detector physical collimation, we implemented a linear forward model to model the detector responses to the fuel rods inside the PGET, to account for the latter. The image reconstruction is performed by solving a regularized linear inverse problem using the fast-iterative shrinkage-thresholding algorithm. We have also implemented the traditional filtered back projection (FBP) method based on the inverse Radon transform for comparison and applied both methods to reconstruct images of simulated mockup fuel assemblies. Higher image resolution and fewer reconstruction artifacts were obtained with the inverse-problem approach, with the mean-square-error reduced by 50%, and the structural-similarity improved by 200%. We then used a convolutional neural network (CNN) to automatically identify the bundle type and extract the pin locations from the images; the estimated activity levels finally being compared with the ground truth. The proposed computational methods accurately estimated the activity levels of the present pins, with an associated uncertainty of approximately 5%.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Model-less Source Location for Forced Oscillation based on Synchrophasor and Moving Fast Fourier Transformation

Forced oscillations in power systems occur when the grid is driven by an external and periodic force. To quickly detect and locate the source of the forced oscillation is critical in terms of ensuring the reliability of an interconnected power grid. This paper explores the electromechanical wave propagation theory and the Fast Fourier Transformation to analyze the forced oscillations. It proposes a model-less, adaptive, fast, and accurate source location algorithm. The proposed algorithm is extensively evaluated through simulation data from a 70k-bus U.S. Eastern Interconnection test system and field-collected synchrophasor data from the distribution-level wide-area monitoring system, FNET/GridEye. The evaluation results demonstrate the correctness and effectiveness of the proposed model-less forced oscillation source location algorithm.

Wang, Weikang↗

Software-Defined Virtual Synchronous Condenser

Synchronous condensers (SCs) play important roles in integrating wind energy into relatively weak power grids. However, the design of SCs usually depends on specific application requirements and may not be adaptive enough to the frequently-changing grid conditions caused by the transition from conventional to renewable power generation. This paper devises a software-defined virtual synchronous condenser (SDViSC) method to address the challenges. Our contributions are fourfold: 1) design of a virtual synchronous condenser (ViSC) to enable full converter wind turbines to provide built-in SC functionalities; 2) engineering SDViSCs to transfer hardware-based ViSC controllers into software services, where a Tustin transformation-based software-defined control algorithm guarantees accurate tracking of fast dynamics under limited communication bandwidth; 3) a software-defined networking-enhanced SDViSC communication scheme to allow enhanced communication reliability and reduced communication bandwidth occupation; and 4) Prototype of SDViSC on our real-time, cyber-in-the-loop digital twin of large-wind-farm in an RTDS environment. Furthermore, extensive test results validate the excellent performance of SDViSC to support reliable and resilient operations of wind farms under various physical and cyber conditions.

17 WIND ENERGY↗

Constrained Local Approximate Ideal Restriction for Advection-Diffusion Problems

Herein this paper focuses on developing a reduction-based algebraic multigrid (AMG) method that is suitable for solving general (non)symmetric linear systems and is naturally robust from pure advection to pure diffusion. Initial motivation comes from a new reduction-based AMG approach, $\ell \text{AIR}$ (local approximate ideal restriction), that was developed for solving advection-dominated problems. Though this new solver is very effective in the advection-dominated regime, its performance degrades in cases where diffusion becomes dominant. This is consistent with the fact that in general, reduction-based AMG methods tend to suffer from growth in complexity and/or convergence rates as the problem size is increased, especially for diffusion-dominated problems in two or three dimensions. Motivated by the success of $\ell \text{AIR}$ in the advective regime, our aim in this paper is to generalize the AIR framework with the goal of improving the performance of the solver in diffusion-dominated regimes. To do so, we propose a novel way to combine mode constraints as used commonly in energy-minimization AMG methods with the local approximation of ideal operators used in $\ell \text{AIR}$. The resulting constrained $\ell \text{AIR}$ algorithm is able to achieve fast scalable convergence on advective and diffusive problems. In addition, it is able to achieve standard low complexity hierarchies in the diffusive regime through aggressive coarsening, something that was previously difficult for reduction-based methods.

97 MATHEMATICS AND COMPUTING↗

Penalized ensemble Kalman filters for high dimensional non-linear systems

The ensemble Kalman filter (EnKF) is a data assimilation technique that uses an ensemble of models, updated with data, to track the time evolution of a usually non-linear system. It does so by using an empirical approximation to the well-known Kalman filter. However, its performance can suffer when the ensemble size is smaller than the state space, as is often necessary for computationally burdensome models. This scenario means that the empirical estimate of the state covariance is not full rank and possibly quite noisy. To solve this problem in this high dimensional regime, we propose a computationally fast and easy to implement algorithm called the penalized ensemble Kalman filter (PEnKF). Under certain conditions, it can be theoretically proven that the PEnKF will be accurate (the estimation error will converge to zero) despite having fewer ensemble members than state dimensions. Further, as contrasted to localization methods, the proposed approach learns the covariance structure associated with the dynamical system. These theoretical results are supported with simulations of several non-linear and high dimensional systems.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

FastACE

SAND2024-01893O The Fast Adaptive Cosine Estimator (FastACE) algorithm modifies the well-known Adaptive Cosine Estimator for target detection in hyperspectral imagery. Specifically, FastACE modifies the computation of the background precision matrix (C_b^-1) under a Vecchia approximation. That is, each spectral band, conditioned on a local neighborhood of bands around that band, is independent of the other spectral bands. The FastACE algorithm leverages a parameterizable, auto-regressive neighborhood for the conditional independence assumption. The underlying math used for computing detection scores is equivalent to ACE, but it is implemented within FastACE. The software implements a target detection algorithm and associated utilities for target detection in hyperspectral imagery. Provided with a hyperspectral image and corresponding target signature, it produces relevant background statistics and the corresponding detection scores for the target signature in the image. The software is designed to integrate into other end-user applications or processing. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

VanderLaan, John↗

Advection algorithms for quantum neutrino moment transport

Neutrino transport in compact objects is an inherently challenging multidimensional problem. Here, this difficulty is compounded if one includes flavor transformation—an intrinsically quantum phenomenon requiring one to follow the coherence between flavors and thus necessitating the introduction of complex numbers. To reduce the computational burden, simulations of compact objects that include neutrino transport often make use of momentum-angle-integrated moments (the lowest order ones being commonly referred to as the energy density and flux) and these quantities can be generalized to include neutrino flavor, i.e., they become quantum moments. Numerous finite-volume approaches to solving the moment evolution equations for classical neutrino transport have been developed based on solving a Riemann problem at cell interfaces. In this paper we describe our generalization of a Riemann solver for quantum moments, specifically decomposing complex numbers in terms of a (signed) magnitude and phase instead of real and imaginary parts. We then test our new algorithm in numerous cases showing a neutrino fast flavor instability, varying from toy models with analytic solutions to snapshots from neutron star merger simulations. Compared to previous algorithms for neutrino transport with flavor mixing, we find uniformly smaller growth rates of the flavor transformation along with concomitantly larger length-scales, and that the results are a better match with the growth rates seen from multiangle codes.

79 ASTRONOMY AND ASTROPHYSICS↗

Recent Development of Frequency Estimation Methods for Future Smart Grid

The frequency estimated by the Phasor Measurement Unit (PMU) is a critical index of power system status and supports many smart grid applications. The future smart grid features high penetration of renewables and more fast-moving power electronics inverters but raises challenges to the reliable frequency estimation. This article presents three methods to address these challenges. First, an enhanced zero-crossing algorithm was developed to track the fast-changing frequency in system dynamics. Second, we propose a technology that can tolerate the system transient and suppress the outliers. Third, an algorithm was developed to export high time-resolution frequency estimations with minimum computational effort. All of the proposed methods are realized in hardware and compared with classical frequency estimation methods. The testing results indicate that the proposed methods have excellent performance. They can be used in future PMUs and provide reliable and high time resolution data for smart grid applications.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Fast and efficient identification of anomalous galaxy spectra with neural density estimation

ABSTRACT Current large-scale astrophysical experiments produce unprecedented amounts of rich and diverse data. This creates a growing need for fast and flexible automated data inspection methods. Deep learning algorithms can capture and pick up subtle variations in rich data sets and are fast to apply once trained. Here, we study the applicability of an unsupervised and probabilistic deep learning framework, the probabilistic auto-encoder, to the detection of peculiar objects in galaxy spectra from the SDSS survey. Different to supervised algorithms, this algorithm is not trained to detect a specific feature or type of anomaly, instead it learns the complex and diverse distribution of galaxy spectra from training data and identifies outliers with respect to the learned distribution. We find that the algorithm assigns consistently lower probabilities (higher anomaly score) to spectra that exhibit unusual features. For example, the majority of outliers among quiescent galaxies are E+A galaxies, whose spectra combine features from old and young stellar population. Other identified outliers include LINERs, supernovae, and overlapping objects. Conditional modelling further allows us to incorporate additional information. Namely, we evaluate the probability of an object being anomalous given a certain spectral class, but other information such as metrics of data quality or estimated redshift could be incorporated as well. We make our code publicly available.

Böhm, Vanessa↗

Technoeconomic Design Optimization for Fast Reactors. Part I: Workflow Development and Case Study for Small LFR District Energy Application

The nuclear industry is developing small reactor designs that can target a variety of deployment locations and energy products. Smaller nuclear designs have traditionally struggled to handle the steep trade-offs between size and cost that have historically incentivized large reactors. This motivates computational optimization of small reactors to minimize costs and quantify the trade-off between size and cost. In this paper, the cost/size trade-off for a small fast reactor is derived using a multi-objective genetic algorithm optimization, with steady-state, transient, and cost analysis of the fast reactor being performed. Specifically, the method is demonstrated on a small 10- to 120-MW(thermal) U-Pu-Zr–fueled lead-cooled fast reactor with a 10-year core life for district energy applications, which can have a thermal load compatible with this range. The results reinforced that fast reactor cores at the lower end of this power range suffer cost penalties due to critical mass considerations. It was found that high power density cores with strong reactivity swings and many control rods were favored over designing to minimize reactivity swing. Furthermore, this contrasts with some traditional configurations designed using engineering judgment and demonstrates that optimizers can find nontraditional but realistic solutions, along with demonstrating the value of incorporating cost functions into whole-reactor design optimization.

Fast reactor↗

Optimizing High Performance Markov Clustering for Pre-Exascale Architectures

HipMCL is a high-performance distributed memory implementation of the popular Markov Cluster Algorithm (MCL) and can cluster large-scale networks within hours using a few thousand CPU-equipped nodes. It relies on sparse matrix computations and heavily makes use of the sparse matrix-sparse matrix multiplication kernel (SpGEMM). The existing parallel algorithms in HipMCL are not scalable to Exascale architectures, both due to their communication costs dominating the runtime at large concurrencies and also due to their inability to take advantage of accelerators that are increasingly popular. In this work, we systematically remove scalability and performance bottlenecks of HipMCL. We enable GPUs by performing the expensive expansion phase of the MCL algorithm on GPU. Additionally, we propose a CPU-GPU joint distributed SpGEMM algorithm called pipelined Sparse SUMMA and integrate a probabilistic memory requirement estimator that is fast and accurate. Furthermore, we develop a new merging algorithm for the incremental processing of partial results produced by the GPUs, which improves the overlap efficiency and the peak memory usage. We also integrate a recent and faster algorithm for performing SpGEMM on CPUs. We validate our new algorithms and optimizations with extensive evaluations. With the enabling of the GPUs and integration of new algorithms, HipMCL is up to 12.4x faster, being able to cluster a network with 70 million proteins and 68 billion connections just under 15 minutes using 1024 nodes of ORNL's Summit supercomputer.

97 MATHEMATICS AND COMPUTING↗

Integral Kernel Methods for Nonlinear Parabolic-Elliptic Systems

Nonlinear parabolic-elliptic systems arise in many physical, biological, and chemical phenomena such as chemotaxis, ion transport, self-gravitating particles, and Brownian vortices. Existing methods struggle with the strong coupling and high nonlinearity and nonlocality of some of these systems, especially the ill-conditioned, convection-dominated problems. To overcome numerical difficulties, current approaches rely on initial guesses, preconditioning, or iterative techniques with no convergence guarantees. They might suffer from poor scalability, large memory usage, and difficulty to parallelize. Inspired by the connection of parabolic-elliptic systems to stochastic processes, we introduce a novel meshless, monolithic, and fully explicit method that naturally encapsulates the elliptic and parabolic operators into a single step which updates each node deterministically with global information. By being fully quadrature-based, it avoids solving systems of discretized equations and does not utilize initial guesses or preconditioning, while requiring little memory and being easy to parallelize. We first derive the method in an integral kernel formulation with quadratic complexity in the number of integration nodes and then leverage kernel-independent fast multipole methods (FMM) to present a scalable algorithm with linear complexity. We provide numerical examples for the Poisson-Nernst-Planck equations in one, two, and three dimensions, together with the derivation of the integral kernel for each case. Furthermore, the examples demonstrate the fast convergence and scalability of the FMM-accelerated algorithm, as well as its suitability for convection-dominated problems, making it competitive against traditional PDE solvers.

PDE systems↗

Examination of Semi-Analytical Solution Methods in the Coarse Operator of Parareal Algorithm for Power System Simulation

With continuing advances in high-performance parallel computing platforms, parallel algorithms have become powerful tools for development of faster than real-time power system dynamic simulations. In particular, it has been demonstrated in recent years that parallel-in-time (Parareal) algorithms have the potential to achieve such an ambitious goal. Here, the selection of a fast and reasonably accurate coarse operator of the Parareal algorithm is crucial for its effective utilization and performance. This paper examines semi-analytical solution (SAS) methods as the coarse operators of the Parareal algorithm and explores performance of the SAS methods to the standard numerical time integration methods. Two promising time-power series-based SAS methods were considered; Adomian decomposition method and Homotopy analysis method with a windowing approach for improving the convergence. Numerical performance case studies on 10-generator 39-bus system and 327-generator 2383-bus system were performed for these coarse operators over different disturbances, evaluating the number of Parareal iterations, computational time, and stability of convergence. All the coarse operators tested with different scenarios have converged to the same corresponding true solution (if they are convergent) and the SAS methods provide comparable computational speed, while having more stable convergence to the true solution in many cases.

97 MATHEMATICS AND COMPUTING↗

Implicit–explicit multirate infinitesimal stage-restart methods

Implicit–Explicit (IMEX) methods are flexible numerical time integration methods which solve an initial-value problem (IVP) that is split into stiff and nonstiff processes with the goal of lower computational costs than a purely implicit or explicit approach. A complementary form of flexible IVP solvers are multirate infinitesimal methods for problems split into fast- and slow-changing dynamics, that solve a multirate IVP by evolving a sequence of “fast” IVPs using any suitably accurate algorithm. This article introduces a new class of high-order implicit–explicit multirate methods that are designed for multirate IVPs in which the slow-changing dynamics are further split in an IMEX fashion. This new class, which we call implicit–explicit multirate infinitesimal stage-restart (IMEX-MRI-SR), both improves upon the previous implicit–explicit multirate infinitesimal generalized-structure additive Runge Kutta (IMEX-MRI-GARK) methods by allowing for far easier creation of new embedded methods, and extends multirate exponential Runge Kutta (MERK) methods by allowing the fast-changing dynamics to be nonlinear and the methods to be implicit. We leverage GARK theory to derive conditions for orders of accuracy up to four, and we provide second- and third-order accurate example methods, which are the first known embedded MRI methods with IMEX structure. We then perform numerical simulations demonstrating convergence rates and computational performance in both fixed-step and adaptive-step settings.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Computation-Efficient Algorithm for Distributed Feedback Optimization of Distribution Grids

Feedback-based optimization algorithms use real-time measurements to update the optimal control for the underlying system which may not be fully identified. Recently, we have developed a distributed feedback-based algorithm [1] that avoids the requirement of fast communication between central computing and local actuator/sensor agents. This paper extends the work by greatly reducing the number of copies of variables involved in the distributed feedback-based algorithm, which results in faster convergence and lower communication requirement. The main idea is to leverage the specific structural properties of the admittance matrix for distribution systems with tree network topology. We also show the effectiveness of the proposed algorithm in simulations.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗