Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “butterfly algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

A Linear-Complexity Tensor Butterfly Algorithm for Compressing High-Dimensional Oscillatory Integral Operators

This paper presents a multilevel tensor compression algorithm called tensor butterfly algorithm for efficiently representing large-scale and high-dimensional oscillatory integral operators, including Green's functions for wave equations and integral transforms such as Radon transforms and Fourier transforms. The proposed algorithm leverages a tensor extension of the so-called complementary low-rank property of existing matrix butterfly algorithms. The algorithm partitions the discretized integral operator tensor into subtensors of multiple levels and factorizes each subtensor at the middle level as a Tucker-type interpolative decomposition, whose factor matrices are formed in a multilevel fashion. For a d-dimensional (d > 1) integral operator discretized into a 2d-mode tensor with n2d entries, the overall CPU time and memory requirement scale as O(nd), in stark contrast to the O(nd log n) complexity of existing matrix algorithms such as matrix butterfly algorithms and fast Fourier transforms (FFTs), where n is the number of points per direction. When comparing with other tensor algorithms such as quantized tensor train (QTT), the proposed algorithm also shows superior CPU and memory performance for tensor contraction. Remarkably, the tensor butterfly algorithm can efficiently model high-frequency Green's function interactions between two unit cubes, each spanning 512 wavelengths per direction, which represents problems of scale over 512× larger than that existing butterfly algorithms can handle, with the same amount of computation resources. On the other hand, for a problem representing 64 wavelengths per direction, which is the largest size existing algebraic matrix algorithms can handle, our tensor butterfly algorithm exhibits 200x speedups and 30× memory reduction compared with existing ones. Moreover, the tensor butterfly algorithm also permits O(nd)-complexity FFTs and Radon transforms up to d = 6 dimensions.

Kielstra, P Michael↗

Sparse Approximate Multifrontal Factorization with Composite Compression Methods

This article presents a fast and approximate multifrontal solver for large sparse linear systems. In a recent work by Liu et al., we showed the efficiency of a multifrontal solver leveraging the butterfly algorithm and its hierarchical matrix extension, HODBF (hierarchical off-diagonal butterfly) compression to compress large frontal matrices. The resulting multifrontal solver can attain quasi-linear computation and memory complexity when applied to sparse linear systems arising from spatial discretization of high-frequency wave equations. To further reduce the overall number of operations and especially the factorization memory usage to scale to larger problem sizes, in this article we develop a composite multifrontal solver that employs the HODBF format for large-sized fronts, a reduced-memory version of the nonhierarchical block low-rank format for medium-sized fronts, and a lossy compression format for small-sized fronts. This allows us to solve sparse linear systems of dimension up to 2.7 × larger than before and leads to a memory consumption that is reduced by 70% while ensuring the same execution time. The code is made publicly available in GitHub.

97 MATHEMATICS AND COMPUTING↗

Sparse Approximate Multifrontal Factorization with Butterfly Compression for High-Frequency Wave Equations

In this work, we present a fast and approximate multifrontal solver for large-scale sparse linear systems arising from finite-difference, finite-volume or finite-element discretization of high-frequency wave equations. The proposed solver leverages the butterfly algorithm and its hierarchical matrix extension for compressing and factorizing large frontal matrices via graph-distance guided entry evaluation or randomized matrix-vector multiplication-based schemes. Complexity analysis and numerical experiments demonstrate $\mathcal{O}(N\log^2 N)$ computation and $\mathcal{O}(N)$ memory complexity when applied to an $N\times N$ sparse system arising from 3D high-frequency Helmholtz and Maxwell problems.

97 MATHEMATICS AND COMPUTING↗

A Butterfly-Accelerated Volume Integral Equation Solver for Broad Permittivity and Large-Scale Electromagnetic Analysis

In this work, a butterfly-accelerated volume integral equation (VIE) solver is proposed for fast and accurate electromagnetic (EM) analysis of scattering from heterogeneous objects. The proposed solver leverages the hierarchical off-diagonal butterfly (HOD-BF) scheme to construct the system matrix and obtain its approximate inverse, used as a preconditioner. Complexity analysis and numerical experiments validate the O(N log 2 N) construction cost of the HOD-BF-compressed system matrix and O(N log 1.5 N) inversion cost for the preconditioner, where N is the number of unknowns in the high-frequency EM scattering problem. For many practical scenarios, the proposed VIE solver requires less memory and computational time to construct the system matrix and obtain its approximate inverse compared to a H matrix-accelerated VIE solver. The accuracy and efficiency of the proposed solver have been demonstrated via its application to the EM analysis of large-scale canonical and real-world structures comprising of broad permittivity values and involving millions of unknowns.

42 ENGINEERING↗

Butterfly Factorization Via Randomized Matrix-Vector Multiplications

This paper presents an adaptive randomized algorithm for computing the butterfly factorization of an m × n matrix with m ≈ n provided that both the matrix and its transpose can be rapidly applied to arbitrary vectors. The resulting factorization is composed of O(log n) sparse factors, each containing O(n) nonzero entries. The factorization can be attained using O(n 3/2 log n) computation and O(n log n) memory resources. Furthermore, the proposed algorithm can be implemented in parallel and can apply to matrices with strong or weak admissibility conditions arising from surface integral equation solvers as well as multi-frontal-based finite-difference, finite-element, or finite-volume solvers. A distributed-memory parallel implementation of the algorithm demonstrates excellent scaling behavior.

97 MATHEMATICS AND COMPUTING↗

A Fast Butterfly-Compressed Hadamard–Babich Integrator for High-Frequency Helmholtz Equations in Inhomogeneous Media with Arbitrary Sources

Here we present a butterfly-compressed representation of the Hadamard-Babich (HB) ansatz for the Green's function of the high-frequency Helmholtz equation in smooth inhomogeneous media. For a computational domain discretized with Nv discretization cells, the proposed algorithm first solves and tabulates the phase and HB coefficients via eikonal and transport equations with observation points and point sources located at the Chebyshev nodes using a set of much coarser computation grids, and then butterfly compresses the resulting HB interactions from all Nv cell centers to each other. The overall CPU time and memory requirement scale as O(Nv log2 Nv) for any bounded two-dimensional (2D) domains with arbitrary excitation sources. A direct extension of this scheme to bounded 3D domains yields an O(Nv4/3) CPU complexity, which can be further reduced to quasi-linear complexities with proposed remedies. The scheme can also efficiently handle scattering problems involving inclusions in inhomogeneous media. Although the current construction of our HB integrator does not accommodate caustics, the resulting HB integrator itself can be applied to certain sources, such as concave-shaped sources, to produce caustic effects. Compared to finite-difference frequency domain methods, the proposed HB integrator is free of numerical dispersion and requires fewer discretization points per wavelength. As a result, it can solve wave propagation problems well beyond the capability of existing solvers. Remarkably, the proposed scheme can accurately model wave propagation in 2D domains with 640 wavelengths per direction and in 3D domains with 54 wavelengths per direction on a state-of-the-art supercomputer at Lawrence Berkeley National Laboratory.

Hadamard--Babich ansatz↗

Synchrotron‐source micro‐x‐ray computed tomography for examining butterfly eyes

Comparative anatomy is an important tool for investigating evolutionary relationships among species, but the lack of scalable imaging tools and stains for rapidly mapping the microscale anatomies of related species poses a major impediment to using comparative anatomy approaches for identifying evolutionary adaptations. We describe a method using synchrotron source micro-x-ray computed tomography (syn-μXCT) combined with machine learning algorithms for high-throughput imaging of Lepidoptera (i.e., butterfly and moth) eyes. Our pipeline allows for imaging at rates of ~15 min/mm 3 at 600 nm 3 resolution. Image contrast is generated using standard electron microscopy labeling approaches (e.g., osmium tetroxide) that unbiasedly labels all cellular membranes in a species-independent manner thus removing any barrier to imaging any species of interest. To demonstrate the power of the method, we analyzed the 3D morphologies of butterfly crystalline cones, a part of the visual system associated with acuity and sensitivity and found significant variation within six butterfly individuals. Despite this variation, a classic measure of optimization, the ratio of interommatidial angle to resolving power of ommatidia, largely agrees with early work on eye geometry across species. We show that this method can successfully be used to determine compound eye organization and crystalline cone morphology. Our novel pipeline provides for fast, scalable visualization and analysis of eye anatomies that can be applied to any arthropod species, enabling new questions about evolutionary adaptations of compound eyes and beyond.

59 BASIC BIOLOGICAL SCIENCES↗

Evolution of the SLATE linear algebra library

SLATE (Software for Linear Algebra Targeting Exascale) is a distributed, dense linear algebra library targeting both CPU-only and GPU-accelerated systems, developed over the course of the Exascale Computing Project (ECP). While it began with several documents setting out its initial design, significant design changes occurred throughout its development. In some cases, these were anticipated: an early version used a simple consistency flag that was later replaced with a full-featured consistency protocol. In other cases, performance limitations and software and hardware changes prompted a redesign. Sequential communication tasks were parallelized; host-to-host MPI calls were replaced with GPU device-to-device MPI calls; more advanced algorithms such as Communication Avoiding LU and the Random Butterfly Transform (RBT) were introduced. Early choices that turned out to be cumbersome, error prone, or inflexible have been replaced with simpler, more intuitive, or more flexible designs. Applications have been a driving force, prompting a lighter weight queue class, nonuniform tile sizes, and more flexible MPI process grids. Of paramount importance has been building a portable library that works across several different GPU architectures – AMD, Intel, and NVIDIA – while keeping a clean and maintainable codebase. Here we explore the evolving design choices and their effects, both in terms of performance and software sustainability.

Gates, Mark↗

Numerical Investigation of Butterfly Valve Performance in Variable Valve Sizes, Positions and Flow Regimes

Reliability and efficiency of valves are necessary for precise control and sufficient heat-flow to heat application plants for the integrated energy systems of nuclear power plants (NPPs). Strategic Management Analysis Requirement and Technology (SMART) valves’ ability to control flow and assess environmental parameters stands out for these requirements. Their ability to sustain the downstream flow rate, prevent reverse flow, and maintain pressure in the heat transport loop is much more efficient with the integration of sensors and intelligent algorithms. For assessing valve performance and monitoring, mechanical design and operating conditions are two important parameters. In this study, the butterfly valves of three different sizes are simulated with water and steam using STAR-CCM+ in various flow regimes and positions to analyze performance parameters to strategize an automated control system for efficiently balancing the heat–transport network. Also, flow behavior is studied using velocity and pressure fields for valve–body geometry optimization. It can be observed, through performance parameters, that the valves are suitable for operation between 30° and 90° positions with significantly low loss coefficients and high flow coefficients, and the performance parameters follow a certain pattern in both water and steam flow in each scenario.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗