Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “partitioned algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Unsupervised learning of representative local atomic arrangements in molecular dynamics data

Molecular dynamics (MD) simulations present a data-mining challenge, given that they can generate a considerable amount of data but often rely on limited or biased human interpretation to examine their information content. By not asking the right questions of MD data we may miss critical information hidden within it. Here we combine dimensionality reduction (UMAP) and unsupervised hierarchical clustering (HDBSCAN) to quantitatively characterize prevalent coordination environments of chemical species within MD data. By focusing on local coordination, we significantly reduce the amount of data to be analyzed by extracting all distinct molecular formulas within a given coordination sphere. We then efficiently combine UMAP and HDBSCAN with alignment or shape-matching algorithms to partition these formulas into structural isomer families indicating their relative populations. The method was employed to reveal details of cation coordination in electrolytes based on molecular liquids.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Autonomous Anomaly Detection For Continuous Streams

The code implements the Isolation Forest (IFML) algorithm within the digital twin (DT) of the AGN-201 nuclear reactor. The DT captures real-time operational data including control rod positions, reactor power, and temperature. The IFML model isolates anomalies by detecting patterns that deviate from expected operational behavior. The algorithm recursively partitions the data and assigns anomaly scores based on the isolation of rare and different events. By tuning parameters specific to the reactor’s operational data, the IFML identifies deviations such as unauthorized material insertions or reactor reactivity shifts. The system streams data using LabView and integrates with the DeepLynx data warehouse for anomaly processing.

Trevino, Eduardo↗

Fast shared-memory streaming multilevel graph partitioning

In this report we show that a fast parallel graph partitioner can benefit many applications by reducing data transfers. The online methods for partitioning graphs have to be fast and they often rely on simple one-pass streaming algorithms, while the offline methods for partitioning graphs contain more involved algorithms and the most successful methods in this category belong to the multilevel approaches. In this work, we assess the feasibility of using streaming graph partitioning algorithms within the multilevel framework. Our end goal is to come up with a fast parallel offline multilevel partitioner that can produce competitive cutsize quality. We rely on a simple but fast and flexible streaming algorithm throughout the entire multilevel framework. This streaming algorithm serves multiple purposes in the partitioning process: a clustering algorithm in the coarsening, an effective algorithm for the initial partitioning, and a fast refinement algorithm in the uncoarsening. Its simple nature also lends itself easily for parallelization. The experiments on various graphs show that our approach is on the average up to 5.1x faster than the multi-threaded MeTiS, which comes at the expense of only 2x worse cutsize.

97 MATHEMATICS AND COMPUTING↗

Overset-Grid Method with Smooth Orbital Partitioning for Molecular Scattering Calculations

To solve molecular photoionization and electron scattering problems, we use an overset-grid representation of electronic continuum functions, which has an extended central spherical grid that overlaps small spherical grids (subgrids) centered on each atom of a polyatomic molecule. Here, in this work, we present an improved algorithm that smoothly partitions the total wave function between the central grid and the atomic subgrids. The smooth partitioning allows one to use approximately one-fourth the number of partial waves on the central grid compared to our previous implementation with switching functions. The resulting numerical method for treating electron scattering and photoionization of polyatomic molecules combines the accuracy and flexibility of pure numerical grid representations with the rapid convergence of hybrid combinations of atom-centered basis-set expansions and grid methods. The overset-grid representation is implemented using the complex Kohn variational principle for scattering and photoionization amplitudes. The faster convergence with respect to the number of central grid partial waves is demonstrated and accuracy is verified by comparisons with the previous implementation and with far more computationally demanding single-center numerical expansions in electron-molecule scattering and photoionization calculations on the neon dimer (Ne 2 ) system, carbon tetrafluoride (CF 4 ) molecule, and the pyridine (C 5 H 5 N) molecule in the static-exchange approximation.

Molecules↗

Quantum Tensor-Product Decomposition from Choi-State Tomography

The Schmidt decomposition is the go-to tool for measuring bipartite entanglement of pure quantum states. Similarly, it is possible to study the entangling features of a quantum operation using its operator-Schmidt or tensor-product decomposition. While quantum technological implementations of the former are thoroughly studied, entangling properties on the operator level are harder to extract in the quantum computational framework because of the exponential nature of sample complexity. Here, we present an algorithm for unbalanced partitions into a small subsystem and a large one (the environment) to compute the tensor-product decomposition of a unitary the effect of which on the small subsystem is captured in classical memory, while the effect on the environment is accessible as a quantum resource. This quantum algorithm may be used to make predictions about operator nonlocality and effective open quantum dynamics on a subsystem, as well as for finding low-rank approximations and low-depth compilations of quantum circuit unitaries. We demonstrate the method and its applications on a time-evolution unitary of an isotropic Heisenberg model in two dimensions. Published by the American Physical Society 2024

Mansuroglu, Refik (ORCID:000000017352513X)↗

A Higher Order, Stable Partitioned Scheme for Fluid-Structure Interaction Problems

Although still a very active area of research with many open questions, there exist highly-accurate and efficient algorithms for numerically estimating solutions to the equations of fluid motion. Similarly, great strides have been made in numerically modeling the motion of solids. However, there exist many practical applications where one must solve both fluid and structure in tandem. There is a gap in the literature for high order unconditionally stable partitioned algorithms for problems of this type. We investigate a new partitioned scheme which shows promise in filling this gap. We use a finite-element-based model with an added mass approach on problems coupling the Euler beam equation with the incompressible Navier-Stokes equations.

97 MATHEMATICS AND COMPUTING↗

Accelerated impurity solver for DMFT and its diagrammatic extensions

Here, we present ComCTQMC, a GPU accelerated quantum impurity solver. It uses the continuous-time quantum Monte Carlo (CTQMC) algorithm wherein the partition function is expanded in terms of the hybridisation function (CT-HYB). ComCTQMC supports both partition and worm-space measurements, and it uses improved estimators and the reduced density matrix to improve observable measurements whenever possible. ComCTQMC efficiently measures all one and two-particle Green's functions, all static observables which commute with the local Hamiltonian, and the occupation of each impurity orbital. ComCTQMC can solve complex-valued impurities with crystal fields that are hybridized to both fermionic and bosonic baths. Most importantly, ComCTQMC utilizes graphical processing units (GPUs), if available, to dramatically accelerate the CTQMC algorithm when the Hilbert space is sufficiently large. We demonstrate acceleration by a factor of over 600 (100) in a simulation of δ-Pu at 600 K with (without) crystal fields. In easier problems, the GPU offers less impressive acceleration or even decelerates the CTQMC. Here we describe the theory, algorithms, and structure used by ComCTQMC in order to achieve this set of features and level of acceleration.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Machine learning-aided inverse design for biogas upgrading through biological CO 2 conversion

The biogas upgrading process through bioconversion of CO 2 to CH 4 by hydrogenotrophic methanogens is an attractive strategy for energy decarbonation. Many studies have optimized operational parameters to improve key performance indicators such as CH 4 % and H 2 utilization efficiency. However, inconsistent laboratory conditions make it challenging to compare results. Existing models for analyzing operating conditions can only assess the impact of individual conditions and lack the ability to simultaneously optimize multiple conditions. To address this, two XGBoost models were built with R 2 of 0.779 and 0.903 with data collected from literatures and were embedded into multi-objective partitive swarm optimization algorithm to optimal operating conditions. Predictions were compared with experimental validations under optimized conditions, revealing an 8.50% and 2.95% relative error in CH 4 % and H 2 conversion rate, respectively. This approach streamlines biogas upgrading processes, offering a data-driven solution to enhance efficiency and consistency in the pursuit of sustainable methane production.

Biogas upgrading↗

Hierarchical Network Partitioning for Solution of Potential-Driven, Steady-State Nonlinear Network Flow Equations

The solution of potential-driven steady-state flow in large networks is a task which manifests in various engineering applications, such as transport of natural gas or water through pipeline networks. The resultant system of nonlinear equations depends on the network topology, and in general, there is no numerical algorithm that offers guaranteed convergence to the solution (assuming a solution exists). Some methods offer guarantees in cases where the network topology satisfies certain assumptions, but these methods fail for larger networks. On the other hand, the Newton-Raphson algorithm offers a convergence guarantee if the starting point lies close to the (unknown) solution. It would be advantageous to compute the solution of the large nonlinear system through the solution of smaller nonlinear sub-systems wherein the solution algorithms (Newton-Raphson or otherwise) are more likely to succeed. Here, this letter proposes and describes such a procedure, a hierarchical network partitioning algorithm that enables the solution of large nonlinear systems corresponding to potential-driven steady-state network flow equations.

42 ENGINEERING↗

A Fast Algorithm for Scanning Transmission Electron Microscopy Imaging and 4D-STEM Diffraction Simulations

Scanning transmission electron microscopy (STEM) is an extremely versatile method for studying materials on the atomic scale. Many STEM experiments are supported or validated with electron scattering simulations. However, using the conventional multislice algorithm to perform these simulations can require extremely large calculation times, particularly for experiments with millions of probe positions as each probe position must be simulated independently. Recently, the plane-wave reciprocal-space interpolated scattering matrix (PRISM) algorithm was developed to reduce calculation times for large STEM simulations. Here, we introduce a new method for STEM simulation: partitioning of the STEM probe into “beamlets,” given by a natural neighbor interpolation of the parent beams. This idea is compatible with PRISM simulations and can lead to even larger improvements in simulation time, as well requiring significantly less computer random access memory (RAM). We have performed various simulations to demonstrate the advantages and disadvantages of partitioned PRISM STEM simulations. We find that this new algorithm is particularly useful for 4D-STEM simulations of large fields of view. We also provide a reference implementation of the multislice, PRISM, and partitioned PRISM algorithms.

97 MATHEMATICS AND COMPUTING↗

Distributed Tomographic Reconstruction with Quantization

Conventional tomographic reconstruction typically depends on centralized servers for both data storage and computation, leading to concerns about memory limitations and data privacy. Distributed reconstruction algorithms mitigate these issues by partitioning data across multiple nodes, reducing server load and enhancing privacy. However, these algorithms often encounter challenges related to memory constraints and communication overhead between nodes. In this paper, we introduce a decentralized Alternating Directions Method of Multipliers (ADMM) with configurable quantization. By distributing local objectives across nodes, our approach is highly scalable and can efficiently reconstruct images while adapting to available resources. To overcome communication bottlenecks, we propose two quantization techniques based on K-means clustering and JPEG compression. Numerical experiments with benchmark images illustrate the tradeoffs between communication efficiency, memory use, and reconstruction accuracy.

Miao, Runxuan↗

Slicing with deep learning models at ProtoDUNE-SP

DUNE is a cutting-edge experiment aiming to study neutrinos in detail, with a special focus on the flavor oscillation mechanism. The prototype of the DUNE Far Detector Single Phase TPC (ProtoDUNE-SP) was built and operated at CERN with a full set of reconstruction tools. To implement these reconstruction tools, Pandora, a multi-algorithm framework, has been developed. A large number of these algorithms, some of them being exploiting traditional clustering, detector physics and deep learning approaches, have been applied to images to gradually build up a picture out of singular events. One of such algorithms is the Pandora slicing algorithm which aims to partition the detector hits of an event in sets called slices. Each slice represents a single interaction in the detector and should identify all the hits related to the interacting particle and its subsequent decay products. We expect the order of tens of slices per event in ProtoDUNE-SP. In this paper we present a deep learning approach to the problem, designing a model able to outperform the state-of-the-art slicing algorithm which is currently implemented within Pandora. We assess the performance of our tool in terms of efficiency and accuracy, while exploiting hardware accelerating setups. The ultimate goal is to incorporate this deep learning approach in the Pandora reconstruction tool.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Communication Lower Bounds and Optimal Algorithms for Symmetric Matrix Computations

In this article, we focus on the communication costs of three symmetric matrix computations: (i) multiplying a matrix with its transpose, known as a symmetric rank-k update (SYRK) (ii) adding the result of the multiplication of a matrix with the transpose of another matrix and the transpose of that result, known as a symmetric rank-2k update (SYR2K) (iii) performing matrix multiplication with a symmetric input matrix (SYMM). All three computations appear in the Level 3 Basic Linear Algebra Subroutines (BLAS) and have wide use in applications involving symmetric matrices. We establish communication lower bounds for these kernels using sequential and distributed-memory parallel computational models, and we show that our bounds are tight by presenting communication-optimal algorithms for each setting. Our lower bound proofs rely on applying a geometric inequality for symmetric computations and analytically solving constrained nonlinear optimization problems. As a result, the symmetric matrix and its corresponding computations are accessed and performed according to a triangular block partitioning scheme in the optimal algorithms.

Al Daas, Hussam [Rutherford Appleton Laboratory, D↗

Scalable Approaches to Selecting Key Entities in Large Networked Infrastructure Systems

This work aims at bringing advances in discrete optimization algorithms to solving practical engineering problems at scale. Often times, in many engineering design problems, there is a need to select a small set of influential or representative elements from a large ground set of entities in an optimal fashion. Submodular optimization provides for a formal way to solve such problems. Common examples with infrastructure systems involve sensor placement and identification of key entities with certain objectives. However, scaling these approaches to large infrastructure systems can be challenging because of the high computational complexity of the overall framework that include the optimization algorithms as well as high-complexity compute-oracles that provide the necessary objective function values. In this work, we explore a well-studied and widely-applicable paradigm, namely leader-selection in a multi-agent networked setting in the context of scalable methodologies. We demonstrate novel frameworks that utilize variations of accelerated submodular optimization algorithms along with linear-algebraic methods that can help accelerate the oracle computations. We further explore this combination in conjunction with graph partitioning paradigms to take advantage of the accelerated algorithms in a distributed setting. Finally we demonstrate the key findings on a practical problem in an operational setting. For this, we leverage an example road network with approximately 18k nodes and 27k edges in a traffic control application, where we seek a limited number of k=200 key intersections. This problem can be solved in a serial setting in just under 5 hours providing more than 2 orders of magnitude speed-up over methods that do not consider acceleration techniques.

Visweswara Sathanur, Arun↗

An Orthogonal Recursive Bisection (ORB) Based Time Advancement Algorithm for CFD-DEM Solvers

The time integration of the granular phase in coupled computational fluid dynamics (CFD) – discrete element method (DEM) simulations presents a unique computational challenge brought about by the large variations in particle collisional time scales. Particles in the dilute regions of the computational domain can be advanced with large time steps while dense regions require much smaller time increments. However, the time step size in most solvers is globally set as the limit for accuracy and stability imposed by the collisions and is typically orders of magnitude less than that required away from collisions. This work addresses this precise issue and provides a strategy to avoid the use of a global conservative small time step size for the entire set of particles.A novel time stepping algorithm for CFD-DEM solvers using a partitioning approach using orthogonal recursive bisection (ORB) that allows for variable time steps among particles is described and its computational performance is compared against baseline explicit methods, typically used in several CFD-DEM solvers. ORB has advantages of being relatively quick and easy to update incrementally and has the required heuristic behavior (i.e., it will split the region in half with a cluster on each side) when groups of particles are well separated (clustered). The algorithm presented in this work uses a local time stepping approach to resolve collisional time scales for subsets of particles that are present at the leaves of the ORB, thereby resulting in substantial reduction of computational cost. The parallel implementation of this method where a ``knapsack” algorithm is used in tandem with ORB for effective load-balancing is also presented, where a best possible partitioning is obtained based on number of particles and local time-stepping costs. The algorithm is tested against benchmark problems with varying particle distributions that include fluidized bed and riser flow scenarios. Preliminary results indicate that the approach is 2-3X faster than traditional explicit methods for problems that involve both dense and dilute regions, while maintaining the same level of accuracy.

adaptive timestepping↗

Shallow-circuit variational quantum eigensolver based on symmetry-inspired Hilbert space partitioning for quantum chemical calculations

Development of resource-friendly quantum algorithms remains highly desirable for noisy intermediate-scale quantum computing. Based on the variational quantum eigensolver (VQE) with unitary coupled-cluster Ansatz, we demonstrate that partitioning of the Hilbert space made possible by the point-group symmetry of the molecular systems greatly reduces the number of variational operators by confining the variational search within a subspace. In addition, we found that instead of including all subterms for each excitation operator, a single-term representation suffices to reach required accuracy for various molecules tested, resulting in an additional shortening of the quantum circuit by a factor of 4–8. With these strategies, VQE calculations on a noise-free quantum simulator achieve energies within a few meVs of those obtained with the full unitary coupled-cluster Ansatz with single and double excitations for the H 4 -square, H 4 -chain, and H 6 -hexagon molecules, while the number of cnot gates, a measure of the quantum-circuit depth, is reduced by a factor of as large as 35. Furthermore, we introduced an efficient “score” parameter to rank the excitation operators, so that the operators causing larger energy reduction can be applied first. Using the H 4 square and H 4 chain as examples, We demonstrated on noisy quantum simulators that the first few variational operators can bring the energy within the chemical accuracy, while additional operators do not improve the energy since the accumulative noise outweighs the gain from the expansion of the variational Ansatz.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

A Mountaintop View Requires Minimal Sorting: A Faster Contour Tree Algorithm

Consider a scalar field f : M → R, where M is a triangulated simplicial mesh in R d . A level set, or contour, at value v is a connected component of f –1 (v). As v is changed, these contours change topology, merge into each other, or split. Contour trees are concise representations of f that track this contour behavior. The vertices of these trees are the critical points of f, where the gradient is zero. The edges represent changes in the topology of contours. It is a fundamental data structure in data analysis and visualization, and there is significant previous work (both theoretical and practical) on algorithms for constructing contour trees. Suppose M has n vertices, N facets, and t critical points. A classic result of Carr, Snoeyink, and Axen (2000) gives an algorithm that takes O(n log n+Nα(N)) time (where α(·) is the inverse Ackermann function). A further improvement to O(t log t + N) time was given by Chiang et al. All these algorithms involve a global sort of the critical points, a significant computational bottleneck. Unfortunately, lower bounds of Ω(t log t) also exist. We present the first algorithm that can avoid the global sort and has a refined time complexity that depends on the contour tree structure. Intuitively, if the tree is short and fat, we get significant improvements in running time. For a partition of the contour tree into a set of descending paths, P, our algorithm runs in O($\Sigma$ pϵP |p| log |p| + tα(t) + N). This is at most O(t log D + N), where D is the diameter of the contour tree. Moreover, it is O(tα(t) + N) for balanced trees, a significant improvement over the previous complexity. Our algorithm requires numerous ideas: partitioning the contour tree into join and split trees, a local growing procedure to iteratively build contour trees, and the use of heavy path decompositions for the time complexity analysis. There is a crucial use of a family of binomial heaps to maintain priorities, ensuring that any comparison made is between comparable nodes of the contour tree. We also prove lower bounds showing that the $\Sigma$ pϵP |p| log |p| complexity is inherent to computing contour trees.

97 MATHEMATICS AND COMPUTING↗

TDAG: Tree-based Directed Acyclic Graph Partitioning for Quantum Circuits

We propose the Tree-based Directed Acyclic Graph (TDAG) partitioning for quantum circuits, a novel quantum circuit partitioning method which partitions circuits by viewing them as a series of binary trees and selecting the tree containing the most gates. TDAG produces results of comparable quality (number of partitions) to an existing method called ScanPartitioner (an exhaustive search algorithm) with an 95% average reduction in execution time. Furthermore, TDAG improves compared to a faster partitioning method called QuickPartitioner by 38% in terms of quality of the results with minimal overhead in execution time.

Clark, Joseph↗