Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “hierarchical algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Real-time Simulation Framework for Hardware-in-the-Loop Testing of Multi-port Autonomous Reconfigurable Solar Power Plant (MARS)

Multiport autonomous reconfigurable solar power plant (MARS) has been proposed for integrated development of photovoltaic (PV) and energy storage system (ESS) that can connect to high-voltage direct current (HVdc) and alternating current (ac) transmission grid. To de-risk the development of this complex integrated system that consists of hundreds to thousands of power electronics modules, a controller hardware-in-the-loop (cHIL) test setup will be extremely beneficial. The cHIL testing can be used for evaluation of modules as well as the hierarchical control system in MARS. With the unique configuration of power electronics modules in MARS, it becomes necessary to develop custom-designed real-time simulation models in the cHIL setup in absence of off-the-shelf models. In this paper, high-fidelity dynamic model of MARS, control algorithms at the lower level, and required communication algorithms are developed and optimized for real-time performance in the cHIL setup. Real-time experimental results from the cHIL are provided.

Dong, Zerui↗

Reverse-mode differentiation in arbitrary tensor network format: with application to supervised learning.

This paper describes an efficient reverse-mode differentiation algorithm for contraction operations of tensor networks that may have arbitrary and unconventional network topologies. The approach leverages the tensor contraction tree of Evenbly and Pfeifer (2014), which provides an instruction set for the contraction sequence of a network. We show that this tree can be efficiently leveraged for differentiation of a full tensor network contraction using a recursive scheme that exploits (1) the bilinear property of contraction and (2) the property that trees have single path from root to leaves. While differentiation of tensor-tensor contraction is already possible in most automatic differentiation packages, we show that exploiting these two additional properties in the specific context of contraction sequences can improve efficiency. Following a description of the algorithm and computational complexity analysis, we investigate its utility for gradient-based supervised learning for low-rank function recovery and for fitting real-world unstructured datasets. We demonstrate improved performance over alternating least-squares optimization approaches and the capability to handle heterogeneous and arbitrary tensor network formats. When compared to alternating minimization algorithms, we find that the gradient-based approach requires a smaller oversampling ratio (number of samples compared to number model parameters) for recovery. This increased efficiency extends to fitting unstructured data of varying dimensionality and when employing a variety of tensor network formats. Here, we show improved learning using the hierarchical Tucker method over the tensor-train in high-dimensional settings on a number of benchmark problems.

97 MATHEMATICS AND COMPUTING↗

Multiscale Normalizing Flows for Gauge Theories

Scale separation is an important physical principle that has previously enabled algorithmic advances such as multigrid solvers. Previous work on normalizing flows has been able to utilize scale separation in the context of scalar field theories, but the principle has been largely unexploited in the context of gauge theories. This work gives an overview of a new method for generating gauge fields using hierarchical normalizing flow models. This method builds gauge fields from the outside in, allowing different parts of the model to focus on different scales of the problem. Numerical results are presented for $U(1)$ and $SU(3)$ gauge theories in 2, 3, and 4 spacetime dimensions.

Abbott, Ryan↗

Identification of crystal plasticity model parameters by multi-objective optimization integrating microstructural evolution and mechanical data

Crystal plasticity models evolve a polycrystalline yield surface using meso-scale descriptions of deformation mechanisms. The activation of deformation mechanisms is governed by crystallography and a set of model parameters, which are typically calibrated through the fitting of mechanical data such as stress–strain curves and elastic lattice strains. Microstructural data such as phase fractions and texture evolution are used for verifying crystal plasticity parameters. In this study, we use a multi-objective genetic algorithm to identify hardening parameters from flow stress curves with an option to incorporate texture into the optimization approach. Robust, generalized objective functions are developed and used to identify sets of parameters pertaining to dislocation density-based hardening laws in visco-plastic and elasto-plastic self-consistent (VPSC and EPSC) homogenization models. First, the parameters are identified for pure Nb directly from texture using an objective function based on generalized spherical harmonics. Since texture evolution is driven by the relative contribution of active slip systems, the parameters governing the evolution of slip resistance ratios can be recovered from fitting discrete textures at a series of strains. Next, a comprehensive set of load reversal data for dual phase (DP) 780 steel is used to fit a hardening law and a back-stress law in EPSC. Finally, parameters pertaining to a complex hardening law for the evolution of slip and twinning in pure α-Ti are identified. Remarkably, using texture as an objective in combination with stress–strain objectives constrains the model of Ti to fully reproduce not only stress–strain and texture evolution but also hierarchical twinning measurements as a function of initial grain size and texture. Furthermore, given an appropriate model fit to representative experimental texture evolution, underlying twin volume fractions contributing to texture evolution can be predicted.

42 ENGINEERING↗

Unsupervised learning of representative local atomic arrangements in molecular dynamics data

Molecular dynamics (MD) simulations present a data-mining challenge, given that they can generate a considerable amount of data but often rely on limited or biased human interpretation to examine their information content. By not asking the right questions of MD data we may miss critical information hidden within it. Here we combine dimensionality reduction (UMAP) and unsupervised hierarchical clustering (HDBSCAN) to quantitatively characterize prevalent coordination environments of chemical species within MD data. By focusing on local coordination, we significantly reduce the amount of data to be analyzed by extracting all distinct molecular formulas within a given coordination sphere. We then efficiently combine UMAP and HDBSCAN with alignment or shape-matching algorithms to partition these formulas into structural isomer families indicating their relative populations. The method was employed to reveal details of cation coordination in electrolytes based on molecular liquids.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Quantum Fast Approximate Synthesis Tool (QFAST) v1.0.0

We present QFAST, a quantum synthesis tool designed to produce short circuits and to scale well in practice. Our contributions are: 1) a novel representation of circuits able to encode placement and topology; 2) a hierarchical approach with an iterative refinement formulation that combines "coarse-grained" fast optimization during circuit structure search with a good, but slower, optimization stage only in the final circuit instantiation stage. When compared against state-of-the-art techniques, although not optimal, QFAST can generate much shorter circuits for "time dependent evolution" algorithms used by domain scientists. We also show the composability and tunability of our formulation in terms of circuit depth and running time. For example, we show how to generate shorter circuits by plugging in the best available third party synthesis algorithm at a given hierarchy level. Composability enables portability across chip architectures, which is missing from the available approaches.

Younis, Ed↗

Modeling and experimental validation of dynamical effects in Bragg coherent x-ray diffractive imaging of finite crystals

Bragg coherent diffractive imaging (BCDI) is a noninvasive microscopy technique that can visualize the shape and internal lattice deviations of crystals with nanoscale spatial resolution and picometer deformation sensitivity. Its strain imaging capability relies on Fourier transform–based iterative phase retrieval algorithms, which are mostly developed under the kinematical approximation. Such approximation prohibits the application of BCDI on larger crystals, which are commonly seen in most emerging functional materials. Understanding the dynamical effect in BCDI, as well as developing a validated method for modeling BCDI at the dynamical diffraction limit, is crucial for applying BCDI to hierarchical systems that contain micron-sized crystals and grains. Thus we report a comparative study on the impact of dynamical diffraction effects by comparing the reconstruction results from two measurements of the same crystal. Forward simulation is implemented to show subtle changes of interference fringes in the diffraction pattern due to the dynamical diffraction, and is compared directly with the experimental data.

36 MATERIALS SCIENCE↗

Block-structured, equal-workload, multi-grid-nesting interface for the Boussinesq wave model FUNWAVE-TVD (Total Variation Diminishing)

Abstract. We describe the development of a block-structured, equal-CPU-load (central processing unit), multi-grid-nesting interface for the Boussinesq wave model FUNWAVE-TVD (Fully Nonlinear Boussinesq Wave Model with Total Variation Diminishing Solver). The new model framework does not interfere with the core solver, and thus the core program, FUNWAVE-TVD, is still a standalone model used for a single grid. The nesting interface manages the time sequencing and two-way nesting processes between the parent grid and child grid with grid refinement in a hierarchical manner. Workload balance in the MPI-based (message passing interface) parallelization is handled by an equal-load scheme. A strategy of shared array allocation is applied for data management that allows for a large number of nested grids without creating additional memory allocations. Four model tests are conducted to verify the nesting algorithm with assessments of model accuracy and the robustness in the application in modeling transoceanic tsunamis and coastal effects.

Choi, Young-Kwang↗

Chemist: A Domain-Specific Language by Chemists for Chemists

Managing the complexity of quantum chemistry (QC) software is key to ensuring it remains accessible, maintainable, and reusable. Noticeably missing from the QC ecosystem are modules targeting bottleneck routines. Here we argue that this is likely due to the difficulty in defining interfaces for such modules. To that end, we introduce the open-source, publicly available Chemist library https://github.com/NWChemEx/Chemist. Chemist is a domain-specific language targeting the QC domain. Chemist has been developed focusing on performance and user-friendliness. Using Chemist, QC tasks are defined using familiar domain concepts such as molecules, wave functions, and operators. The domain objects are hierarchical to ensure a systematic encapsulation of information. Key features of Chemist include: extensibility, the ability to alias existing data, and the ability to succinctly define many common QC tasks. The usefulness of Chemist is demonstrated by discussing the interface of NWChemEx’s Fock build module and by showcasing a proof-of-concept self-consistent field algorithm containing uncertainty propagation.

Algorithms↗

Accelerating Bilevel Optimization With Hierarchical Many-Threaded Parallel Differential Evolution

Bilevel optimization is encountered in many relevant real-world applications. The main feature of this type of problem is that an upper-level optimization problem is constrained by a nested lower-level optimization problem. Because of this nested structure, bilevel problems (BLPs) are usually computationally expensive to solve. Differential evolution (DE) has demonstrated promising results in solving BLPs of relatively small scales. As the problem scale increases, the decision space becomes intrinsically larger, requiring a growing number of function evaluations for the method to work properly. In this context, heavy parallelization and high-performance computing techniques are indispensable to enable the resolution of more complex and challenging optimization problems. Hence, we propose a hierarchical many-threaded parallel DE approach for BLPs, where both levels are parallelized. The computational experiments demonstrate that the parallel implementation achieved runtime speeds ranging from 44 to 2559 times faster than the sequential version on a well-known scalable SMD benchmark test problem when executed on an NVIDIA A100 GPU. The findings indicate that the algorithm’s convergence is strongly influenced by the number of both upper- and lower-level generations. Moreover, the success of experiments with large-scale problems is closely linked to the choice of small population sizes.

Dufek, Amanda S↗

Tetrahedral Trees: A Family of Hierarchical Spatial Indexes for Tetrahedral Meshes

In this work, we address the problem of performing efficient spatial and topological queries on large tetrahedral meshes with arbitrary topology and complex boundaries. Such meshes arise in several application domains, such as 3D Geographic Information Systems (GISs), scientific visualization, and finite element analysis. To this aim, we propose Tetrahedral trees, a family of spatial indexes based on a nested space subdivision (an octree or a kD-tree) and defined by several different subdivision criteria. We provide efficient algorithms for spatial and topological queries on Tetrahedral trees and compare to state-of-the-art approaches. Our results indicate that Tetrahedral trees are an improvement over R*-trees for querying tetrahedral meshes; they are more compact, faster in many queries, and stable at variations of construction thresholds. They also support spatial queries on more general domains than topological data structures, which explicitly encode adjacency information for efficient navigation but have difficulties with domains with a non-trivial geometric or topological shape.

97 MATHEMATICS AND COMPUTING↗

kynema-fmb [SWR-23-07]

Kynema-FMB (FKA: Kynema) is an open-source performance portable flexible multibody (FMB) dynamics solver designed for time-domain simulations. While originally tailored for wind turbine structural dynamics, the formulation and implementation are those of a general flexible-multidbody dynamics solver that can readily be applied to a wide range of systems. Kynema was designed with a narrow focus, namely to provide a lightweight, fast, accurate FMD solver for coupling to computational-fluid-dynamics (CFD) codes, especially the CFD codes in the Kynema suite, for fluid-structure-interaction (FSI) simulations. Kynema-FMB is equipped to model systems that can be represented as a collection of beams and rigid bodies that are connected through constraints. Degrees of freedom are defined in the inertial/global frame of reference and include displacements and rotations (formally as rotation matrices, but stored as quaternions). The underlying formulation is built on a Lie-group time integrator designed for index-3 differential-algebraic equations, which is second-order accurate in time (Bruls et al., 2012). Beam models are based on geometrically exact beam theory and are discretized as high-order spectral finite elements similar to those in BeamDyn (Wang et al., 2017). The governing equations for a FMD system like a wind turbine constitute a highly nonlinear system of constrained partial-differential equations. Kynema-FMB uses analytical Jacobians in the nonlinear-system solves in each time step. Linear systems use sparse storage and several third-party sparse-linear-system solvers are enabled. Ill conditioning of linear systems is mitigated with preconditioning described in Bottasso et al, 2008. Kynema-FMB is integrated with a simple open-source controller (ROSCO). There is an application programming interface (API) for coupling to geometry-resolved CFD (like that in Sharma et al., 2023) and actuator-force CFD (like that in Kuhn et al., 2025). In the latter, for actuator-line models, Kynema-FMB includes an internal blade-element solver that depends on user-provided lookup tables for coefficients of lift and drag, i.e., aerodynamic polars. Kynema-FMB is written in C++ and leverages Kokkos and Kokkos-Kernels (KokkosEcosystem) as its performance portability layer enabling simulations on both CPU and GPU systems. The repository is equipped with extensive automated testing at the unit and regression/system levels. The following describes the high-level development objectives conceived for Kynema: *Kynema will follow modern software development best practices, including test-driven development (TDD), version control, hierarchical automated testing, and continuous integration (CI) for a robust development environment. *The core data structures are memory efficient and enable vectorization and parallelization at multiple levels. *Data structures are data-oriented to exploit methods for accelerated computing including high utilization of chip resources (e.g., single instruction multiple data (SIMD) instruction sets) and parallelization using GP-GPUs. *The computational algorithms incorporate robust open-source libraries for mathematical operations, resource allocation, and data management. *The API design considers multiple stakeholder needs and ensure integration with existing and future ecosystems for data science, machine learning, and AI. *Kynema-FMB is written in modern C++ and leverages Kokkos as its performance-portability library with inspiration from the kynema stack.

Sprague, MichaelA.↗

Distributed Transient Safety Verification via Robust Control Invariant Sets: A Microgrid Application

Modern safety-critical energy infrastructures are increasingly operated in a hierarchical and modular control framework which allows for limited data exchange between the modules. In this context, it is important for each module to synthesize and communicate constraints on the values of exchanged information in order to assure system-wide safety. To ensure transient safety in inverter-based microgrids, we develop a set invariance-based distributed safety verification algorithm for each inverter module. Applying Nagumo's invariance condition, we construct a robust polynomial optimization problem to jointly search for safety-admissible set of control set-points and design parameters, under allowable disturbances from neighbors. We use sum-of-squares (SOS) programming to solve the verification problem and we perform numerical simulations using grid-forming inverters to illustrate the algorithm.

Bouvier, Jean-Baptiste H.↗

U-splines: Splines over unstructured meshes

U-splines are a novel approach to the construction of a spline basis for representing smooth objects in Computer-Aided Design (CAD) and Computer-Aided Engineering (CAE). A spline is a piecewise-defined function that satisfies continuity constraints between adjacent cells in a mesh. U-splines differ from existing spline constructions, such as Non-Uniform Rational B-splines (NURBS), subdivision surfaces, T-splines, and hierarchical B-splines, in that they can accommodate local variation in cell size, polynomial degree, and smoothness simultaneously over more varied mesh configurations. Mixed cell types (e.g., triangle and quadrilateral cells in the same mesh) and T-junctions are also supported, although the continuity of interfaces with triangle and tetrahedral cells is limited in the present work. The U-spline algorithm introduces a new technique for using local null space solutions to construct basis functions for the global spline null space problem. The U-spline construction is presented for curves, surfaces, and volumes with higher dimensional generalizations possible. Lastly, a set of requirements are given to ensure that the U-spline basis is positive, forms a partition of unity, is complete, and is locally linearly independent.

42 ENGINEERING↗

Multi-Federate Co-Convergence with HELICS

In co-simulation studies, convergence refers to the ability of different simulation tools involved to achieve a consistent and stable solution. Convergence is a critical aspect of co-simulation as it determines the accuracy of the results obtained from the simulation. Convergence in co-simulation studies depends on several factors, this includes system complexity, the accuracy of the models used, and the numerical methods employed by the simulation tools. It is essential to ensure that the coupling interfaces between the different tools are designed to allow for data exchange in a consistent and accurate manner. Hierarchical Engine for Large-scale Infrastructure Co-Simulation (HELICS) is an open-source co-simulation framework developed for the energy domain. This paper explores the convergence performance of a set of co-simulation use cases. We further explore the use of a co-convergence helper federate to help with co-simulation convergence. The convergence efficacy of several algorithms (both gradient-based and gradient-free) is tested against these use cases. Finally, the sensitivity of these algorithms to several factors, such as system scaling and others, is tested and detailed in this paper. Our results show that for a subset of use cases, the co-convergence helper federate is able to improve co-simulation convergence significantly.

co-convergence↗

QBTNs - Quantum Boolean Tensor Networks

We develop algorithms and software that uses the D-Wave 2000Q quantum annealer to solve several types of Boolean tensor factorization problems. Boolean tensor factorization refers to the problem of representing a high-dimensional tensor filled with Boolean values as a product of smaller Boolean core tensors and Boolean matrices. We consider different tensor factorization models, including Boolean Tensor Train, Boolean Tucker, and Boolean Hierarchical Tucker. As an exact decomposition of a given type may not exist in the general case, the objective is to minimize the difference between the input high-dimensional tensor and the product of the lower-dimensional tensors of the proposed factorization, using a specified tensor norm. In our approach, we reduce the Boolean tensor factorization problem to a sequence of quadratic unconstrained binary optimization problems suitable for the D-Wave 2000Q quantum annealer. Although current quantum technology is still fairly restricted in the problems it can tackle, we show that complex tensor factorization problems as the ones addressed by us can be solved efficiently and accurately.

Alexandrov, Boian↗

Memory effect and phase transition in a hierarchical trap model for spin glasses

Here, we introduce an efficient dynamical tree method that enables us to explicitly demonstrate the thermoremanent magnetization memory effect in a hierarchical energy landscape. Our simulation nicely reproduces the nontrivial waiting-time and waiting-temperature dependences in this nonequilibrium phenomenon. We further investigate the condensation effect, in which a small set of microstates dominates the thermodynamic behavior in the multilayer trap model. Importantly, a structural phase transition of the multilayer tree model is shown to coincide with the onset of the condensation phenomenon. Our results underscore the importance of hierarchical structure and demonstrate the intimate relation between the glassy behavior and structure of barrier trees.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Bit-GraphBLAS: Bit-Level Optimizations of Matrix-Centric Graph Processing on GPU

In the graph data structure like adjacency matrix, the connectivity of two nodes can be sufficiently represented using only 1 bit, but they are generally treated as 32-bit full-precision in state-of-the-art graph frameworks to adopt common sparse format such as CSR. Meanwhile, bit-level parallelism has recently be explored to have high-performance potential and low storage requirement on GPUs with dense bit-tiles. To fill the gap, our solution is a hierarchical storage format that contains the bit-indexing base and dense bit-tile units. Inherently, the granularity of the bit-tile is an essential factor in achieving both storage compression and GPU parallelism. How to find a sweet spot that trades off between avoiding sparsity and exploiting is comprehensively researched in this work. In the experiment, we evaluate the proposed storage format and algorithms on modern generation GPUs, including Pascal and Volta, to figure out critical software co-designs in conjunction with existing hardware-specific optimization.

Chen, Jou-An↗