Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “fast algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Technoeconomic Design Optimization for Fast Reactors. Part I: Workflow Development and Case Study for Small LFR District Energy Application

The nuclear industry is developing small reactor designs that can target a variety of deployment locations and energy products. Smaller nuclear designs have traditionally struggled to handle the steep trade-offs between size and cost that have historically incentivized large reactors. This motivates computational optimization of small reactors to minimize costs and quantify the trade-off between size and cost. In this paper, the cost/size trade-off for a small fast reactor is derived using a multi-objective genetic algorithm optimization, with steady-state, transient, and cost analysis of the fast reactor being performed. Specifically, the method is demonstrated on a small 10- to 120-MW(thermal) U-Pu-Zr–fueled lead-cooled fast reactor with a 10-year core life for district energy applications, which can have a thermal load compatible with this range. The results reinforced that fast reactor cores at the lower end of this power range suffer cost penalties due to critical mass considerations. It was found that high power density cores with strong reactivity swings and many control rods were favored over designing to minimize reactivity swing. Furthermore, this contrasts with some traditional configurations designed using engineering judgment and demonstrates that optimizers can find nontraditional but realistic solutions, along with demonstrating the value of incorporating cost functions into whole-reactor design optimization.

Fast reactor↗

Optimizing High Performance Markov Clustering for Pre-Exascale Architectures

HipMCL is a high-performance distributed memory implementation of the popular Markov Cluster Algorithm (MCL) and can cluster large-scale networks within hours using a few thousand CPU-equipped nodes. It relies on sparse matrix computations and heavily makes use of the sparse matrix-sparse matrix multiplication kernel (SpGEMM). The existing parallel algorithms in HipMCL are not scalable to Exascale architectures, both due to their communication costs dominating the runtime at large concurrencies and also due to their inability to take advantage of accelerators that are increasingly popular. In this work, we systematically remove scalability and performance bottlenecks of HipMCL. We enable GPUs by performing the expensive expansion phase of the MCL algorithm on GPU. Additionally, we propose a CPU-GPU joint distributed SpGEMM algorithm called pipelined Sparse SUMMA and integrate a probabilistic memory requirement estimator that is fast and accurate. Furthermore, we develop a new merging algorithm for the incremental processing of partial results produced by the GPUs, which improves the overlap efficiency and the peak memory usage. We also integrate a recent and faster algorithm for performing SpGEMM on CPUs. We validate our new algorithms and optimizations with extensive evaluations. With the enabling of the GPUs and integration of new algorithms, HipMCL is up to 12.4x faster, being able to cluster a network with 70 million proteins and 68 billion connections just under 15 minutes using 1024 nodes of ORNL's Summit supercomputer.

97 MATHEMATICS AND COMPUTING↗

Integral Kernel Methods for Nonlinear Parabolic-Elliptic Systems

Nonlinear parabolic-elliptic systems arise in many physical, biological, and chemical phenomena such as chemotaxis, ion transport, self-gravitating particles, and Brownian vortices. Existing methods struggle with the strong coupling and high nonlinearity and nonlocality of some of these systems, especially the ill-conditioned, convection-dominated problems. To overcome numerical difficulties, current approaches rely on initial guesses, preconditioning, or iterative techniques with no convergence guarantees. They might suffer from poor scalability, large memory usage, and difficulty to parallelize. Inspired by the connection of parabolic-elliptic systems to stochastic processes, we introduce a novel meshless, monolithic, and fully explicit method that naturally encapsulates the elliptic and parabolic operators into a single step which updates each node deterministically with global information. By being fully quadrature-based, it avoids solving systems of discretized equations and does not utilize initial guesses or preconditioning, while requiring little memory and being easy to parallelize. We first derive the method in an integral kernel formulation with quadratic complexity in the number of integration nodes and then leverage kernel-independent fast multipole methods (FMM) to present a scalable algorithm with linear complexity. We provide numerical examples for the Poisson-Nernst-Planck equations in one, two, and three dimensions, together with the derivation of the integral kernel for each case. Furthermore, the examples demonstrate the fast convergence and scalability of the FMM-accelerated algorithm, as well as its suitability for convection-dominated problems, making it competitive against traditional PDE solvers.

PDE systems↗

Examination of Semi-Analytical Solution Methods in the Coarse Operator of Parareal Algorithm for Power System Simulation

With continuing advances in high-performance parallel computing platforms, parallel algorithms have become powerful tools for development of faster than real-time power system dynamic simulations. In particular, it has been demonstrated in recent years that parallel-in-time (Parareal) algorithms have the potential to achieve such an ambitious goal. Here, the selection of a fast and reasonably accurate coarse operator of the Parareal algorithm is crucial for its effective utilization and performance. This paper examines semi-analytical solution (SAS) methods as the coarse operators of the Parareal algorithm and explores performance of the SAS methods to the standard numerical time integration methods. Two promising time-power series-based SAS methods were considered; Adomian decomposition method and Homotopy analysis method with a windowing approach for improving the convergence. Numerical performance case studies on 10-generator 39-bus system and 327-generator 2383-bus system were performed for these coarse operators over different disturbances, evaluating the number of Parareal iterations, computational time, and stability of convergence. All the coarse operators tested with different scenarios have converged to the same corresponding true solution (if they are convergent) and the SAS methods provide comparable computational speed, while having more stable convergence to the true solution in many cases.

97 MATHEMATICS AND COMPUTING↗

Implicit–explicit multirate infinitesimal stage-restart methods

Implicit–Explicit (IMEX) methods are flexible numerical time integration methods which solve an initial-value problem (IVP) that is split into stiff and nonstiff processes with the goal of lower computational costs than a purely implicit or explicit approach. A complementary form of flexible IVP solvers are multirate infinitesimal methods for problems split into fast- and slow-changing dynamics, that solve a multirate IVP by evolving a sequence of “fast” IVPs using any suitably accurate algorithm. This article introduces a new class of high-order implicit–explicit multirate methods that are designed for multirate IVPs in which the slow-changing dynamics are further split in an IMEX fashion. This new class, which we call implicit–explicit multirate infinitesimal stage-restart (IMEX-MRI-SR), both improves upon the previous implicit–explicit multirate infinitesimal generalized-structure additive Runge Kutta (IMEX-MRI-GARK) methods by allowing for far easier creation of new embedded methods, and extends multirate exponential Runge Kutta (MERK) methods by allowing the fast-changing dynamics to be nonlinear and the methods to be implicit. We leverage GARK theory to derive conditions for orders of accuracy up to four, and we provide second- and third-order accurate example methods, which are the first known embedded MRI methods with IMEX structure. We then perform numerical simulations demonstrating convergence rates and computational performance in both fixed-step and adaptive-step settings.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Computation-Efficient Algorithm for Distributed Feedback Optimization of Distribution Grids

Feedback-based optimization algorithms use real-time measurements to update the optimal control for the underlying system which may not be fully identified. Recently, we have developed a distributed feedback-based algorithm [1] that avoids the requirement of fast communication between central computing and local actuator/sensor agents. This paper extends the work by greatly reducing the number of copies of variables involved in the distributed feedback-based algorithm, which results in faster convergence and lower communication requirement. The main idea is to leverage the specific structural properties of the admittance matrix for distribution systems with tree network topology. We also show the effectiveness of the proposed algorithm in simulations.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Exact and locally implicit source term solvers for multifluid-Maxwell systems

Recently, a family of models that couple multifluid systems to the full Maxwell equations have been used in laboratory, space, and astrophysical plasma modeling. These models are more complete descriptions of the plasma than reduced models like magnetohydrodynamic (MHD) since they are derived more closely from the full kinetic Vlasov-Maxwell system, without assumptions like quasi-neutrality, negligible electron mass, etc. Thus these models naturally retain non-ideal MHD effects like electron inertia, Hall term, pressure anisotropy/nongyrotropy, displacement current, among others. One obstacle to broader application of these model is that an explicit treatment of their source terms leads to the need to resolve rapid processes like plasma oscillation and electron cyclotron motion, even when these are not important. In this paper, we suggest two ways to address this issue. First, we derive the analytic solutions to the source update equations, which can be implemented as a practical, but less generic solver. We then develop a time-centered, locally implicit algorithm to update the source terms, allowing stepping over the fast kinetic time-scales. For a plasma with S species, the locally implicit algorithm involves inverting a local (3 S + 3) × (3 S + 3) matrix only, thus is very efficient. The performance can be further increased by using the direct update formulas to skip null calculations. In this paper, we present benchmarks illustrating the exact energy-conservation of the locally implicit solver, as well as its efficiency and robustness for both small-scale, idealized problems and largescale, complex systems. The locally implicit algorithm can be also easily extended to include other local sources, like collisions and ionization, which are difficult to solve analytically.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

A Butterfly-Accelerated Volume Integral Equation Solver for Broad Permittivity and Large-Scale Electromagnetic Analysis

In this work, a butterfly-accelerated volume integral equation (VIE) solver is proposed for fast and accurate electromagnetic (EM) analysis of scattering from heterogeneous objects. The proposed solver leverages the hierarchical off-diagonal butterfly (HOD-BF) scheme to construct the system matrix and obtain its approximate inverse, used as a preconditioner. Complexity analysis and numerical experiments validate the O(N log 2 N) construction cost of the HOD-BF-compressed system matrix and O(N log 1.5 N) inversion cost for the preconditioner, where N is the number of unknowns in the high-frequency EM scattering problem. For many practical scenarios, the proposed VIE solver requires less memory and computational time to construct the system matrix and obtain its approximate inverse compared to a H matrix-accelerated VIE solver. The accuracy and efficiency of the proposed solver have been demonstrated via its application to the EM analysis of large-scale canonical and real-world structures comprising of broad permittivity values and involving millions of unknowns.

42 ENGINEERING↗

An Online Joint Optimization–Estimation Architecture for Distribution Networks

Here in this article, we propose an optimal joint optimization-estimation architecture for distribution networks, which jointly solves the optimal power flow (OPF) problem and static state estimation (SE) problem through an online gradient-based feedback algorithm. The main objective is to enable a fast and timely interaction between the OPF decisions and state estimators with limited sensor measurements. First, convergence and optimality of the proposed algorithm are analytically established. Then, the proposed gradient-based algorithm is modified by introducing statistical information of the inherent estimation and linearization errors for an improved and robust performance of the online OPF decisions. Overall, the proposed method eliminates the traditional separation of operation and monitoring, where optimization and estimation usually operate at distinct layers and different time scales. Hence, it enables a computationally affordable, efficient, and robust online operational framework for distribution networks under time-varying settings.

24 POWER TRANSMISSION AND DISTRIBUTION↗

A Robust Numerical Treatment of Solid-Phase Diffusion in Pseudo Two-Dimensional Lithium-Ion Battery Models

Solid-phase diffusion in active materials of lithium-ion batteries significantly affects charging and safety-related behavior of lithium-ion batteries. Therefore, it is essential to develop an efficient and robust numerical algorithm for solving solid-phase diffusion equations in physics-based battery models. In this work, we discuss the origins of numerical instabilities that can occur when solving the solid-phase diffusion equations using iterative methods. Then, in order to resolve such issues, we propose a simple numerical treatment to the surface flux term of discretized solid-phase diffusion equations. To demonstrate its numerical robustness, the proposed method is implemented into a pseudo two-dimensional (P2D) physics-based battery model and simulations are conducted at wide ranges of operating conditions. Even with extremely poor initial guesses for the Li+ concentrations of the active materials, computations using the proposed method do not diverge and the their computational speeds are comparable to those with conventional initial guesses. Comprehensive tests of the proposed method are also performed with a dynamic current profile based on US06 driving profile and a multi-stage charging profile with very high initial C-rate (12C).

battery modeling↗

Real-Time Distributed Control of Smart Inverters for Network-level Optimization

The limitations of centralized optimization methods in managing electric power distribution systems operations have led to the distributed paradigm of computing and decision-making. Unfortunately, the existing distributed optimization algorithms are limited in their applicability to managing fast varying phenomena such as those resulting from highly variable Distributed Energy Resource (DER) generation patterns. They require a large number of communication rounds (in the order of 10 2 to 10 3 ) among the computing agents to solve one instance of the optimization problem. Related real-time distributed control methods are equally limited in their applications to power distribution systems with fast-changing DER generation; they require hundreds of rounds of communication and thus are slow in tracking the network-level optimal solutions. In this paper, we propose a novel distributed voltage controller that provides a fast-tracking of rapidly varying DER generation profiles while simultaneously converging to network-level optimal solutions within a few communication rounds. The proposed control algorithm leverages the radial topology of the system, which reduces the required communication rounds to reach the network-level optimum solution by order of magnitude. The novelty lies in carefully reducing the electrical network model from the perspective of each distributed controller and enabling appropriate data sharing among upstream and downstream nodes to achieve fast convergence. The simulation results demonstrate the effectiveness of the proposed approach in minimizing the feeder losses while maintaining the node voltage within the pre-specified limits.

voltage control, optimization, reactive power, inv↗

Verification of the DIF3D Software to Support Fast Reactor Analysis

Ongoing design activities at Argonne National Laboratory are requiring a thorough verification of the Argonne Reactor Computation codes be performed. DIF3D is central to this system. The driver for this effort requires the 3D Cartesian, triangular-Z, and hexagonal-Z core geometry options of DIF3D be verified. Previous work identified the DIF3D features required to be verified to support current design activities, features of which are generally applicable to hexagonal-Z fast reactor designs. The scope of this verification effort includes verifying DIF3D’s ability to correctly translate the user’s model in to DIF3D’s preferred format, verifying that options planned for use have the desired effect, and verifying the correctness of the eigenvalue, fixed-source, forward, and adjoint solvers in DIF3D-FD and DIF3D-VARIANT. This manuscript provides the verification tasks and their results with respect to the features needed for current design activities. Since analytic solutions of the neutron diffusion and transport equations are either limited in scope or not possible, multiple tiers of problems unique to each solver and geometry type were implemented. Each of these tiers tests features independent and complementary arguments for why the separate testing of functionalities is acceptable. Finally, this separate testing was also supplemented with a high-level integral check of each the diffusion and transport capabilities and applicable geometries. To accommodate cases which an analytic solution is not feasible, MCNP6.2 was relied upon to provide a higher-order reference solution. This therefore required that the capabilities within MCNP6.2 which were relied upon for this work are also verified in this work. No MCNP discrepancies were noted in this effort. Note that the MCNP6.2 verification included in this work does not stand as a full verification of MCNP6.2, but merely verifies the features used in verifying DIF3D. The verification effort identified no issues that are debilitating or otherwise impactful to design usage of DIF3D, and thus DIF3D version 11.0, release 3012 is considered verified. As some additional changes have been made to the ARC software since this point all versions between release 3012 and 3266 can be considered verified as version 3253 was used for all updates in this revision. The types of issues that were identified were predominantly in the areas of: unclear documentation, software bugs which were inconsequential to final results, editing options which were ignored in favor of printing more information than requested, bugs in the outputs of intermediate results, or secondary output binary file information which was not present. While not a bug, this verification report also identified that the algorithm used to evaluate the peak fast flux in a nodal transport solution can be quite unreliable due to the methodology used and the location of the peak within the mesh. The authors of the report therefore recommend the usage of the EvaluateFlux software (distributed with ARC) as a more robust alternative noting that DIF3D will properly notify the user when the peaking values it is providing are potentially incorrect.

97 MATHEMATICS AND COMPUTING↗

Verification of the DIF3D Software to Support Fast Reactor Analysis (Rev. 3)

Ongoing design activities at Argonne National Laboratory are requiring a thorough verification of the Argonne Reactor Computation codes be performed. DIF3D is central to this system. The driver for this effort requires the 3D Cartesian, triangular-Z, and hexagonal-Z core geometry options of DIF3D be verified. Previous work identified the DIF3D features required to be verified to support current design activities, features of which are generally applicable to hexagonal-Z fast reactor designs. The scope of this verification effort includes verifying DIF3D’s ability to correctly translate the user’s model in to DIF3D’s preferred format, verifying that options planned for use have the desired effect, and verifying the correctness of the eigenvalue, fixed-source, forward, and adjoint solvers in DIF3D-FD and DIF3D-VARIANT. This manuscript provides the verification tasks and their results with respect to the features needed for current design activities. Since analytic solutions of the neutron diffusion and transport equations are either limited in scope or not possible, multiple tiers of problems unique to each solver and geometry type were implemented. Each of these tiers tests features independent and complementary arguments for why the separate testing of functionalities is acceptable. Finally, this separate testing was also supplemented with a high-level integral check of each the diffusion and transport capabilities and applicable geometries. To accommodate cases which an analytic solution is not feasible, MCNP6.2 was relied upon to provide a higher-order reference solution. This therefore required that the capabilities within MCNP6.2 which were relied upon for this work are also verified in this work. No MCNP discrepancies were noted in this effort. Note that the MCNP6.2 verification included in this work does not stand as a full verification of MCNP6.2, but merely verifies the features used in verifying DIF3D. The verification effort identified no issues that are debilitating or otherwise impactful to design usage of DIF3D, and thus DIF3D version 11.0, release 3012 is considered verified. As some additional changes have been made to the ARC software since this point all versions between release 3012 and 3266 can be considered verified as version 3253 was used for all updates in this revision. The types of issues that were identified were predominantly in the areas of: unclear documentation, software bugs which were inconsequential to final results, editing options which were ignored in favor of printing more information than requested, bugs in the outputs of intermediate results, or secondary output binary file information which was not present. While not a bug, this verification report also identified that the algorithm used to evaluate the peak fast flux in a nodal transport solution can be quite unreliable due to the methodology used and the location of the peak within the mesh. The authors of the report therefore recommend the usage of the EvaluateFlux software (distributed with ARC) as a more robust alternative noting that DIF3D will properly notify the user when the peaking values it is providing are potentially incorrect.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Long-time simulations for fixed input states on quantum hardware

Publicly accessible quantum computers open the exciting possibility of experimental dynamical quantum simulations. While rapidly improving, current devices have short coherence times, restricting the viable circuit depth. Despite these limitations, we demonstrate long-time, high fidelity simulations on current hardware. Specifically, we simulate an XY-model spin chain on Rigetti and IBM quantum computers, maintaining a fidelity over 0.9 for 150 times longer than is possible using the iterated Trotter method. Our simulations use an algorithm we call fixed state Variational Fast Forwarding (fsVFF). Recent work has shown an approximate diagonalization of a short time evolution unitary allows a fixed-depth simulation. fsVFF substantially reduces the required resources by only diagonalizing the energy subspace spanned by the initial state, rather than over the total Hilbert space. We further demonstrate the viability of fsVFF through large numerical simulations, and provide an analysis of the noise resilience and scaling of simulation errors.

97 MATHEMATICS AND COMPUTING↗

VoroClust

SAND2025-11465O VoroClust, also known as Voronoi Clustering, is a fast, density-based unsupervised clustering algorithm applicable to high-resolution and high-dimensional data. It operates as quickly as distance-based clustering methods while effectively capturing complex regional geometries, matching the performance of current density-based methods. VoroClust employs a data-centered sphere cover to reduce computational demands while preserving data topology. It propagates clusters outward from local density peaks. Although supervised machine learning is powerful for applications like image classification and segmentation, it requires comprehensive, consistent datasets, which many applications lack. Unsupervised clustering algorithms analyze the structure of each dataset rather than relying on similarities with other examples, making them well-suited for practical applications with insufficient or inappropriate data for supervised learning. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Ebeida, Mohamed [Sandia National Lab. (SNL-CA), Li↗

PhaseGAN: a deep-learning phase-retrieval approach for unpaired datasets

Phase retrieval approaches based on deep learning (DL) provide a framework to obtain phase information from an intensity hologram or diffraction pattern in a robust manner and in real-time. However, current DL architectures applied to the phase problem rely on i) paired datasets, i. e., they arc only applicable when a satisfactory solution of the phase problem has been found, and ii) the fact that most of them ignore the physics of the imaging process. Here, we present PhaseGAN, a new DL approach based on Generative Adversarial Networks, which allows the use of unpaired datasets and includes the physics of image formation. The performance of our approach is enhanced by including the image formation physics and a novel Fourier loss function, providing phase reconstructions when conventional phase retrieval algorithms fail, such as ultra-fast experiments. Thus, PhaseGAN offers the opportunity to address the phase problem in real-time when no phase reconstructions but good simulations or data from other experiments are available.

47 OTHER INSTRUMENTATION↗

Longitudinal Phase Space Tomography for the Booster Synchrotron

Efforts in the study of the longitudinal behavior of charged particles in the Fermilab Booster can be catalyzed with an image of the two-dimensional phase space distribution. In the past, tomography has been extensively employed in the reconstruction of the phase space in accelerators such as the Recycler at Fermilab and the Proton Synchrotron Booster at CERN. However, such a capability had yet to realize for the Fermilab Booster synchrotron. In this work, the first successful tomographic phase space reconstruction of a low-energy Booster bunch is presented along with validation metrics. A numerical turn-by-turn model of the longitudinal particle dynamics in the Booster has been implemented, which utilizes a fast, map-based particle transport algorithm. Using a sinogram generated from the Wall Current Monitor, the iterative reconstruction algorithm recovers a discretized image of the original phase space distribution at variable resolution. The reconstruction result shows low root-mean-square error and a rapid convergence toward the solution, providing strong evidence of accuracy. Future and ongoing work includes modeling high-energy bunches above transition and using tomography to infer certain machine parameters such as synchronous phase, peak gap voltage, and synchronous energy in addition to the phase space distribution.

Ebeid, Safi [Unlisted]↗