Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “fast algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Performance of modern color decompositions for standard candle LHC tree amplitudes

In the last decade, developments of matrix element and phase space generators have focused on providing good efficiency and maximal flexibility and automation for a wide range of physical processes. However, as recent studies have shown, they are a major bottleneck in the established Monte Carlo event generator toolchains. With the advent of the HL-LHC and ever rising precision requirements, future developments will need to focus on computational performance, especially at intermediate to large jet multiplicities. We present the novel BlockGen family of fast matrix element algorithms that are amenable for GPU acceleration, making use of modern, minimal color decompositions. Moreover, we discuss the performance achieved for standard candle processes such as V +jets and tt̄+jets production.

Bothmann, E. [Gottingen U.]↗

Reconstruction of signal amplitudes in the CMS electromagnetic calorimeter in the presence of overlapping proton-proton interactions

A template fitting technique for reconstructing the amplitude of signals produced by the lead tungstate crystals of the CMS electromagnetic calorimeter is described. This novel approach is designed to suppress the contribution to the signal of the increased number of out-of-time interactions per beam crossing following the reduction of the accelerator bunch spacing from 50 to 25 ns at the start of Run 2 of the LHC. Execution of the algorithm is sufficiently fast for it to be employed in the CMS high-level trigger. It is also used in the offline event reconstruction. Results obtained from simulations and from Run 2 collision data (2015–2018) demonstrate a substantial improvement in the energy resolution of the calorimeter over a range of energies extending from a few GeV to several tens of GeV.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Exact and fast calculation of the X-ray pair distribution function

A fast and exact algorithm to calculate the powder pair distribution function (PDF) for the case of periodic structures is presented. The new algorithm calculates the PDF by a detour via reciprocal space. The calculated normalized total powder diffraction pattern is transferred into the PDF via the sine Fourier transform. The calculation of the PDF via the powder pattern avoids the conventional simplification of X-ray and electron atomic form factors. It is thus exact for these types of radiation, as is the conventional calculation for the case of neutron diffraction. The new algorithm further improves the calculation speed. Additional advantages are the improved detection of errors in the primary data, the handling of preferred orientation, the ease of treatment of magnetic scattering and a large improvement to accommodate more complex instrumental resolution functions.

36 MATERIALS SCIENCE↗

Rolling Root Mean Square Based Multimodal Anomaly Detection for Real Time Monitoring of Smart Grid

Reliable real-time monitoring is valuable for maintaining the operational integrity of modern electrical smart grids. Deployment of heterogeneous sensing technologies in substations has enabled high-resolution, multichannel waveform monitoring, but also introduces challenges for anomaly detection due to noise, baseline drift, and modality-dependent signal characteristics. In this work, we present a computationally efficient unsupervised method for multimodal event detection based on Rolling Root Mean Square based Event Detection (RRMSED). The method is developed using in-house, field deployed sensors collecting data at a utility substation. The sensing system comprises voltage and current sensors, triaxial accelerometers, and magnetometers, collectively capturing electrical, vibrational, and magnetic waveform measurements at high temporal resolution. RRMSED operates by extracting rolling RMS energy features and their first-order temporal differences from consecutive waveform segments for each channel and then applying channel-specific statistical thresholds learned from historical data. A persistence-based exceedance logic is employed to robustly identify transient events while suppressing impulsive noise, and to provide precise temporal localization with high resolution. The framework is designed for continuous server-side operation and can be deployed in real time without requiring complex models. Experiments on simulated waveform data with known ground truth demonstrate low false positive (FP) and false negative (FN) rates. Application to real substation data shows RRMSED to identify events that are not captured by conventional monitoring indicators including fast transient detection algorithm currently deployed in the system. These results indicate that rolling RMS based features provide an effective and practical basis for real-time multimodal event detection in smart-grid substations.

Mukherjee, Subrata [ORNL] (ORCID:0000000309930338)↗

Computationally Efficient Decompositions of Oblique Projection Matrices

Oblique projection matrices arise in problems in weighted least squares, signal processing, and optimization. While these matrices can be potentially very large, their low-rank structure can be exploited for efficient computation. Here, we propose fast and scalable algorithms for computing their eigendecomposition and singular value decomposition (SVD). Numerical experiments that compare our proposed approaches to existing methods, including randomized SVD, are presented. In addition, we test their accuracy on linear systems from equality constrained optimization problems.

97 MATHEMATICS AND COMPUTING↗

Can changes in deformation regimes be inferred from crystallographic preferred orientations in polar ice?

Creep due to ice flow is generally thought to be the main cause for the formation of crystallographic preferred orientations (CPOs) in polycrystalline anisotropic ice. However, linking the development of CPOs to the ice flow history requires a proper understanding of the ice aggregate's microstructural response to flow transitions. In this contribution the influence of ice deformation history on the CPO development is investigated by means of full-field numerical simulations at the microscale. We simulate the CPO evolution of polycrystalline ice under combinations of two consecutive deformation events up to high strain, using the code VPFFT (visco-plastic fast Fourier transform algorithm) within ELLE. A volume of ice is first deformed under coaxial boundary conditions, which results in a CPO. The sample is then subjected to different boundary conditions (coaxial or non-coaxial) in order to observe how the deformation regime switch impacts the CPO. The model results indicate that the second flow event tends to destroy the first, inherited fabric with a range of transitional fabrics. However, the transition is slow when crystallographic axes are critically oriented with respect to the second imposed regime. Therefore, interpretations of past deformation events from observed CPOs must be carried out with caution, particularly in areas with complex deformation histories.

54 ENVIRONMENTAL SCIENCES↗

Power and Limitations of Linear Programming Decoder for Quantum LDPC Codes

Decoding quantum error-correcting codes is a key challenge in enabling fault-tolerant quantum computation. In the classical setting, linear programming (LP) decoders offer provable performance guarantees and can leverage fast practical optimization algorithms. Although LP decoders have been proposed for quantum codes, their performance and limitations remain relatively underexplored. In this work, we uncover a key limitation of LP decoding for quantum low-density parity-check (LDPC) codes: certain constant-weight error patterns lead to ambiguous fractional solutions that cannot be resolved through independent rounding. To address this issue, we incorporate a post-processing technique known as ordered statistics decoding (OSD), which significantly enhances LP decoding performance in practice. Our results show that LP decoding, when augmented with OSD, can outperform belief propagation with the same post-processing for intermediate code sizes of up to hundreds of qubits. These findings suggest that LP-based decoders, equipped with effective post-processing, offer a promising approach for decoding near-term quantum LDPC codes.

Gu, Shouzhen [Yale U.]↗

Power and Limitations of Linear Programming Decoder for Quantum LDPC Codes

Decoding quantum error-correcting codes is a key challenge in enabling fault-tolerant quantum computation. In the classical setting, linear programming (LP) decoders offer provable performance guarantees and can leverage fast practical optimization algorithms. Although LP decoders have been proposed for quantum codes, their performance and limitations remain relatively underexplored. In this work, we uncover a key limitation of LP decoding for quantum low-density parity-check (LDPC) codes: certain constant-weight error patterns lead to ambiguous fractional solutions that cannot be resolved through independent rounding. To address this issue, we incorporate a post-processing technique known as ordered statistics decoding (OSD), which significantly enhances LP decoding performance in practice. Our results show that LP decoding, when augmented with OSD, can outperform belief propagation with the same post-processing for intermediate code sizes of up to hundreds of qubits. These findings suggest that LP-based decoders, equipped with effective post-processing, offer a promising approach for decoding near-term quantum LDPC codes.

Gu, Shouzhen [Yale U.]↗

Power and Limitations of Linear Programming Decoder for Quantum LDPC Codes

Decoding quantum error-correcting codes is a key challenge in enabling fault-tolerant quantum computation. In the classical setting, linear programming (LP) decoders offer provable performance guarantees and can leverage fast practical optimization algorithms. Although LP decoders have been proposed for quantum codes, their performance and limitations remain relatively underexplored. In this work, we uncover a key limitation of LP decoding for quantum low-density parity-check (LDPC) codes: certain constant-weight error patterns lead to ambiguous fractional solutions that cannot be resolved through independent rounding. To address this issue, we incorporate a post-processing technique known as ordered statistics decoding (OSD), which significantly enhances LP decoding performance in practice. Our results show that LP decoding, when augmented with OSD, can outperform belief propagation with the same post-processing for intermediate code sizes of up to hundreds of qubits. These findings suggest that LP-based decoders, equipped with effective post-processing, offer a promising approach for decoding near-term quantum LDPC codes.

Gu, Shouzhen [Yale U.]↗

Optimizing Fast Charging and Wetting in Lithium-Ion Batteries with Optimal Microstructure Patterns Identified by Genetic Algorithm

To sustain the high-rate current required for fast charging electric vehicle batteries, electrodes must exhibit sufficiently high effective ionic diffusion. Additionally, to reduce battery manufacturing costs, wetting time must decrease. Both of these issues can be addressed by structuring the electrodes with mesoscale pore channels. However, their optimal spatial distribution, or patterns, is unknown. Herein, a genetic algorithm has been developed to identify these optimal patterns using a CPU-cheap proxy distance-based model to evaluate the impact of the added pore networks. Both coin-cell and pouch cell form factors have been considered for the wetting analysis, with their respective electrolyte infiltration mode. Regular hexagonal and mud-crack-like patterns, respectively, for fast charging and fast wetting were found to be optimal and have been compared with pre-determined, easier to manufacture, patterns. The model predicts that using cylindrical channels arranged in a regular hexagonal pattern is ∼6.25 times more efficient for fast charging as compared to grooved lines with both structuring strategies being restricted to a 5% electrode total volume loss. The model also shows that only a very limited electrode volume loss (1%–2%) is required to dramatically improve the wetting (5–20 times) compared to an unstructured electrode.

25 ENERGY STORAGE↗

Electromagnetic Transient (EMT) Simulation Algorithms for Evaluation of Large-Scale Extreme Fast Charging Systems (T&D Models)

Simulation of high-fidelity models of extreme fast charging (XFC) systems and large-area power grids with many XFCs can be time consuming in traditional simulators. Traditional simulators use a single method of discretization for all the components that results in imposing a large computational burden of inverting a large matrix as well as increased computations related to single method of discretization (that is typically a trapezoidal method). To overcome the problem of simulating large-area power grids with many XFCs, in this paper, advanced numerical simulation algorithms are applied for the first time together to reduce the dimension of matrix inversion. Here, the algorithms include numerical stiffness-based segregation, time constant-based segregation, clustering and aggregation on differential algebraic equations (DAEs), and multi-order integration approaches. These algorithms apply multiple discretization algorithms rather than a single discretization algorithm that further reduces the computational burden. The approaches mentioned here have resulted in speed-up of up to 18x in the simulation of a single distribution system with 15 XFCs and of up to 271x in the simulation of a transmission-distribution system with 300 XFCs in multiple distribution feeders with respect to conventional simulators (like power systems computer aided design [PSCAD]).

42 ENGINEERING↗

Online Data-Enabled Predictive Control

We develop an online data-enabled predictive (ODeePC) control method for trajectory tracking of unknown systems, building upon the recently proposed DeePC. Our proposed ODeePC method leverages a primal-dual algorithm with real-time measurement feedback to iteratively compute the corresponding real-time optimal control policy as system conditions change. Specifically, our developed ODeePC: a) records data from the unknown system and updates the underlying primal-dual algorithm dynamically, b) can track changes in the system's operating point and adjust the control inputs, and c) is computationally efficient as it deploys a Fast Fourier Transform-based algorithm enabling the fast computation of the product of a non-square Hankel matrix with a vector. We provide theoretical guarantees regarding the asymptotic behavior of ODeePC and demonstrate its performance through a power system application.

61 RADIATION PROTECTION AND DOSIMETRY↗

Electromagnetic Transient Simulation Algorithms for Evaluation of Large-Scale Extreme Fast Charging Systems (Distribution Grid Models)

The distribution and transmission grids are observing an increased penetration of power electronics in loads and generations. For example, there is increasing interest in integrating in extreme fast charging (XFC) systems for fast charging of electrical vehicles. As these systems are integrated, developing high-fidelity electromagnetic transient model of XFC systems in distribution grids and evaluating their interactions with the power grid would be of significant interest. This model will be utilized for design of XFC systems, to identify upgrades in distribution and/or transmission grids, for planning purposes by transmission planners or operators or owners, among others. It can also be utilized in operations for improved reliable performance of the grid and/or XFC station. The challenge with simulating these models is the high computational complexity introduced by the large number of states present in the system and the time-step needed to simulate the system. In this paper, advanced simulations algorithms are applied to reduce the computational complexity of simulating large-scale XFC systems. The algorithms include numerical stiffness-based segregation, time constant-based segregation, clustering and aggregation on differential algebraic equations (DAEs), and multi-order integration approaches. While the first three algorithms split the matrix that needs to be inverted from a large matrix to much smaller matrices, the final algorithm reduces the computational burden of applying higher-order integration approaches in the complete system. The comparison made in the previous sentence is with respect to use of homogeneous integration approaches used in conventional electromagnetic transient simulators like power systems computer aided design (PSCAD). The approaches mentioned here have resulted in speed-up of 36x in the simulation of a single distribution system with 15 XFCs.

Debnath, Suman↗

Fast shared-memory streaming multilevel graph partitioning

In this report we show that a fast parallel graph partitioner can benefit many applications by reducing data transfers. The online methods for partitioning graphs have to be fast and they often rely on simple one-pass streaming algorithms, while the offline methods for partitioning graphs contain more involved algorithms and the most successful methods in this category belong to the multilevel approaches. In this work, we assess the feasibility of using streaming graph partitioning algorithms within the multilevel framework. Our end goal is to come up with a fast parallel offline multilevel partitioner that can produce competitive cutsize quality. We rely on a simple but fast and flexible streaming algorithm throughout the entire multilevel framework. This streaming algorithm serves multiple purposes in the partitioning process: a clustering algorithm in the coarsening, an effective algorithm for the initial partitioning, and a fast refinement algorithm in the uncoarsening. Its simple nature also lends itself easily for parallelization. The experiments on various graphs show that our approach is on the average up to 5.1x faster than the multi-threaded MeTiS, which comes at the expense of only 2x worse cutsize.

97 MATHEMATICS AND COMPUTING↗

Spectrally accurate, reverse-mode differentiable bounce-averaging algorithm and its applications

We present a fast, spectrally (exponentially) accurate, automatically differentiable bounce-averaging algorithm that is used to simplify kinetic models. Using this algorithm, implemented in the DESC stellarator optimisation suite, we can perform efficient optimisation of many objectives to improve stellarator performance, such as the effective ripple 𝜖 eff metric for the neoclassical transport coefficient in the low collisionality regime and proxies for energetic particle confinement. For the first time, we optimise a finite-beta stellarator to directly reduce neoclassical ripple transport using reverse-mode differentiation. This ensures the computational cost of differentiation is independent of the number of controllable parameters.

fusion plasma↗

Fast BLT Code

This note discusses numerical algorithmic software design considerations and performance estimates for a fast BLT coupling code [1] written in c++. The aim of this code is to conduct faster parameter studies over line orientations. The original matlab code was written by Mike Rivera. Art Barnes ported this code to julia. I rewrote portions of the code for speed improvement, mainly to eliminate some redundant computation when calculating many line orientations. But this code is still far from optimal.

97 MATHEMATICS AND COMPUTING↗

Fast and Scalable FFT-Based GPU-Accelerated Algorithms for Block-Triangular Toeplitz Matrices with Application to Linear Inverse Problems Governed by Autonomous Dynamical Systems

In this work, we present an efficient and scalable algorithm for performing matrix-vector multiplications (matvecs) for block Toeplitz matrices. Such matrices, which are shift-invariant with respect to their blocks, arise in the context of solving inverse problems governed by autonomous systems, and time-invariant systems in particular. In this article, we consider inverse problems that infer unknown parameters from observational data of a linear time-invariant dynamical system given in the form of partial differential equations (PDEs). Matrix-free Newton-conjugate-gradient methods are often the gold standard for solving these inverse problems, but they require numerous actions of the Hessian on a vector. Matrix-free adjoint-based Hessian matvecs require solution of a pair of linearized forward/adjoint PDE solves per Hessian action, which may be prohibitive for large-scale inverse problems. Time invariance of the forward PDE problem leads to a block Toeplitz structure of the discretized parameter-to-observable (p2o) map defining the mapping from inputs (parameters) to outputs (observables) of the PDEs. This block Toeplitz structure enables us to exploit two key properties: (1) compact storage of the p2o map and its adjoint, and (2) efficient fast Fourier transform–based Hessian matvecs. The proposed algorithm is mapped onto large multi-GPU clusters and achieves more than 80% of peak bandwidth on NVIDIA A100 GPUs. Excellent weak scaling is shown for up to 48 A100 GPUs. For the targeted problems, the implementation executes Hessian matvecs within fractions of a second, which is orders of magnitude faster than can be achieved by conventional matrix-free Hessian matvecs via forward/adjoint PDE solves.

97 MATHEMATICS AND COMPUTING↗

A New Proposal Generalized Predictive Control Algorithm With Polynomial Reference Tracking Applied for Sodium Fast Reactors

This paper proposes a generalized predictive control (GPC) with constraints and orthonormal Laguerre functions using the simplified model of the primary system (reactor core and intermediate heat exchanger (IHX)) of a prototypical sodium fast reactor (SFR). This paper develops a multiple-input multiple-output (MIMO) GPC with input constraints able to track polynomial references of any degree applied in coolant temperature difference across the core and fractional power. The manipulated variables of the GPC-SFR are the reactivity and the sodium flow rate of the primary and secondary pipes. Moreover, orthonormal Laguerre functions and step down condition number techniques were also applied to avoid the numerical ill-conditioning issue in quadratic programming of large systems. Thus, a GPC type-2 was designed to control fractional power, coolant temperature difference across the core and sodium tank temperature of the SFR primary system when temperature references change according to a linear ramp after reaching their steady-state operation, sustaining 100% power operation on the reactor. In order to analyze the load tracking capability of the GPC-SFR type-2, the load following from 100% fractional power (FP) to 60% FP at 0.8% FP/min rate is simulated. Constraints on the rate of coolant temperature difference across the core and reactivity were applied for the design safety. For comparison criteria, this paper compares the GPC-SFR type-2 with the GPC-SFR type-1, i.e., standard model predictive control (MPC), to verify the viability and superior performance of the proposal regarding: (a) ramp-tracking capability of temperature and load; (b) the rejections of a reactivity disturbance of -1 cent and a secondary sodium inlet temperature disturbance of +10°F; and (c) a simulation with uncertainty in reactor design. The simulations show that the GPC-SFR type-2 overcome the GPC-SFR type-1 robustness and performance.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗