Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “fast algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

High-Fidelity Models and Fast EMT Simulation Algorithms for Isolated Multi-port Autonomous Reconfigurable Solar power plant (MARS)

The integration of hybrid photovoltaic (PV) and energy storage system (ESS) based plants has become a promising way of solving the intermittency of PV plants and providing frequency support to the power grid. The multi-port autonomous reconfigurable solar power plant (MARS) can integrate the PV systems and ESSs to an ac grid and dc lines. The proposed isolated MARS incorporates an isolated converter that connects to the PV arrays and is based on the dual active bridge (DAB) converter. The high frequency switching in the DAB and the means to control the DAB converter using delays between switching signals lead to the need for a small timestep in simulations. Moreover, several hundreds of modules that include a DAB converter are present in the MARS. The small timestep and the presence of several hundreds of modules lead to a significant rise in the overall simulation time. To address this issue, simulation algorithms like numerical stiffness-based hybrid discretization and the hysteresis relaxation technique are applied to the switched system model of isolated MARS. Additionally, an event-driven interpolating method is introduced to help increase the minimum timestep to simulate the conventional DAB converter model while maintaining high accuracy in the simulation results. The developed model is validated by comparison with its reference model built in the PSCAD/EMTDC and MATLAB software environments using library components.

Xia, Qian↗

Piecewise linear approximation with minimum number of linear segments and minimum error: A fast approach to tighten and warm start the hierarchical mixed integer formulation

In several areas of economics and engineering, it is often necessary to fit discrete data points or approximate nonlinear functions with continuous functions. Piecewise linear (PWL) functions are a convenient way to achieve this. PWL functions can be modeled in mathematical problems using only linear and integer variables. Moreover, there is a computational benefit in using PWL functions that have the least possible number of segments. This work proposes a novel hierarchical mixed integer linear programming (MILP) formulation that identifies a continuous PWL approximation with minimum number of linear segments for a given target maximum error. The proposed MILP formulation also identifies the solution with the least maximum error among the solutions with minimum number of segments. Then, this work proposes a fast iterative algorithm that identifies non necessarily continuous PWL approximations by solving O(S log N) linear programming (LP) problems, where N is the number of data points and S is the minimum number of segments in the non necessarily continuous case. This work demonstrates that tight bounds for the MILP problem can be derived from these approximations. Next, a fast algorithm is introduced to transform a non necessarily continuous PWL approximation into a continuous one. Finally, the tight bounds and the continuous PWL approximations are used to tighten and warm start the MILP problem. The tightened formulation is shown in experimental results to be more efficient, especially for large data sets, with a solution time that is up to two orders of magnitude less than the existing literature.

97 MATHEMATICS AND COMPUTING↗

Fast tree-based algorithms for DBSCAN for low-dimensional data on GPUs

DBSCAN is a well-known density-based clustering algorithm to discover arbitrary shape clusters. While conceptually simple in serial, the algorithm is challenging to efficiently parallelize on manycore GPU architectures. Common pitfalls, such as asynchronous range query calls, result in high thread execution divergence in many implementations. In this paper, we propose a new framework for GPU-accelerated DBSCAN, and describe two tree-based algorithms within that framework. Both algorithms fuse the search for neighbors with updating cluster information, but differ in their treatment of dense regions of the data. We show that the time taken to compute clusters is at most twice that of determination of the neighbors. We compare the proposed algorithms with existing CPU and GPU implementations, and demonstrate their competitiveness and performance using a fast traversal structure (bounding volume hierarchy) for low dimensional data. We also show that the memory usage can be reduced by processing object neighbors dynamically without storing them.

Prokopenko, Andrey↗

Butterfly Factorization Via Randomized Matrix-Vector Multiplications

This paper presents an adaptive randomized algorithm for computing the butterfly factorization of an m × n matrix with m ≈ n provided that both the matrix and its transpose can be rapidly applied to arbitrary vectors. The resulting factorization is composed of O(log n) sparse factors, each containing O(n) nonzero entries. The factorization can be attained using O(n 3/2 log n) computation and O(n log n) memory resources. Furthermore, the proposed algorithm can be implemented in parallel and can apply to matrices with strong or weak admissibility conditions arising from surface integral equation solvers as well as multi-frontal-based finite-difference, finite-element, or finite-volume solvers. A distributed-memory parallel implementation of the algorithm demonstrates excellent scaling behavior.

97 MATHEMATICS AND COMPUTING↗

A Study of Model-Based Protective Fast-Charging and Associated Degradation in Commercial Smartphone Cells: Insights on Cathode Degradation as a Result of Lithium Depositions on the Anode

The ever expanding mobile consumer electronic market has accelerated the need for safe and efficient fast-charging approaches that improve the overall speed of battery charging without hastened deterioration of the battery performance. Herein, the impact of a resource inexpensive, physics-based, electrochemically optimized fast-charging algorithm (charging time < 2 h) for mobile devices is investigated. A critical difference in the amount and morphology of lithium deposits on the anode for cells fast-charged without an optimized algorithm is observed and found to be the main cause of capacity decay. Furthermore, an in-depth study of the LiCoO 2 cathode regions opposite to pronounced lithium deposits on the anode reveals a “mirroring” phenomenon, i.e., a frozen monoclinic phase, and inactivity to relithiation. In operando hard X-ray absorption spectroscopy reveals that degraded spots on harvested cathodes seem to be activated again and participate in the intercalation process when lithiated at low rates from lithium foil counter electrodes. On the other hand, tests at higher C-rates, closer to the actual fast-charging rate, reveal only negligible oxidation state changes and therefore poor performance.

25 ENERGY STORAGE↗

Distribution Feeder-Scale Fast Frequency Response via Optimal Coordination of Net-load Resources Part II: Large-Scale Demonstration

This work is the second of a two-part series in which we develop and experimentally demonstrate a hierarchical control solution for optimally coordinating thousands of deferrable loads and distributed energy resources (DERs) to provide fast frequency response (FFR) from an entire distribution feeder. In Part I, we developed and proved practical algorithms for fast, cost-based optimal dispatch and for determining the optimal amount of headroom to operate solar inverters with to support FFR dispatch while minimizing opportunity cost. Simulation results in Part I demonstrated the advantages of the hierarchical dispatch approach in being able to maintain fast solution times needed for FFR even when the problem size increases. In Part II, we implement the algorithms developed in Part I in a novel, large-scale power hardware-in-the-loop experiment including embedded controllers and more than 100 powered appliance loads and DER connected to a simulated real-world distribution system with more than 10,000 controlled devices. Experimental results from multiple scenarios confirm that the optimal FFR dispatch approach scales well and can optimally coordinate more than 10,000 net-load resources across a distribution network while achieving hardware response times within 500 ms, which is not possible using state-of-the-art optimal coordination approaches.

24 POWER TRANSMISSION AND DISTRIBUTION↗

A Linear-Complexity Tensor Butterfly Algorithm for Compressing High-Dimensional Oscillatory Integral Operators

This paper presents a multilevel tensor compression algorithm called tensor butterfly algorithm for efficiently representing large-scale and high-dimensional oscillatory integral operators, including Green's functions for wave equations and integral transforms such as Radon transforms and Fourier transforms. The proposed algorithm leverages a tensor extension of the so-called complementary low-rank property of existing matrix butterfly algorithms. The algorithm partitions the discretized integral operator tensor into subtensors of multiple levels and factorizes each subtensor at the middle level as a Tucker-type interpolative decomposition, whose factor matrices are formed in a multilevel fashion. For a d-dimensional (d > 1) integral operator discretized into a 2d-mode tensor with n2d entries, the overall CPU time and memory requirement scale as O(nd), in stark contrast to the O(nd log n) complexity of existing matrix algorithms such as matrix butterfly algorithms and fast Fourier transforms (FFTs), where n is the number of points per direction. When comparing with other tensor algorithms such as quantized tensor train (QTT), the proposed algorithm also shows superior CPU and memory performance for tensor contraction. Remarkably, the tensor butterfly algorithm can efficiently model high-frequency Green's function interactions between two unit cubes, each spanning 512 wavelengths per direction, which represents problems of scale over 512× larger than that existing butterfly algorithms can handle, with the same amount of computation resources. On the other hand, for a problem representing 64 wavelengths per direction, which is the largest size existing algebraic matrix algorithms can handle, our tensor butterfly algorithm exhibits 200x speedups and 30× memory reduction compared with existing ones. Moreover, the tensor butterfly algorithm also permits O(nd)-complexity FFTs and Radon transforms up to d = 6 dimensions.

Kielstra, P Michael↗

Optimizing Mu2e Spill Regulation System Algorithms

A slow extraction system is being developed for the Fermilab’s Delivery Ring to deliver protons to the Mu2e experiment. During the extraction, the beam on target experiences small intensity variations owing to many factors. Various adaptive learning algorithms will be employed for beam regulation to achieve the required spill quality. We discuss here preliminary results of the slow and fast regulation algorithms validation through the computer simulations before their implementation in the FPGA. Particle tracking with sextupole resonance was used to determine the fine shape of the spill profile. Fast semi-analytical simulation schemes and Machine Learning models were used to optimize the fast regulation loop.

43 PARTICLE ACCELERATORS↗

Absolute contrast estimation for soft X-ray photon fluctuation spectroscopy using a variational droplet model

Abstract X-ray photon fluctuation spectroscopy using a two-pulse mode at the Linac Coherent Light Source has great potential for the study of quantum fluctuations in materials as it allows for exploration of low-energy physics. However, the complexity of the data analysis and interpretation still prevent recovering real-time results during an experiment, and can even complicate post-analysis processes. This is particularly true for high-spatial resolution applications using CCDs with small pixels, which can decrease the photon mapping accuracy resulting from the large electron cloud generation at the detector. Droplet algorithms endeavor to restore accurate photon maps, but the results can be altered by their hyper-parameters. We present numerical modeling tools through extensive simulations that mimic previous x-ray photon fluctuation spectroscopy experiments. By modification of a fast droplet algorithm, our results demonstrate how to optimize the precise parameters that lift the intrinsic counting degeneracy impeding accuracy in extracting the speckle contrast. These results allow for an absolute determination of the summed contrast from multi-pulse x-ray speckle diffraction, the modus operandi by which the correlation time for spontaneous fluctuations can be measured.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Assessing Historical Variability of South Asian Monsoon Lows and Depressions With an Optimized Tracking Algorithm

Abstract Cyclonic low‐pressure systems (LPS) produce abundant rainfall in South Asia, where they are traditionally categorized as monsoon lows, monsoon depressions, and more intense cyclonic storms. The India Meteorological Department (IMD) has tracked monsoon depressions for over a century, finding a large decline in their number in recent decades, but their methods have changed over time and do not include monsoon lows. This study presents a fast, objective algorithm for identifying monsoon LPS and uses it to assess interannual variability and trends in reanalyses. Variables and thresholds used in the algorithm are selected to best match a subjectively analyzed LPS data set while minimizing disagreement between four reanalyses in a training period. The stream function of 850 hPa horizontal wind is found to be optimal in this sense; it is less noisy than vorticity and represents the complete nondivergent wind, even when flow is not geostrophic. Using this algorithm, LPS statistics are computed for five reanalyses, and none show a detectable trend in monsoon depression counts since 1979. Both the Japanese 55‐year Reanalysis (JRA‐55) and the IMD data set show a step‐like reduction in depression counts when they began using geostationary satellite data, in 1979 and 1982, respectively; the 1958–2018 linear trend in JRA‐55, however, is smaller than in the IMD data set, and its error bar includes 0. There are more LPS in seasons with above‐average monsoon rainfall and in La Niña years, but few other large‐scale modes of interannual variability are found to modulate LPS counts, lifetimes, or track length consistently across reanalyses.

54 ENVIRONMENTAL SCIENCES↗

Lattice Green’s Functions for High-Order Finite Difference Stencils

Lattice Green's Functions (LGFs) are fundamental solutions to discretized linear operators, and as such they are a useful tool for solving discretized elliptic PDEs on domains that are unbounded in one or more directions. The majority of existing numerical solvers that make use of LGFs rely on a second-order discretization and operate on domains with free-space boundary conditions in all directions. Under these conditions, fast expansion methods are available that enable precomputation of 2D or 3D LGFs in linear time, avoiding the need for brute-force multi-dimensional quadrature of numerically unstable integrals. Here we focus on higher-order discretizations of the Laplace operator on domains with more general boundary conditions, by (1) providing an algorithm for fast and accurate evaluation of the LGFs associated with high-order dimension-split centered finite differences on unbounded domains, and (2) deriving closed-form expressions for the LGFs associated with both dimension-split and Mehrstellen discretizations on domains with one unbounded dimension. Through numerical experiments we demonstrate that these techniques provide LGF evaluations with near machine-precision accuracy, and that the resulting LGFs allow for numerically consistent solutions to high-order discretizations of the Poisson's equation on fully or partially unbounded 3D domains.

97 MATHEMATICS AND COMPUTING↗

Distribution Feeder-Scale Fast Frequency Response via Optimal Coordination of Net-load Resources Part I: Solution Design

This work is the first of a two-part series that develops and experimentally demonstrates a first-of-its-kind hierarchical control solution for optimally dispatching thousands of deferrable loads and distributed energy resources (DERs) across a distribution feeder to provide fast frequency response (FFR) within 500 ms to the bulk power system. This approach rapidly coordinates resources online after a frequency event occurs, allowing fast-changing, behind-the-meter (BTM) resources to be incorporated and aggregate FFR power set points to be achieved more quickly and accurately than existing approaches. We also present a solution for determining the optimal amount of headroom to operate solar inverters with to minimize opportunity cost while ensuring the FFR response viability of a building with the inverter and deferrable loads. In Part I, we develop practical algorithms for fast, cost-based optimal dispatch at multiple aggregation scales (single building, multiple buildings, and full distribution feeder), establish their optimality, and demonstrate via simulation that they are faster than state-of-the-art, coordinated frequency response approaches. In Part II, the entire platform is implemented and experimentally verified using a unique power hardware-in-the-loop demonstration, including more than 100 powered loads and DERs connected to a real-world distribution network model and over 10,000 net-load resources dispatched.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Towards a Quantum Algorithm for the Incompressible Nonlinear Navier-Stokes Equations

In this work, we present novel concepts for quantum algorithms to solve transient, nonlinear partial differential equations (PDEs). The challenge lies in how to effectively represent, encode, process, and evolve the nonlinear system of PDEs on quantum computers. We will discuss the new techniques using the incompressible Navier-Stokes equations as an example, because it represents the fundamental nonlinear feature and yet removes certain complexity in physics, allowing us to focus on the design of quantum algorithms. Previous attempts solving nonlinear PDEs in quantum computation have often involved storing multiple copies of solutions or employing linearizations. Neither is practical due to exponential scaling with evolution time or insufficient solution accuracy. We propose a new framework based on matrix product states (MPSs) and matrix product operators (MPOs), in addition to the Krylov subspace methods. For example, the solution variables of the Navier-Stokes equations are represented by MPSs, and the linear and nonlinear terms are processed by MPOs. The time evolution of the operators is attained by a fast-forwarding algorithm using Krylov subspace methods. Furthermore, we discuss various techniques for efficient encoding of MPSs, measurement reduction for MPOs, and use of tensor operations to treat multi-variate, multi-physics characteristics of Navier-Stokes.

Gopalakrishnan Meena, Murali [ORNL] (ORCID:0000000↗

High-performance data format for scientific data storage and analysis

Here, in this article, we present the High-Performance Output (HiPO) data format developed at Jefferson Laboratory for storing and analyzing data from Nuclear Physics experiments. The format was designed to efficiently store large amounts of experimental data, utilizing modern fast compression algorithms. The purpose of this development was to provide organized data in the output, facilitating access to relevant information within the large data files. The HiPO data format has features that are suited for storing raw detector data, reconstruction data, and the final physics analysis data efficiently, eliminating the need to do data conversions through the lifecycle of experimental data. The HiPO data format is implemented in C++ and JAVA, and provides bindings to FORTRAN, Python, and Julia, providing users with the choice of data analysis frameworks to use. In this paper, we will present the general design and functionalities of the HiPO library and compare the performance of the library with more established data formats used in data analysis in High Energy and Nuclear Physics (such as ROOT and Parquete). In columnar data analysis, HiPO surpasses established data formats in performance and can be effectively applied to data analysis in other scientific fields.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Location-Dependent Cobalt Deposition in Smartphone Cells upon Long-Term Fast-Charging Visualized by Synchrotron X-ray Fluorescence

In this work, we investigate the transition-metal dissolution of the layered cathode material LiCoO 2 upon repeated fast-charging of three smartphone batteries from different manufacturers using synchrotron micro X-ray fluorescence (μ-XRF). Using this spatially resolved technique, dissolution of Co and subsequent location-dependent deposition on the anode are observed. μ-XRF mapping of selected parts of the anode electrode sheets, such as electrode folds and edges of the jelly roll, reveals the difference in the way Co is deposited on specific regions of the anode electrode. While some folds show no depositions, edges of the anode show gradually accumulating Co depositions. Furthermore, careful quantification of the dissolved Co reveals that the capacity loss scales with the amount of deposited Co on the anode, that is, total Co loss from within the cathode. Soft X-ray absorption spectroscopy of the Co depositions on the anode shows that Co is mainly deposited in a reduced 2 + state. While optimization of the fast-charging protocol mitigates Li plating on the anode, no significant difference in the amount of deposited Co can be observed between an optimized and a nonoptimized fast-charging algorithm.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Efficient atmospheric, solar, and supernova neutrino propagation through the Earth

Algorithms for computing neutrino oscillation probabilities in sharply varying matter potentials such as the Earth are becoming increasingly important. As the next generation of experiments, DUNE and HyperK as well as the IceCube upgrade and KM3NeT, come online, the computational cost for atmospheric and solar neutrinos will continue to increase. To address these issues, we expand upon our previous algorithm for long-baseline calculations to efficiently handle probabilities through the Earth for atmospheric, nighttime solar, and supernova neutrinos. The algorithm is fast, flexible, and accurate. It can handle arbitrary Earth models with two different schemes for varying density profiles. We also provide a c ++ implementation of the code called NUF ast- E arth along with a detailed user manual. The code intelligently keeps track of repeated calculations and only recalculates what is needed on each successive call which can also help provide significant speed-ups.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Laminography as a tool for imaging large-size samples with high resolution

Despite the increased brilliance of the new generation synchrotron sources, there is still a challenge with high-resolution scanning of very thick and absorbing samples, such as a whole mouse brain stained with heavy elements, and, extending further, brains of primates. Samples are typically cut into smaller parts, to ensure a sufficient X-ray transmission, and scanned separately. Compared with the standard tomography setup where the sample would be cut into many pillars, the laminographic geometry operates with slab-shaped sections significantly reducing the number of sample parts to be prepared, the cutting damage and data stitching problems. In this work, a laminography pipeline for imaging large samples (>1 cm) at micrometre resolution is presented. The implementation includes a low-cost instrument setup installed at the 2-BM micro-CT beamline of the Advanced Photon Source. Additionally, sample mounting, scanning techniques, data stitching procedures, a fast reconstruction algorithm with low computational complexity, and accelerated reconstruction on multi-GPU systems for processing large-scale datasets are presented. The applicability of the whole laminography pipeline was demonstrated by imaging four sequential slabs throughout an entire mouse brain sample stained with osmium, in total generating approximately 12 TB of raw data for reconstruction.

47 OTHER INSTRUMENTATION↗