Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “fast algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

TEQUILA: a platform for rapid development of quantum algorithms

Variational quantum algorithms are currently the most promising class of algorithms for deployment on near-term quantum computers. In contrast to classical algorithms, there are almost no standardized methods in quantum algorithmic development yet, and the field continues to evolve rapidly. As in classical computing, heuristics play a crucial role in the development of new quantum algorithms, resulting in a high demand for flexible and reliable ways to implement, test, and share new ideas. In this paper, inspired by this demand, we introduce TEQUILA, a development package for quantum algorithms in PYTHON, designed for fast and flexible implementation, prototyping and deployment of novel quantum algorithms in electronic structure and other fields. TEQUILA operates with abstract expectation values which can be combined, transformed, differentiated, and optimized. On evaluation, the abstract data structures are compiled to run on state of the art quantum simulators or interfaces.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Genetic algorithm-based optimisation of the few-group structure for lead fast reactors analysis

The optimal choice of the few-group structure for full-core transient analyses is still an open issue in reactor physics, especially for fast system like the lead fast reactor. One possible approach to select the group boundaries is represented by heuristic search algorithms, such as evolutionary ones. In this paper, a genetic algorithm coupled with the SIMMER code is employed to determine optimized six-group boundaries for the analysis of the ALFRED reactor. The Serpent Monte Carlo code is adopted to produce both the fine-group cross section library and the fine-group flux, used as a figure of merit to drive the genetic optimisation. The results show that the algorithm is indeed able to find satisfactory solutions that comply with the set objectives and can be reasonably interpreted in light of the underlying physics of the considered core. (authors)

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

TURBOMOLE: Modular program suite for ab initio quantum-chemical and condensed-matter simulations

TURBOMOLE is a collaborative, multi-national software development project aiming to provide highly efficient and stable computational tools for quantum chemical simulations of molecules, clusters, periodic systems, and solutions. The TURBOMOLE software suite is optimized for widely available, inexpensive, and resource-efficient hardware such as multi-core workstations and small computer clusters. TURBOMOLE specializes in electronic structure methods with outstanding accuracy–cost ratio, such as density functional theory including local hybrids and the random phase approximation (RPA), GW-Bethe–Salpeter methods, second-order Møller–Plesset theory, and explicitly correlated coupled-cluster methods. TURBOMOLE is based on Gaussian basis sets and has been pivotal for the development of many fast and low-scaling algorithms in the past three decades, such as integral-direct methods, fast multipole methods, the resolution-of-the-identity approximation, imaginary frequency integration, Laplace transform, and pair natural orbital methods. This review focuses on recent additions to TURBOMOLE’s functionality, including excited-state methods, RPA and Green’s function methods, relativistic approaches, high-order molecular properties, solvation effects, and periodic systems. A variety of illustrative applications along with accuracy and timing data are discussed. Moreover, available interfaces to users as well as other software are summarized. TURBOMOLE’s current licensing, distribution, and support model are discussed, and an overview of TURBOMOLE’s development workflow is provided. Challenges such as communication and outreach, software infrastructure, and funding are highlighted.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Fast and Universal Kohn-Sham Density Functional Theory Algorithm for Warm Dense Matter to Hot Dense Plasma

Understanding many processes, e.g., fusion experiments, planetary interiors, and dwarf stars, depends strongly on microscopic physics modeling of warm dense matter and hot dense plasma. This complex state of matter consists of a transient mixture of degenerate and nearly free electrons, molecules, and ions. This regime challenges both experiment and analytical modeling, necessitating predictive ab initio atomistic computation, typically based on quantum mechanical Kohn-Sham density functional theory (KS-DFT). However, cubic computational scaling with temperature and system size prohibits the use of DFT through much of the warm dense matter regime. A recently developed stochastic approach to KS-DFT can be used at high temperatures, with the exact same accuracy as the deterministic approach, but the stochastic error can converge slowly and it remains expensive for intermediate temperatures (< 50 eV). Here we have developed a universal mixed stochastic-deterministic algorithm for DFT at any temperature. This approach leverages the physics of KS-DFT to seamlessly integrate the best aspects of these different approaches. We demonstrate that this method significantly accelerated self-consistent field calculations for temperatures from 3 to 50 eV, while producing stable molecular dynamics and accurate diffusion coefficients.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Finding Fast Transients in Real Time Using a Novel Light-curve Analysis Algorithm

The current data acquisition rate of astronomical transient surveys and the promise for significantly higher rates in the next decade necessitate the development of novel approaches to analyze astronomical data sets and promptly detect objects of interest. The Deeper, Wider, Faster (DWF) program is a survey focused on the identification of fast-evolving transients, such as fast radio bursts, gamma-ray bursts, and supernova shock breakouts. It employs multifrequency simultaneous coverage of the same part of the sky over several orders of magnitude. Using the Dark Energy Camera mounted on the 4 m Blanco telescope, DWF captures a 20 s g -band exposure every minute, at a typical seeing of ~1'' and an air mass of ~1.5. These optical data are collected simultaneously with observations conducted over the entire electromagnetic spectrum—from radio to γ -rays—as well as cosmic-ray observations. In this paper, we present a novel real-time light-curve analysis algorithm, designed to detect transients in the DWF optical data; this algorithm functions independently from, or in conjunction with, image subtraction. We present a sample of fast transients detected by our algorithm, as well as a false-positive analysis. Our algorithm is customizable and can be tuned to be sensitive to transients evolving over different timescales and flux ranges.

79 ASTRONOMY AND ASTROPHYSICS↗

Fused x-ray and fast neutron CT reconstruction for imaging large and dense objects

Megavolt x-ray computed tomography (CT) is a powerful tool for three-dimensional characterization. However, its utility is limited for large objects composed of high-atomic number (Z) materials, where x rays fail to penetrate. Information from fast neutron CT (FNCT) can complement x-ray CT reconstructions since fast neutrons can more readily penetrate high-Z objects. In this work, we demonstrate a method for combining FNCT and x-ray CT data to create a single reconstruction, more accurate than could be achieved with either x rays or fast neutrons alone. The algorithm was tested on an exemplar comprising multiple concentric, nested cylinders of different materials. Simulated and empirical x-ray CT data were acquired for the exemplar using a 9 MV bremsstrahlung spectrum. Additional simulated and empirical FNCT data were acquired using an accelerator based fast neutron source. The FNCT data were used to synthesize x-ray CT data and augment the x-ray CT data missing due to lack of penetration. This approach mitigates artifacts that would otherwise negatively affect the accuracy and resolution of a single-modality reconstructed volume.

47 OTHER INSTRUMENTATION↗

Hls4ml Synthesis Testing

HLS4ml (high level synthesis for machine learning) Is a Python package used to translate commonly used open-source machine learning models into HLS. This is useful in machine learning applications on FPGAs. Machine learning algorithms are only as fast as the hardware that they are used on, and some applications require high speed without sacrificing accuracy. In these situations, an FPGA is a good choice since it is faster than a CPU or a GPU, but programming an FPGA is difficult. This is where HLS4ml can be used to simplify the process, as a well-known learning model can be converted to HLS and more easily deployed onto an FPGA. There are many use cases for a machine learning algorithm running on an FPGA. For example, detectors in a particle accelerator cannot keep every event that they detect, and so a computer must decide which events to keep and which to discard. Using an FPGA with a machine learning algorithm would be a good way to keep as many events as possible.

Swanson, Caiden↗

hls4ml

hls4ml (high level synthesis for machine learning) Is a Python package used to translate commonly used open-source machine learning models into HLS. This is useful in machine learning applications on FPGAs. Machine learning algorithms are only as fast as the hardware that they are used on, and some applications require high speed without sacrificing accuracy. In these situations, an FPGA is a good choice since it is faster than a CPU or a GPU, but programming an FPGA is difficult. This is where hls4ml can be used to simplify the process, as a well-known learning model can be converted to HLS and more easily deployed onto an FPGA. There are many use cases for a machine learning algorithm running on an FPGA. For example, detectors in a particle accelerator cannot keep every event that they detect, and so a computer must decide which events to keep and which to discard. Using an FPGA with a machine learning algorithm would be a good way to keep as many events as possible.

Swanson, Caiden↗

Reconstruction of the 4D beam matrix

The widely used transverse parameters characterizing particle beams are the Twiss parameters. These parameters can be measured experimentally but they do not fully characterize the beam since they do not account for possible correlations in particle distribution between two transverse coordinates. These correlations may occur due to uncompensated magnetic field at the cathode or misalignment of focusing quadrupoles in the transport beamline. We test a novel diagnostic for diagnosing full 4D beam matrix which may be used to identify such imperfections. The diagnostic is based on transporting the beam through the beamline which includes a quadrupole and a skew quadrupole magnets and measuring the resulting 2D beam distribution at the screen downstream. Such a measurement can be viewed as measuring a 2D projection of the 4D distribution. Different settings of the quads provide measurements of different slices of the phase space. The reconstruction of the original beam matrix from a number of measurements is done using machine learning algorithm, which provides a fast and reliable way of reconstruction for an arbitrary configuration of the scanning beamline. In August 2024, we set up the diagnostic beamline to perform a quadrupole scan of the beam. The setup includes a skew quadrupole, a regular quadrupole, and a screen. The images on the screen were post-processed to remove experimental artifacts and enhance contrast by eliminating background noise outside the core of the distribution=. The rms parameters of the distribution were then calculated and used as inputs for the reconstruction algorithm. This algorithm attempts to determine the initial beam matrix that produces expected images on the screen closely matching the observed images across all quadrupole settings. The algorithm found a solution in which the expected rms parameters closely align with the observations. Validation of the results is planned for FY25.

43 PARTICLE ACCELERATORS↗

Randomized Algorithms for Low-Rank Matrix and Tensor Decompositions

This paper surveys randomized algorithms in numerical linear algebra for low-rank decompositions of matrices and tensors. The survey begins with a review of classical matrix algorithms that can be accelerated by randomized dimensionality reduction, such as the singular value decomposition (SVD) or interpolative (ID) and CUR decompositions. Recent advances in randomized dimensionality reduction are discussed, including new methods of fast matrix sketching and sampling techniques, which are incorporated into classical matrix algorithms for fast low-rank matrix approximations. The extension of randomized matrix algorithms to tensors is then explored for several low-rank tensor decompositions in the CP and Tucker formats, including the higher-order SVD, ID, and CUR decomposition.

Pearce, Katherine J. [The University of Texas at A↗

Advancing the Prediction of MS/MS Spectra Using Machine Learning

Tandem mass spectrometry (MS/MS) is an important tool for the identification of small molecules and metabolites where resultant spectra are most commonly identified by matching them with spectra in MS/MS reference libraries. While popular, this strategy is limited by the contents of existing reference libraries. In response to this limitation, various methods are being developed for the in silico generation of spectra to augment existing libraries. Recently, machine learning and deep learning techniques have been applied to predict spectra with greater speed and accuracy. Here, in this work, we investigate the challenges these algorithms face in achieving fast and accurate predictions on a wide range of small molecules. The challenges are often amplified by the use of generic machine learning benchmarking tactics, which lead to misleading accuracy scores. Curating data sets, only predicting spectra for sufficiently high collision energies, and working more closely with experimental mass spectrometrists are recommended strategies to improve overall prediction accuracy in this nuanced field.

47 OTHER INSTRUMENTATION↗

Using Artificial Neural Networks to Predict Physical Properties of Membrane Polymers

Membrane polymers are a promising technology for use in many challenging gas separation applications. Here, the techniques of computer-aided molecular design can be used to search through the massive molecular space of heteropolymers and develop a set of likely candidate repeat units matching specific physical property targets. However, reasonably accurate property prediction algorithms are needed, but these algorithms must be very fast in order to be combined with an optimization framework. Artificial neural networks (ANNs), a branch of machine learning, are applied in this work to predict the physical properties of polymers. All of the physical properties investigated were found to be predicted by ANNs with R 2 scores exceeding 0.82.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Flexible dynamic boundary microgrid operation considering network and load unbalances

Flexible microgrids with dynamic boundaries have recently been introduced in the literature. With the ability to reconfigure the topology of the microgrids dynamically through remotely controlled switches, flexible microgrids with dynamic boundaries can further improve the resiliency and energy efficiency of microgrids with distributed energy resources (DERs). This paper focuses on the optimal operation considering one of the predominant characteristics of microgrids and distribution systems – unbalanced networks and loads. In existing literature, balanced modeling of microgrids is more common due to its attractive simplicity. The three-phase power unbalance has not been considered as a constraint on the generation units in a microgrid. Further, negative sequence constraints have also been neglected. In this article, we propose a set of constraints that is specifically related to the capabilities of inverter interfaced resources to supply unbalanced current/power when the microgrid is islanded from the main distribution grid. We incorporate the new set of constraints into two optimization formulations leveraging two convex relaxations of the three-phase power flow equations: mixed-integer linear programming (MILP) and mixed-integer semidefinite programming (MISDP) that optimize the dispatch of controllable switches and DERs in the microgrid. The algorithms are then extended to networked microgrids with grid-forming sources. We test the algorithms on a realistic community microgrid model in Puerto Rico as well as standardized IEEE distribution test feeders. The testing results demonstrate the performance of the proposed algorithms. The MILP is fast and scalable, and the MISDP enforces the negative sequence voltage constraints.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Efficient Fourier transforms for transverse momentum dependent distributions

Hadron production at low transverse momenta in semi-inclusive deep inelastic scattering can be described by transverse momentum dependent (TMD) factorization. This formalism has also been widely used to study the Drell-Yan process and back-to-back hadron pair production in $e^+e^-$ collisions. These processes are the main ones for extractions of TMD parton distribution functions and TMD fragmentation functions, which encode important information about nucleon structure and hadronization. One of the most widely used TMD factorization formalism in phenomenology formulates TMD observables in coordinate $b_\perp$-space, the conjugate space of the transverse momentum. The Fourier transform from $b_\perp$-space back into transverse momentum space is sufficiently complicated due to oscillatory integrands that it requires a careful and computationally intensive numerical treatment in order to avoid potentially large numerical errors. Within the TMD formalism, the azimuthal angular dependence is analytically integrated and the two-dimensional $b_\perp$ integration reduces to a one-dimensional integration over the magnitude $b_\perp$. In this paper we develop a fast numerical Hankel transform algorithm for such a $b_\perp$-integration that improves the numerical accuracy of TMD calculations in all standard processes. Libraries for this algorithm are implemented in Python 2.7 and 3, C++, as well as FORTRAN77. All packages are made available open source.

97 MATHEMATICS AND COMPUTING↗

Efficient Decision Trees for Tensor Regressions

Here, we proposed the tensor-input tree (TT) method for scalar-on-tensor and tensor-on-tensor regression problems. We first address scalar-on-tensor problem by proposing scalar-output regression tree models whose input variables are tensors (i.e., multi-way arrays). We devised and implemented fast randomized and deterministic algorithms for efficient fitting of scalar-on-tensor trees, making TT competitive against tensor-input GP models (Yu, Li, and Liu; Sun et al.). Based on scalar-on-tensor tree models, we extend our method to tensor-on-tensor problems using additive tree ensemble approaches. Theoretical justification and extensive experiments, including testing robustness to entrywise input tensor noise, are provided on real and synthetic datasets to illustrate the performance of TT. Our implementation is provided at https://github.com/hrluo/TensorDecisionTreeRegressor. Supplementary materials for this article are available online.

Decision tree regressions↗

Photometric Redshifts and Galaxy Clusters for DES DR2, DESI DR9, and HSC-SSP PDR3 Data

Photometric redshift (photoz) is a fundamental parameter for multi-wavelength photometric surveys, while galaxy clusters are important cosmological probes and ideal objects for exploring the dense environmental impact on galaxy evolution. We extend our previous work on estimating photoz and detecting galaxy clusters to the latest data releases of the Dark Energy Spectroscopic Instrument (DESI) imaging surveys, Dark Energy Survey (DES) and Hyper Suprime-Cam Subaru Strategic Program (HSC-SSP) imaging surveys and make corresponding catalogs publicly available for more extensive scientific applications. The photoz catalogs include accurate measurements of photoz and stellar mass for about 320, 293 and 134 million galaxies with r < 23, i < 24 and i < 25 in DESI DR9, DES DR2 and HSC-SSP PDR3 data, respectively. The photoz accuracy is about 0.017, 0.024 and 0.029 and the general redshift coverage is z < 1, z < 1.2 and z < 1.6, respectively for those three surveys. Furthermore, the uncertainty of the logarithmic stellar mass that is inferred from stellar population synthesis fitting is about 0.2 dex. With the above photoz catalogs, galaxy clusters are detected using a fast cluster-finding algorithm. A total of 532,810, 86,963 and 36,566 galaxy clusters with the number of members larger than 10 is discovered for DESI, DES and HSC-SSP, respectively. Their photoz accuracy is at the level of 0.01. The total mass of our clusters is also estimated by using the calibration relations between the optical richness and the mass measurement from X-ray and radio observations. The photoz and cluster catalogs are available at ScienceDB (https://www.doi.org/10.11922/sciencedb.o00069.00003) and PaperData Repository (https://doi.org/10.12149/101089).

79 ASTRONOMY AND ASTROPHYSICS↗

A Holistic Algorithmic Approach to Improving Accuracy, Robustness, and Computational Efficiency for Atmospheric Dynamics

Atmospheric weather and climate models must perform simulations very quickly to be useful. Therefore, modelers have traditionally focused on reducing computations as much as possible. However, in our new era of increasingly compute-capable hardware, data movement is now the prohibiting expense. This study examines the computational benefits of a new algorithmic approach to modeling atmospheric dynamics on scales relevant to weather and climate simulation. Rather than minimizing computations, this new approach considers the larger problem more holistically, including spatial accuracy, temporal accuracy, robustness (i.e., oscillations), on-node efficiency, and internode data transfers together at once. Numerical experiments demonstrate how computations can be strategically increased to simultaneously address each of these constraints while reducing data movement to adapt to modern accelerated hardware. The new algorithm can achieve at times up to 80% peak floating point throughput in single precision on the Nvidia Tesla V100 GPU, where the traditional approach is shown to only achieve single-digit floating point efficiency. Further, the new algorithm is twice as fast as a standard Runge--Kutta time integrator, and high-order accuracy with Weighted Essentially Non-Oscillatory (WENO) limiting came at less than 30% additional runtime cost on a GPU, thus increasing the accuracy per degree of freedom.

54 ENVIRONMENTAL SCIENCES↗

Fast and Accurate Intersections on a Sphere

We introduce a fast, high-precision algorithm for calculating intersections between great circle arcs and lines of constant latitude on the unit sphere. We first propose a simplified intersection point formula with improved speed and numerical robustness over the ones traditionally implemented in geoscience software. We then show how algorithms based on the concept of error-free transformations (EFT) can be applied to evaluate this formula within a relative error bound that is on the order of machine precision. Here, we demonstrate that, with a vectorized and parallelized implementation, this enhanced accuracy is achieved with no compute time overhead compared to a direct calculation in hardware floating point, making our algorithm suitable for performance-sensitive applications like regridding of high-resolution climate data. In contrast, evaluating our formula using high-precision data types like quadruple precision and arbitrary precision, or using the robust intersection computation routines from the Computational Geometry Algorithms Library, leads to significant computational overhead, especially since these alternatives inhibit vectorization. More generally, our work demonstrates how EFT techniques can be combined and extended to implement nontrivial geometric calculations with high accuracy and speed.

Environmental sciences↗