Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “runtime”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23

Quantum approximate optimization of the long-range Ising model with a trapped-ion quantum simulator

Quantum computers and simulators may offer significant advantages over their classical counterparts, providing insights into quantum many-body systems and possibly improving performance for solving exponentially hard problems, such as optimization and satisfiability. Here, we report the implementation of a low-depth Quantum Approximate Optimization Algorithm (QAOA) using an analog quantum simulator. We estimate the ground-state energy of the Transverse Field Ising Model with long-range interactions with tunable range, and we optimize the corresponding combinatorial classical problem by sampling the QAOA output with high-fidelity, single-shot, individual qubit measurements. We execute the algorithm with both an exhaustive search and closed-loop optimization of the variational parameters, approximating the ground-state energy with up to 40 trapped-ion qubits. We benchmark the experiment with bootstrapping heuristic methods scaling polynomially with the system size. We observe, in agreement with numerics, that the QAOA performance does not degrade significantly as we scale up the system size and that the runtime is approximately independent from the number of qubits. We finally give a comprehensive analysis of the errors occurring in our system, a crucial step in the path forward toward the application of the QAOA to more general problem instances.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Threadsafe Dynamic Neighbor Lists for Monte Carlo Ray Tracing

Monte Carlo (MC) transport codes offer high-fidelity modeling of particle transport physics, but their high computational cost makes them impractical for many applications. For some applications such as multiphysics and depletion that use finely discretized geometries, a large portion of this computational cost is attributable to ray tracing. Neighbor lists are a well-known method for accelerating ray-tracing calculations in a MC code, but despite their prevalence, little work has been published on the details of their implementation. The fine details can have a significant impact on performance, particularly when using shared-memory parallelism. This paper addresses these details of implementation with a discussion of different neighbor list schemes and their impact on software runtime. Performance tests were run by using OpenMC on a pin-cell problem discretized with up to 200 axial regions. The results demonstrate that switching from surface-based to cell-based neighbor lists leads to a 10 faster calculation rate for the most fine discretization. Finally, using a threadsafe shared-memory data structure results in a 20% faster calculation rate versus simple threadprivate neighbor lists. Results here show that a data structure that is contiguous in memory improves performance by only 1% to 2% over noncontiguous linked lists.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

A Code-Agnostic Driver Application for Coupled Neutronics and Thermal-Hydraulic Simulations

While the literature has numerous examples of Monte Carlo and computational fluid dynamics (CFD) coupling, most are hard-wired codes intended primarily for research rather than as standalone, general-purpose applications. In this work, we describe an open source application, ENRICO, that enables coupled neutronic and thermal-hydraulic simulations between multiple codes that can be chosen at runtime (as opposed to a coupling between two specific codes). The application has been designed such that the control flow logic, domain mapping, nonlinear fixed-point iteration, solution transfers, and convergence checks are all agnostic to the underlying physics solvers used. Special emphasis has also been placed on enabling efficient execution on distributed-memory computing environments. The transfer of solution fields between solvers is performed in memory rather than through filesystem I/O. Additionally, solvers can be configured to run on overlapping or disjoint sets of processes. To date, coupling with the OpenMC and Shift Monte Carlo codes, the Nek5000 CFD code, and a simplified heat diffusion and subchannel solver has been implemented in ENRICO. We present results for coupled simulations of a single light-water reactor fuel assembly based on the NuScale reactor using various combinations of the physics solvers. For this problem, the coupled simulations are shown to converge in about four Picard iterations. A comparison of the heat source and temperature distributions computed by ENRICO using OpenMC coupled with Nek5000 and Shift coupled with Nek5000 illustrates remarkable agreement between the codes.

42 ENGINEERING↗

Application of Fuel Depletion Chain Simplification to Experiment Analysis in the Advanced Test Reactor

An irradiation experiment analysis can be informed by high-fidelity reactor engineering depletion results, but this comes at a computational cost. Applying depletion chain simplification to the advanced test reactor driver fuel before performing experiment depletions permits their programmatic parameters to be calculated faster, with a small penalty to accuracy. Here, this work contrasts the results of two irradiation experiments with different neutronic characteristics. Overall, the simplified nuclide library produced using a simple one-group microscopic cross-section library for a pressurized water reactor in the depletion chain simplification process performed comparably in terms of accuracy and runtime to the simplified nuclide library produced using a three-group microscopic cross-section library generated specifically for the advanced test reactor experiments being modeled. This is attributed to the additional nuclides and transmutation pathways preserved in the one-group cross-section library, which has data for 297 nuclides, compared to the three-group cross-section library, which has data for 217 nuclides. This indicates that a cross-section library with more nuclides is better than a cross-section library with fewer nuclides for the depletion chain simplification process, even if the cross-section library with fewer nuclides better represents the flux spectrum of the system being considered.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Development of near-optimal advanced control sequences for chiller plants with water-side economizers in U.S. Climates (ASHRAE RP-1661)

Various advanced control sequences for chiller plants with water-side economizers (WSE) have been proposed in literature, but the evaluation and optimization of those controls is limited. It is possible to maximize energy savings by selecting different sequences and related parameters based on the plant configuration, load, and climate. This paper addresses this gap by developing near-optimal advanced control sequences for chiller plants with WSEs. First, advanced control sequences for chiller plants with WSEs are categorized into condenser water, chilled water, and hybrid controls and representative sequences from each category are identified. Next, 504 different scenarios are optimized. These scenarios represent all possible combinations of two plant configurations, a constant or variable load profile, three advanced control sequences, and seven optimization parameter combinations in six climate zones. The results show the recommended near-optimal sequences can reduce energy consumption by up to 15% relative to the baseline depending on the configuration, load profile, and climate. Specifically, the CW-CHW sequence is recommended for the majority of systems because it is often the most energy efficient and/or reduces the runtime of chillers. The methodology in this paper provides practical guidance for achieving energy savings through near-optimal control of chiller plants with WSEs.

42 ENGINEERING↗

Benchmark for two-dimensional large scale coherent structures in partially magnetized E × B plasmas—community collaboration & lessons learned

Low-temperature plasmas (LTPs) are essential to both fundamental scientific research and critical industrial applications. As in many areas of science, numerical simulations have become a vital tool for uncovering new physical phenomena and guiding technological development. Code benchmarking remains crucial for verifying implementations and evaluating performance. This work continues the Landmark benchmark initiative, a series specifically designed to support the verification of LTP codes. In this study, seventeen simulation codes from a collaborative community of nineteen international institutions modeled a partially magnetized E × B Penning discharge. The emergence of large scale coherent structures, or rotating plasma spokes, endows this configuration with an enormous range of time scales, making it particularly challenging to simulate. The codes showed excellent agreement on the rotation frequency of the spoke as well as key plasma properties, including time-averaged ion density, plasma potential, and electron temperature profiles. Achieving this level of agreement came with challenges, and we share lessons learned on how to conduct future benchmarking campaigns. Comparing code implementations, computational hardware, and simulation runtimes also revealed interesting trends, which are summarized with the aim of guiding future plasma simulation software development.

benchmarking↗

A fast particle-mesh simulation of non-linear cosmological structure formation with massive neutrinos

Quasi-N-body simulations, such as FastPM, provide a fast way to simulate cosmological structure formation, but have yet to adequately include the effects of massive neutrinos. In this work, we present a method to include neutrino particles in FastPM, enabling computation of the CDM and total matter power spectra to percent-level accuracy in the non-linear regime. The CDM-neutrino cross-power can also be computed at a sufficient accuracy to constrain cosmological observables. To avoid the shot noise that typically plagues neutrino particle simulations, we employ a quasi-random algorithm to sample the relevant Fermi-Dirac distribution when setting the initial neutrino thermal velocities. We additionally develop an effective distribution function to describe a set of non-degenerate neutrinos as a single particle to speed up non-degenerate simulations. The simulation is accurate for the full range of physical interest, M ν ≲ 0.6eV, and applicable to redshifts z ≲ 2. Such accuracy can be achieved by initializing particles with the two-fluid approximation transfer functions (using the REPS package). Convergence can be reached in ~ 25 steps, with a starting redshift of z=99. Probing progressively smaller scales only requires an increase in the number of CDM particles being simulated, while the number of neutrino particles can remain fixed at a value less than or similar to the number of CDM particles. In turn, the percentage increase in runtime-per-step due to neutrino particles is between ~ 5-20% for runs with 1024 3 CDM particles, and decreases as the number of CDM particles is increased. The code has been made publicly available, providing an invaluable resource to produce fast predictions for cosmological surveys and studying reconstruction.

79 ASTRONOMY AND ASTROPHYSICS↗

Particle hit clustering and identification using point set transformers in liquid argon time projection chambers

Liquid argon time projection chambers are often used in neutrino physics and dark-matter searches because of their high spatial resolution. The images generated by these detectors are extremely sparse, as the energy values detected by most of the detector are equal to 0, meaning that despite their high resolution, most of the detector is unused in a particular interaction. Instead of representing all of the empty detections, the interaction is usually stored as a sparse matrix, a list of detection locations paired with their energy values. Traditional machine learning methods that have been applied to particle reconstruction such as convolutional neural networks (CNNs), however, cannot operate over data stored in this way and therefore must have the matrix fully instantiated as a dense matrix. Operating on dense matrices requires a lot of memory and computation time, in contrast to directly operating on the sparse matrix. We propose a machine learning model using a point set neural network that operates over a sparse matrix, greatly improving both processing speed and accuracy over methods that instantiate the dense matrix, as well as over other methods that operate over sparse matrices. Compared to competing state-of-the-art methods, our method improves classification performance by 14%, segmentation performance by more than 22%, while taking 80% less time and using 66% less memory. Compared to state-of-the-art CNN methods, our method improves classification performance by more than 86%, segmentation performance by more than 71%, while reducing runtime by 91% and reducing memory usage by 61%.

calibration and fitting methods↗

XACC: a system-level software infrastructure for heterogeneous quantum–classical computing

Quantum programming techniques and software have advanced significantly over the past five years, with a majority focusing on high-level language frameworks targeting remote REST library APIs. As quantum computing architectures advance and become more widely available, lower-level, system software infrastructures will be needed to enable tighter, co-processor programming and access models. In this work, we present XACC, a system-level software infrastructure for quantum–classical computing that promotes a service-oriented architecture to expose interfaces for core quantum programming, compilation, and execution tasks. Additionally, we detail XACC's interfaces, their interactions, and its implementation as a hardware-agnostic framework for both near-term and future quantum–classical architectures. We provide concrete examples demonstrating the utility of this framework with paradigmatic tasks. Our approach lays the foundation for the development of compilers, associated runtimes, and low-level system tools tightly integrating quantum and classical workflows.

97 MATHEMATICS AND COMPUTING↗

Adaptive pruning-based optimization of parameterized quantum circuits

Abstract Variational hybrid quantum–classical algorithms are powerful tools to maximize the use of noisy intermediate-scale quantum devices. While past studies have developed powerful and expressive ansatze, their near-term applications have been limited by the difficulty of optimizing in the vast parameter space. In this work, we propose a heuristic optimization strategy for such ansatze used in variational quantum algorithms, which we call ‘parameter-efficient circuit training (PECT)’. Instead of optimizing all of the ansatz parameters at once, PECT launches a sequence of variational algorithms, in which each iteration of the algorithm activates and optimizes a subset of the total parameter set. To update the parameter subset between iterations, we adapt the Dynamic Sparse Reparameterization scheme which was originally proposed for training deep convolutional neural networks. We demonstrate PECT for the Variational Quantum Eigensolver, in which we benchmark unitary coupled-cluster ansatze including UCCSD and k -UpCCGSD, as well as the Low-Depth Circuit Ansatz (LDCA), to estimate ground state energies of molecular systems. We additionally use a layerwise variant of PECT to optimize a hardware-efficient circuit for the Sycamore processor to estimate the ground state energy densities of the one-dimensional Fermi-Hubbard model. From our numerical data, we find that PECT can enable optimizations of certain ansatze that were previously difficult to converge and more generally can improve the performance of variational algorithms by reducing the optimization runtime and/or the depth of circuits that encode the solution candidate(s).

Physics↗

Linear-depth quantum circuits for loading Fourier approximations of arbitrary functions

Abstract The ability to efficiently load functions on quantum computers with high fidelity is essential for many quantum algorithms, including those for solving partial differential equations and Monte Carlo estimation. In this work, we introduce the Fourier series loader (FSL) method for preparing quantum states that exactly encode multi-dimensional Fourier series using linear-depth quantum circuits. Specifically, the FSL method prepares a (Dn)-qubit state encoding the 2 Dn -point uniform discretization of aD-dimensional function specified by aD-dimensional Fourier series. A free parameter,m, which must be less thann, determines the number of Fourier coefficients, 2 D ( m + 1 ) , used to represent the function. The FSL method uses a quantum circuit of depth at most 2 ( n − 2 ) + ⌈ log 2 ( n − m ) ⌉ + 2 D ( m + 1 ) + 2 − 2 D ( m + 1 ) , which is linear in the number of Fourier coefficients, and linear in the number of qubits (Dn) despite the fact that the loaded function’s discretization is over exponentially many (2 Dn ) points. The FSL circuit consists of at most D n + 2 D ( m + 1 ) + 1 − 1 single-qubit and D n ( n + 1 ) / 2 + 2 D ( m + 1 ) + 1 − 3 D ( m + 1 ) − 2 two-qubit gates; we present a classical compilation algorithm with runtime O ( 2 3 D ( m + 1 ) ) to determine the FSL circuit for a given Fourier series. The FSL method allows for the highly accurate loading of complex-valued functions that are well-approximated by a Fourier series with finitely many terms. We report results from noiseless quantum circuit simulations, illustrating the capability of the FSL method to load various continuous 1D functions, and a discontinuous 1D function, on 20 qubits with infidelities of less than 10 −6 and 10 −3 , respectively. We also demonstrate the practicality of the FSL method for near-term quantum computers by presenting experiments performed on the Quantinuum H1-1 and H1-2 trapped-ion quantum computers: we loaded a complex-valued function on 3 qubits with a fidelity of over 95 % , as well as various 1D real-valued functions on up to 6 qubits with classical fidelities ≈99%, and a 2D function on 10 qubits with a classical fidelity ≈94%.

Physics↗

Reducing measurement costs by recycling the Hessian in adaptive variational quantum algorithms

Abstract Adaptive protocols enable the construction of more efficient state preparation circuits in variational quantum algorithms (VQAs) by utilizing data obtained from the quantum processor during the execution of the algorithm. This idea originated with Adaptive Derivative-Assembled Problem-Tailored variational quantum eigensolver (ADAPT-VQE), an algorithm that iteratively grows the state preparation circuit operator by operator, with each new operator accompanied by a new variational parameter, and where all parameters acquired thus far are optimized in each iteration. In ADAPT-VQE and other adaptive VQAs that followed it, it has been shown that initializing parameters to their optimal values from the previous iteration speeds up convergence and avoids shallow local traps in the parameter landscape. However, no other data from the optimization performed at one iteration is carried over to the next. In this work, we propose an improved quasi-Newton optimization protocol specifically tailored to adaptive VQAs. The distinctive feature in our proposal is that approximate second derivatives of the cost function are recycled across iterations in addition to optimal parameter values. We implement a quasi-Newton optimizer where an approximation to the inverse Hessian matrix is continuously built and grown across the iterations of an adaptive VQA. The resulting algorithm has the flavor of a continuous optimization where the dimension of the search space is augmented when the gradient norm falls below a given threshold. We show that this inter-optimization exchange of second-order information leads the approximate Hessian in the state of the optimizer to be consistently closer to the exact Hessian. As a result, our method achieves a superlinear convergence rate even in situations where the typical implementation of a quasi-Newton optimizer converges only linearly. Our protocol decreases the measurement costs in implementing adaptive VQAs on quantum hardware as well as the runtime of their classical simulation.

Ramôa, Mafalda (ORCID:0000000302187801)↗

Characterization and thermometry of dissipatively stabilized steady states

In this work we study the properties of dissipatively stabilized steady states of noisy quantum algorithms, exploring the extent to which they can be well approximated as thermal distributions, and proposing methods to extract the effective temperature T. We study an algorithm called the relaxational quantum eigensolver (RQE), which is one of a family of algorithms that attempt to find ground states and balance error in noisy quantum devices. In RQE, we weakly couple a second register of auxiliary ‘shadow’ qubits to the primary system in Trotterized evolution, thus engineering an approximate zero-temperature bath by periodically resetting the auxiliary qubits during the algorithm’s runtime. Balancing the infinite temperature bath of random gate error, RQE returns states with an average energy equal to a constant fraction of the ground state. We probe the steady states of this algorithm for a range of base error rates, using several methods for estimating both T and deviations from thermal behavior. In particular, we both confirm that the steady states of these systems are often well-approximated by thermal distributions, and show that the same resources used for cooling can be adopted for thermometry, yielding a fairly reliable measure of the temperature. These methods could be readily implemented in near-term quantum hardware, and for stabilizing and probing Hamiltonians where simulating approximate thermal states is hard for classical computers.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Lightweight jet reconstruction and identification as an object detection task

We apply object detection techniques based on deep convolutional blocks to end-to-end jet identification and reconstruction tasks encountered at the CERN large hadron collider (LHC). Collision events produced at the LHC and represented as an image composed of calorimeter and tracker cells are given as an input to a Single Shot Detection network. The algorithm, named PFJet-SSD performs simultaneous localization, classification and regression tasks to cluster jets and reconstruct their features. This all-in-one single feed-forward pass gives advantages in terms of execution time and an improved accuracy w.r.t. traditional rule-based methods. A further gain is obtained from network slimming, homogeneous quantization, and optimized runtime for meeting memory and latency constraints of a typical real-time processing environment. We experiment with 8-bit and ternary quantization, benchmarking their accuracy and inference latency against a single-precision floating-point. We show that the ternary network closely matches the performance of its full-precision equivalent and outperforms the state-of-the-art rule-based algorithm. Finally, we report the inference latency on different hardware platforms and discuss future applications.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Magnetic frame-dragging correction to the electromagnetic solution of a compact neutron star

ABSTRACT Neutron stars are usually modelled as spherical, rotating perfect conductors with a predominant intrinsic dipolar magnetic field anchored to their stellar crust. Due to their compactness, General Relativity corrections must be accounted for in Maxwell’s equations, leading to modified interior and exterior electromagnetic solutions. We present analytical solutions for slowly rotating magnetized neutron stars, taking into account the magnetic frame-dragging correction. For typical compactness values, i.e. Rs ∼ 0.5 [R*], we show that the new terms lead to a per cent order correction in the magnetic field orientation and strength compared to the case, with no magnetic frame-dragging correction. Also, we obtain a self-consistent redistribution of the surface azimuthal current. We verify the validity of the derived solution through two-dimensional particle-in-cell simulations of an isolated neutron star. Defining the azimuthal electric and magnetic field amplitudes during the transient phase as observables, we prove that the magnetic frame-dragging correction reduces the transient wave amplitude, as expected from the analytical solution. We show that simulations are more accurate and stable, when we include all first-order terms. The increased accuracy at lower spatiotemporal resolutions translates into a reduction in simulation runtimes.

Torres, R. (ORCID:0000000291820228)↗

Deeplasmid: deep learning accurately separates plasmids from bacterial chromosomes

Plasmids are mobile genetic elements that play a key role in microbial ecology and evolution by mediating horizontal transfer of important genes, such as antimicrobial resistance genes. Many microbial genomes have been sequenced by short read sequencers and have resulted in a mix of contigs that derive from plasmids or chromosomes. New tools that accurately identify plasmids are needed to elucidate new plasmid-borne genes of high biological importance. We have developed Deeplasmid, a deep learning tool for distinguishing plasmids from bacterial chromosomes based on the DNA sequence and its encoded biological data. It requires as input only assembled sequences generated by any sequencing platform and assembly algorithm and its runtime scales linearly with the number of assembled sequences. Deeplasmid achieves an AUC–ROC of over 89%, and it was more accurate than five other plasmid classification methods. Finally, as a proof of concept, we used Deeplasmid to predict new plasmids in the fish pathogen Yersinia ruckeri ATCC 29473 that has no annotated plasmids. Deeplasmid predicted with high reliability that a long assembled contig is part of a plasmid. Using long read sequencing we indeed validated the existence of a 102 kb long plasmid, demonstrating Deeplasmid's ability to detect novel plasmids.

59 BASIC BIOLOGICAL SCIENCES↗

Adaptive time stepping for the two-time integro-differential Kadanoff-Baym equations

The nonequilibrium Green's function gives access to one-body observables for quantum systems. Of particular interest are quantities such as density, currents, and absorption spectra which are important for interpreting experimental results in quantum transport and spectroscopy. We present an integration scheme for the Green's function's equations of motion, the Kadanoff-Baym equations (KBE), which is both adaptive in the time integrator step size and method order as well as the history integration order. We analyze the importance of solving the KBE self-consistently and show that adapting the order of history integral evaluation is important for obtaining accurate results. To examine the efficiency of our method, we compare runtimes to a state-of-the-art fixed time step integrator for several test systems and show an order of magnitude speedup at similar levels of accuracy. Published by the American Physical Society 2024

97 MATHEMATICS AND COMPUTING↗

Monte Carlo control loops for cosmic shear cosmology with DES Year 1 data

Weak lensing by large-scale structure is a powerful probe of cosmology and of the dark universe. This cosmic shear technique relies on the accurate measurement of the shapes and redshifts of background galaxies and requires precise control of systematic errors. Monte Carlo control loops (MCCL) is a forward modeling method designed to tackle this problem. It relies on the ultra fast image generator (UFig) to produce simulated images tuned to match the target data statistically, followed by calibrations and tolerance loops. Here, we present the first end-to-end application of this method, on the Dark Energy Survey (DES) Year 1 wide field imaging data. We simultaneously measure the shear power spectrum $C_ℓ$ and the redshift distribution $n(z)$ of the background galaxy sample. The method includes maps of the systematic sources, point spread function (PSF), an approximate Bayesian computation (ABC) inference of the simulation model parameters, a shear calibration scheme, and a fast method to estimate the covariance matrix. We find a close statistical agreement between the simulations and the DES Y1 data using an array of diagnostics. In a nontomographic setting, we derive a set of $C_ℓ$ and $n(z)$ curves that encode the cosmic shear measurement, as well as the systematic uncertainty. Following a blinding scheme, we measure the combination of $Ω_m$, $σ_8$, and intrinsic alignment amplitude $A_{IA}$, defined as $S_8D_{IA}=σ_8(Ω_m/0.3)^{0.5}D_{IA}$, where $D_{IA}=1-0.11(A_{IA}-1)$. We find $S_8D_{IA}=0.8954_{-0.039}^{+0.054}$, where systematics are at the level of roughly 60% of the statistical errors. We discuss these results in the context of earlier cosmic shear analyses of the DES Y1 data. Our findings indicate that this method and its fast runtime offer good prospects for cosmic shear measurements with future wide-field surveys.

79 ASTRONOMY AND ASTROPHYSICS↗