Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “finite precision”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

A GPU Accelerated Mixed‐Precision Finite Difference Informed Random Walker (FDiRW) Solver for Strongly Inhomogeneous Diffusion Problems

In nature, many complex multi‐physics coupling problems exhibit significant diffusivity inhomogeneity, where one process occurs several orders of magnitude faster than others temporally. Simulating rapid diffusion alongside slower processes demands intensive computational resources due to the necessity for small time steps. To address these computational challenges, we have developed an efficient numerical solver named Finite Difference informed Random Walker (FDiRW). In this study, we propose a GPU‐accelerated, mixed‐precision configuration for the FDiRW solver to maximize efficiency through GPU multi‐threaded parallel computation and lower precision computation. Numerical evaluation results reveal that the proposed GPU‐accelerated mixed‐precision FDiRW solver can achieve a 117× speedup over the CPU baseline, while an additional 1.75× speedup is achieved by employing lower precision GPU computation. Notably, for large model sizes, the GPU‐accelerated mixed‐precision FDiRW solver demonstrates strong scaling with the number of nodes used in simulation. When simulating radionuclide absorption processes by porous wasteform particles with a medium‐sized model of 192 × 192 × 192, this approach reduces the total computational time to 10 min, enabling the simulation of larger systems with strongly inhomogeneous diffusivity.

97 MATHEMATICS AND COMPUTING↗

Mixed precision s –step Lanczos and conjugate gradient algorithms

Compared to the classical Lanczos algorithm, the s-step Lanczos variant has the potential to improve performance by asymptotically decreasing the synchronization cost per iteration. However, this comes at a price; despite being mathematically equivalent, the s-step variant may behave quite differently in finite precision, potentially exhibiting greater loss of accuracy and slower convergence relative to the classical algorithm. It has previously been shown that the errors in the s-step version follow the same structure as the errors in the classical algorithm, but are amplified by a factor depending on the square of the condition number of the O(s)-dimensional Krylov bases computed in each outer loop. As the condition number of these s-step bases grows (in some cases very quickly) with s, this limits the s values that can be chosen and thus can limit the attainable performance. In this work, we show that if a select few computations in s-step Lanczos are performed in double the working precision, the error terms then depend only linearly on the conditioning of the s-step bases. This has the potential for drastically improving the numerical behavior of the algorithm with little impact on per-iteration performance. Our numerical experiments demonstrate the improved numerical behavior possible with the mixed precision approach, and also show that this improved behavior extends to mixed precision s-step CG. Here, we present preliminary performance results on NVIDIA V100 GPUs that show that the overhead of extra precision is minimal if one uses precisions implemented in hardware.

97 MATHEMATICS AND COMPUTING↗

ZFP: A compressed array representation for numerical computations

HPC trends favor algorithms and implementations that reduce data motion relative to FLOPS. We investigate the use of lossy compressed data arrays in place of traditional IEEE floating point arrays to store the primary data of calculations. Simulation is fundamentally an exercise in controlled approximation, and error introduced by finite-precision arithmetic (or lossy compression) is just one of several sources of error that need to be managed to ensure sufficient accuracy in a computed result. We describe ZFP, a compressed numerical format designed for in-memory storage of multidimensional arrays, and summarize theoretical results that demonstrate that the error of repeated lossy compression can be bounded and controlled. Furthermore, we establish a relationship between grid resolution and compression-induced errors and show that, contrary to conventional floating point, ZFP reduces finite-difference errors with finer grids. We present example calculations that demonstrate data reduction by 4x or more with negligible impact on solution accuracy. Our results further demonstrate several orders-of-magnitude increase in accuracy using ZFP over IEEE floating point and Posits for the same storage budget.

Lindstrom, Peter↗

Noise effects on Padé approximants and conformal maps

Here, we analyze the properties of Padé and conformal map approximants for functions with branch points, in the situation where the expansion coefficients are only known with finite precision or are subject to noise. We prove that there is a universal scaling relation between the strength of the noise and the expansion order at which Padé or the conformal map breaks down. We illustrate this behavior with some physically relevant model test functions and with two non-trivial physical examples where the relevant Riemann surface has complicated structure.

97 MATHEMATICS AND COMPUTING↗

Distributed Inference with Sparse and Quantized Communication

Here, we consider the problem of distributed inference where agents in a network observe a stream of private signals generated by an unknown state, and aim to uniquely identify this state from a finite set of hypotheses. We focus on scenarios where communication between agents is costly, and takes place over channels with finite bandwidth. To reduce the frequency of communication, we develop a novel event-triggered distributed learning rule that is based on the principle of diffusing low beliefs on each false hypothesis. Building on this principle, we design a trigger condition under which an agent broadcasts only those components of its belief vector that have adequate innovation, to only those neighbors that require such information. We prove that our rule guarantees convergence to the true state exponentially fast almost surely despite sparse communication, and that it has the potential to significantly reduce information flow from uninformative agents to informative agents. Next, to deal with finite-precision communication channels, we propose a distributed learning rule that leverages the idea of adaptive quantization. We show that by sequentially refining the range of the quantizers, every agent can learn the truth exponentially fast almost surely, while using just 1 bit to encode its belief on each hypothesis. For both our proposed algorithms, we rigorously characterize the trade-offs between communication-efficiency and the learning rate.

42 ENGINEERING↗

Identification of $^{3}$He–$^{3}$H clusters in the $^{6}$Li+$^{89}$Y experiment using particle-$\gamma$ coincidence measurement

The 6 Li+ 89 Y experiment was performed to explore the reaction mechanism induced by a weakly bound nucleus 6 Li and its cluster configuration. Here, the particle-$\gamma$ coincidence method was used to identify the different reaction channels. The $\gamma$-rays coincident with 3 He/ 3 H indicate that the 3 H/ 3 He stripping reaction plays a significant role in the formation of Zr/Nb isotopes. The obtained results support the existence of a 3 He- 3 H cluster in 6 Li. Direct and sequential transfer reactions are adequately discussed, and the FRESCO code is used to perform precise finite-range cyclic redundancy check calculations. In the microscopic calculation, direct cluster transfer is more predominant than sequential transfer in 3 H transfer. However, the direct cluster transfer is of comparable magnitude to the sequential transfer in the 3 He transfer.

CRC calculations↗

Interplay Between Time and Energy in Bosonic Noisy Quantum Metrology

Quantum entanglement and coherence often allow for protocols that outperform classical ones in estimating a system’s parameter. When using infinite-dimensional probes (such as a bosonic mode), one could, in principle, obtain infinite precision in a finite time for both classical and quantum protocols, which makes it hard to quantify potential quantum advantage. However, such a situation is unphysical, as it would require infinite resources, so one needs to impose some additional constraint: typically the average energy employed by the probe is finite. Here we treat both energy and time as a resource, showing that, in the presence of noise, there is a nontrivial interplay between the average energy and the time devoted to the estimation. Our results are valid for the most general metrological schemes (e.g., adaptive schemes, which may involve entanglement with external ancillae or any kind of continuous measurement). We apply recently derived precision bounds for all parameters characterizing the paradigmatic case of a bosonic mode, subject to Lindbladian noise. We show how the time employed in the estimation should be partitioned in order to achieve the best possible precision. In most cases, the optimal performance may be obtained without the necessity of adaptivity or entanglement with ancilla. We compare results with classical strategies. Interestingly, for temperature estimation, applying a fast-prepare-and-measure protocol with Fock states provides better scaling with the number of photons than any classical strategy.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Precise 3D reactor core calculation using spherical harmonics and discontinuous Galerkin finite element methods

We study the use of P{sub N} method in angle and discontinuous Galerkin is space to solve 3D neutron transport problem. P{sub N} method consists in developing the angular flux on truncated spherical harmonics basic. In this paper, we couple this method with the discontinuous finite elements in space to obtain a complete discretization of the multigroup neutron transport equation. To investigate its precision, the method was applied to Takeda and C5G7 benchmark problems. These calculations point out that the proposed P{sub N}-DG method is capable of producing accurate solutions in small computational time, and that it is able to handle complex 3D geometries. (authors)

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Microcanonical Kinetics of Water-Mediated Proton Transfer in 4ABAH + ·(H 2 O) n = 4–6 Clusters (ABA = Aminobenzoic Acid): A Model System for Size-Dependent Relaxation to Ergodic Behavior

Here, we leverage the unique properties of the 4ABAH + · (H 2 O) n clusters (ABA = 4-aminobenzoic acid, n = 4−6) to quantitatively address how a finite, isolated system evolves into an ergodic condition starting from localized arrangements in configuration space. This system adopts two distinct structural isomers in which water molecules cluster around the cationic centers of its two protomers with widely separated positive charge centers. These isomers arise from excess proton attachment to either the acid (O) or amino (N) group on opposite sides of the benzene ring. Both forms are captured and kinetically trapped using cryogenic ion methods and then selectively vibrationally excited through their mutually exclusive IR bands involving NH and OH stretching fundamentals. Because the IR excitation lies below the water binding energy, the system can evolve to explore slow, rare events that lead to the interconversion between the two isomers. The rates of these intracluster reactions are determined by using a pump−probe scheme involving ∼5 ns IR pump and UV probe lasers. The rates occur on the microsecond time scale, leading to steady state populations of the isomers, thus revealing the cluster size-dependent fractionation between the two species at microcanonical equilibrium. The steady state distributions are correlated with the expected trend in the cluster size-dependent reaction energetics, which are in turn consistent with changes in the relative densities of states of the two species. These results thus provide an unusually clear example in which complex, protic-solvent-mediated chemical transformations are captured within a finite system at a precisely determined internal energy.

Rana, Abhijit [Yale Univ., New Haven, CT (United S↗

Scale setting of SU⁡(𝑁) Yang–Mills theory, topology and large-𝑁 volume independence

We set the scale of SU⁡(𝑁) Yang-Mills theories for 𝑁 =3, 5, 8 and in the large-𝑁 limit via gradient flow, as a first step towards the computation of the large-𝑁 Λ-parameter using step scaling. We adopt twisted boundary conditions to achieve large-𝑁 volume reduction and the Parallel Tempering on Boundary Conditions algorithm to tame topological freezing. This setup allows accurate determinations of the gradient-flow scales down to lattice spacings as fine as ∼0.025 fm for all the explored values of 𝑁, a regime that has never been reached with ergodic algorithms. Moreover, we are able to precisely estimate the finite-size systematics related to topological freezing, and to show the suppression of finite-volume effects expected by virtue of large-𝑁 twisted volume reduction.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Precisely computing phonons via irreducible derivatives

Computing phonons from first principles is typically considered a solved problem, yet inadequacies in existing techniques continue to yield deficient results in systems with sensitive phonons. Here, in this study, we circumvent this issue using the lone irreducible derivative (LID) and bundled irreducible derivative (BID) approaches to computing phonons via finite displacements, where the former optimizes precision via energy derivatives and the latter provides the most efficient algorithm using force derivatives. A condition number optimized basis for BID is derived which guarantees the minimum amplification of error. Additionally, a hybrid LID-BID approach is formulated, in which select irreducible derivatives computed using LID replace BID results. We illustrate our approach on two prototypical systems with sensitive phonons: the shape memory alloy AuZn and metallic lithium. Comparing our resulting phonons in the aforementioned crystals to calculations in the literature reveals nontrivial inaccuracies. Our approaches can be fully automated, making them well suited for both niche systems of interest and high-throughput approaches.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Explicit block encodings of boundary value problems for many-body elliptic operators

Simulation of physical systems is one of the most promising use cases of future digital quantum computers. In this work we systematically analyze the quantum circuit complexities of block encoding the discretized elliptic operators that arise extensively in numerical simulations for partial differential equations, including high-dimensional instances for many-body simulations. When restricted to rectangular domains with separable boundary conditions, we provide explicit circuits to block encode the many-body Laplacian with separable periodic, Dirichlet, Neumann, and Robin boundary conditions, using standard discretization techniques from low-order finite difference methods. To obtain high-precision, we introduce a scheme based on periodic extensions to solve Dirichlet and Neumann boundary value problems using a high-order finite difference method, with only a constant increase in total circuit depth and subnormalization factor. We then present a scheme to implement block encodings of differential operators acting on more arbitrary domains, inspired by Cartesian immersed boundary methods. We then block encode the many-body convective operator, which describes interacting particles experiencing a force generated by a pair-wise potential given as an inverse power law of the interparticle distance. This work provides concrete recipes that are readily translated into quantum circuits, with depth logarithmic in the total Hilbert space dimension, that block encode operators arising broadly in applications involving the quantum simulation of quantum and classical many-body mechanics.

Kharazi, Tyler [University of California, Berkeley↗

Dynamics of the O ( 4 ) critical point in QCD: Critical pions and diffusion in model G

We present a detailed study of the finite momentum dynamics of the O ( 4 ) critical point of QCD, which lies in the dynamic universality class of “model G.” The critical scaling of the model is analyzed in multiple dynamical channels. For instance, the finite momentum analysis allows us to precisely extract the pion dispersion curve below the critical point. The pion velocity is in striking agreement with the predictions relation and static universality. The pion damping rate and velocity are both consistent with the dynamical critical exponent ζ = 3 / 2 of model G. Similarly, although the critical amplitude for the diffusion coefficient of the conserved O ( 4 ) charges is small, it is clearly visible both in the restored phase and with finite explicit symmetry breaking, and its dynamical scaling is again consistent with ζ = 3 / 2 . We determine a new set of universal dynamical critical amplitude ratios relating the diffusion coefficient to a suitably defined order parameter relaxation time. We also show that in a finite volume simulation, the chiral condensate diffuses on the coset manifold in a manner consistent with dynamical scaling, and with a diffusion coefficient that is determined by the transport coefficients of hydrodynamic pions. Finally, the amplitude ratios (together with other nonuniversal amplitudes also reported here) compile all relevant information for further studies of model G both in and out of equilibrium. Published by the American Physical Society 2024

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Adaptive finite differencing in high accuracy electronic structure calculations

Abstract A multi-order Adaptive Finite Differencing (AFD) method is developed for the kinetic energy operator in real-space, grid-based electronic structure codes. It uses atomic pseudo orbitals produced by the corresponding pseudopotential codes to optimize the standard finite difference (SFD) operators for improved precision. Results are presented for a variety of test systems and Bravais lattice types, including the well-known Δ test for 71 elements in the periodic table, the Mott insulator NiO, and borax decahydrate, which contains covalent, ionic, and hydrogen bonds. The tests show that an 8th-order AFD operator leads to the same average Δ value as that achieved by plane-wave codes and is typically far more accurate and has a much lower computational cost than a 12th-order SFD operator. The scalability of real-space electronic calculations is demonstrated for a 2016-atom NiO cell, for which the computational time decreases nearly linearly when scaled from 18 to 144 CPU-GPU nodes.

Briggs, E. L. (ORCID:0000000343983492)↗

Interactions of two and three mesons including higher partial waves from lattice QCD

We study two- and three-meson systems composed either of pions or kaons at maximal isospin using Monte Carlo simulations of lattice QCD. Utilizing the stochastic LapH method, we are able to determine hundreds of two- and three-particle energy levels, in nine different momentum frames, with high precision. We fit these levels using the relativistic finite-volume formalism based on a generic effective field theory in order to determine the parameters of the two- and three-particle K-matrices. We find that the statistical precision of our spectra is sufficient to probe not only the dominant s-wave interactions, but also those in d waves. In particular, we determine for the first time a term in the three-particle K-matrix that contains two-particle d waves. We use three N f = 2 + 1 CLS ensembles with pion masses of 200, 280, and 340 MeV. This allows us to study the chiral dependence of the scattering observables, and compare to the expectations of chiral perturbation theory.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Gaussian Process Regression under Computational and Epistemic Misspecification

Gaussian process regression is a classical kernel method for function estimation and data interpolation. In large data applications, computational costs can be reduced using low-rank or sparse approximations of the kernel. This paper investigates the effect of such kernel approximations on the interpolation error. We introduce a unified framework to analyze Gaussian process regression under important classes of computational misspecification: Karhunen-Loève expansions that result in low-rank kernel approximations, multiscale wavelet expansions that induce sparsity in the covariance matrix, and finite element representations that induce sparsity in the precision matrix. Furthermore, our theory also accounts for epistemic misspecification in the choice of kernel parameters.

Gaussian process regression↗

Topological protection of coherence in disordered open quantum systems

Here, we consider topological protection mechanisms in dissipative quantum systems in the presence of quenched disorder, with the intent to prolong the coherence time of a fiducial qubit. The qubit is part of a network of other qubits and dissipative cavities whose coupling parameters are tunable, such that topological edge states can be stabilized. The evolution of the fiducial qubit is entirely determined by a non-Hermitian Hamiltonian which thus emerges from a bona fide physical process. Even in the presence of disorder, a winding number W can be defined and evaluated in real space, as long as certain symmetries are preserved. Hence we can construct the topological phase diagrams of noisy open quantum models, such as the non-Hermitian disordered Su-Schrieffer-Heeger dimer model and a trimer model that includes longer-range couplings. For finite-size systems we find that there are precisely W modes localized at one end of the chain. In such topological phases the qubit's coherence lifetime is exponentially large in the system size. In the presence of competing disorder parameters, interesting reentrance phenomena of topologically nontrivial sectors are observed. This means that in certain parameter regions, increasing disorder drastically increases the coherence time of the fiducial qubit.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗