Engineering PapersSearch

SEARCH · Engineering Papers

Results for “precision”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Mixed-precision numerics in scientific applications: survey and perspectives

The explosive demand for artificial intelligence (AI) workloads has led to a significant increase in silicon area dedicated to lower-precision computations on recent high-performance computing hardware designs. However, mixed-precision capabilities, which can achieve performance improvements of up to 8x compared to double-precision in extreme compute-intensive workloads, remain largely untapped in most scientific applications. A growing number of efforts have shown that mixed-precision algorithmic innovations can deliver superior performance without sacrificing accuracy. These developments should prompt computational scientists to seriously consider whether their scientific modeling and simulation applications could benefit from the acceleration offered by new hardware and mixed-precision algorithms. In this survey, we (1) review progress across diverse scientific domains—fluid dynamics, weather and climate, quantum chemistry, and computational genomics—that have begun adopting mixed-precision strategies; (2) examine state-of-the-art algorithmic techniques such as iterative refinement, splitting and emulation schemes, and adaptive precision solvers; (3) assess their implications for accuracy, performance, and resource utilization; and (4) survey the emerging software ecosystem that enables mixed-precision methods at scale. We conclude with perspectives and recommendations on cross-cutting opportunities, domain-specific challenges, and the role of co-design between application scientists, numerical analysts, and computer scientists. Collectively, this survey underscores that mixed-precision numerics can reshape computational science by aligning algorithms with the evolving landscape of hardware capabilities.

Graphics processing units

Scaling the memory wall using mixed-precision - HPG-MxP on an exascale-class machine

Mixed-precision algorithms have been proposed as a way for scientific computing to benefit from some of the gains seen for AI on recent high performance computing (HPC) platforms. A few applications dominated by dense matrix operations have seen substantial speedups by utilizing low precision formats such as FP16. However, a majority of scientific simulation applications are memory bandwidth limited. Beyond preliminary studies, the practical gain from using mixed-precision algorithms on a given high-performance computing (HPC) system is largely unclear. The High Performance GMRES Mixed Precision (HPG-MxP) benchmark has been proposed to measure the useful performance of a HPC system on sparse matrix-based mixed-precision applications. In this work, we present an implementation of the HPG-MxP benchmark for an exascale system and describe our algorithm enhancements. We show for the first time a speedup of 1.6x using a combination of double- and single-precision keeping the same residual level on modern GPU-based supercomputers.

Kashi, Aditya [ORNL] (ORCID:0000000325893792)

Low Precision for Lower Energy Consumption: Preprint

Low-precision numeric types offer significant efficiency and energy benefits for computing applications. Mixed-precision algorithms, combining low and high precision types, maintain accuracy while improving performance. Despite advantages, there exist challenges on adapting existing mixed-precision algorithms to new technologies, such as new hardware architectures and new low-precision data types. This paper presents current challenges and opportunities to advance science in this domain targetting more energy efficient solutions.

energy efficiency

A GPU Accelerated Mixed‐Precision Finite Difference Informed Random Walker (FDiRW) Solver for Strongly Inhomogeneous Diffusion Problems

In nature, many complex multi‐physics coupling problems exhibit significant diffusivity inhomogeneity, where one process occurs several orders of magnitude faster than others temporally. Simulating rapid diffusion alongside slower processes demands intensive computational resources due to the necessity for small time steps. To address these computational challenges, we have developed an efficient numerical solver named Finite Difference informed Random Walker (FDiRW). In this study, we propose a GPU‐accelerated, mixed‐precision configuration for the FDiRW solver to maximize efficiency through GPU multi‐threaded parallel computation and lower precision computation. Numerical evaluation results reveal that the proposed GPU‐accelerated mixed‐precision FDiRW solver can achieve a 117× speedup over the CPU baseline, while an additional 1.75× speedup is achieved by employing lower precision GPU computation. Notably, for large model sizes, the GPU‐accelerated mixed‐precision FDiRW solver demonstrates strong scaling with the number of nodes used in simulation. When simulating radionuclide absorption processes by porous wasteform particles with a medium‐sized model of 192 × 192 × 192, this approach reduces the total computational time to 10 min, enabling the simulation of larger systems with strongly inhomogeneous diffusivity.

97 MATHEMATICS AND COMPUTING

Creating high-precision reference gas standards of 85Kr for groundwater age-dating

Absolute gas counting (AGC) was applied to two gas blends of 85Kr in argon-methane (P10) counting gas to establish a high-precision specific activity (Bq/cm3) reference value for characterizing 85Kr detection efficiency for groundwater age dating measurements. The AGC or length-compensated technique has been utilized by the metrology community for decades and is an accepted method for developing radioactive gas standards. The AGC capability at Pacific Northwest National Laboratory (PNNL) uses a set of nine unequal-length proportional counters with precisely-measured internal volumes, and a gas loading system with high-precision pressure and temperature sensors. A series of AGC measurements were collected at multiple pressures to determine the inverse pressure relationship (1/P) for 85Kr and define a wall-effect correction that accounts for events decaying into the detector wall and not depositing sufficient energy in the gas to be detected. In addition to the wall-effect, two additional corrections were evaluated and are discussed in detail. Specifically, the threshold effect which accounts for events deposited below the analysis threshold and a detection efficiency as a function of detector volume effect that was observed during analysis. A robust uncertainty model was developed using the Guide to the expression of Uncertainty in Measurements (GUM) approach. The combination of carefully scrutinized correction factors, precise measurements of pressure, temperature and detector volume, and robust counting statistics resulted in the determination of high-precision specific activity values with 0.50% or less total combined uncertainty for two Kr-in-P10 reference gas standards (KP10) that will enable new groundwater age-dating measurements at PNNL.

Kr-85

Efficient Mixed-Precision Matrix Factorization of the Inverse Overlap Matrix in Electronic Structure Calculations with AI-Hardware and GPUs

In recent years, a new kind of accelerated hardware has gained popularity in the artificial intelligence (AI) community which enables extremely high-performance tensor contractions in reduced precision for deep neural network calculations. In this article, we exploit Nvidia Tensor cores, a prototypical example of such AI-hardware, to develop a mixed precision approach for computing a dense matrix factorization of the inverse overlap matrix in electronic structure theory, S –1 . This factorization of S –1 , written as ZZT = S –1 , is used to transform the general matrix eigenvalue problem into a standard matrix eigenvalue problem. Here we present a mixed precision iterative refinement algorithm where Z is given recursively using matrix–matrix multiplications and can be computed with high performance on Tensor cores. To understand the performance and accuracy of Tensor cores, comparisons are made to GPU-only implementations in single and double precision. Additionally, we propose a nonparametric stopping criteria which is robust in the face of lower precision floating point operations. The algorithm is particularly useful when we have a good initial guess to Z, for example, from previous time steps in quantum-mechanical molecular dynamics simulations or from a previous iteration in a geometry optimization.

36 MATERIALS SCIENCE

High-precision measurement of the W boson mass with the CMS experiment

In the standard model of particle physics, the masses of the W and Z bosons, the carriers of the weak interaction, are uniquely related. A precise determination of their masses is important because quantum loops of heavy, undiscovered particles could modify this relationship. Although the Z mass is known to the remarkable precision of 22 parts per million (2.0 MeV), the W mass is known much less precisely. A global fit to measured electroweak observables predicts the W mass with 6 MeV uncertainty [1$-$3]. Reaching a comparable experimental precision would be a sensitive and fundamental test of the standard model, made even more urgent by a recent challenge to the global fit prediction by a measurement from the CDF Collaboration at the Fermilab Tevatron collider [4]. Here we report the measurement of the W mass by the CMS Collaboration at the CERN LHC, based on a large data sample of $W \to \mu \nu$ events collected in 2016 at the proton-proton collision energy of 13 TeV. The measurement exploits a high-granularity maximum likelihood fit to the kinematic properties of muons produced in W decays. By combining an accurate determination of experimental effects with marked in situ constraints of theoretical inputs, we reach a precise measurement of the W mass, of 80 360.2 $\pm$ 9.9 MeV, in agreement with the standard model prediction.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS

High precision measurements of the proton elastic electromagnetic form factors and their ratio at 𝑄 2 = 0.50, 2.64, 3.20, and 4.10 GeV 2

The advent of high-intensity, high-polarization electron beams led to significantly improved measurements of the ratio of the proton’s charge to electric form factors, 𝐺 𝐸⁢ 𝑝 ⁡/𝐺 𝑀⁢ 𝑝 . However, high-𝑄 2 measurements of this ratio yielded significant disagreement with extractions based on unpolarized scattering measurements, raising questions about the reliability of the measurements and consistency of the techniques. Jefferson Lab experiment E01-001 was designed to provide a high precision extraction of 𝐺 𝐸⁢ 𝑝 ⁡/𝐺 𝑀⁢ 𝑝 from unpolarized cross-section measurements using a modified version of the Rosenbluth separation technique to allow for a more precise comparison with polarization data. Rosenbluth separations involve precise measurements of the angular dependence of the elastic 𝑒−𝑝 cross section at fixed momentum transfer, 𝑄 2 . Conventional Rosenbluth separations detect the scattered electron, requiring the comparisons of measurements with very different detected electron energy and rate for electrons at different angles. Our ‘‘super-Rosenbluth’’ measurement detected the struck proton, rather than the scattered electron to extract the elastic 𝑒−𝑝 cross section. This yielded a fixed momentum for the detected particle and dramatically reduced variation of the cross section with angle, significantly reducing rate- and momentum-dependent corrections and uncertainties. We measure the cross section vs angle with high relative precision, allowing for extremely high precision extractions of 𝐺 𝐸⁢ 𝑝 ⁡/𝐺 𝑀⁢ 𝑝 at 𝑄 2 = 2.64, 3.20, and 4.10 GeV 2 . Our results are consistent with traditional Rosenbluth extractions, but with much smaller corrections and systematic uncertainties, comparable to the uncertainties from polarization measurements. Our data confirm the discrepancy between Rosenbluth and polarization extractions of the proton form factor ratio using an improved Rosenbluth extraction that yields smaller and less-correlated uncertainties than those typical of previous Rosenbluth extractions. Here, we compare our results to calculations of two-photon exchange effects and find that the observed discrepancy can be relatively well explained by such effects.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS

Clock Precision beyond the Standard Quantum Limit at 10 −18 Level

Optical atomic clocks with unrivaled precision and accuracy have advanced the frontier of precision measurement science and opened new avenues for exploring fundamental physics. A fundamental limitation on clock precision is the standard quantum limit (SQL), which stems from the uncorrelated projection noise of each atom. State-of-the-art optical lattice clocks interrogate large ensembles to minimize the SQL, but density-dependent frequency shifts pose challenges to scaling the atom number. The SQL can be surpassed, however, by leveraging entanglement, though it remains an open problem to achieve quantum advantage from spin squeezing at state-of-the-art stability levels. Here, we demonstrate clock performance beyond the SQL, achieving a fractional frequency precision of 1.1 × 10 −18 for a single spin-squeezed clock. With cavity-based quantum nondemolition measurements, we prepare two spin-squeezed ensembles of ∼30 000 strontium atoms confined in a two-dimensional optical lattice. A synchronous clock comparison with an interrogation time of 61 ms achieves a metrological improvement of 2.0(2) dB beyond the SQL, after correcting for state preparation and measurement errors. These results establish the most precise entanglement-enhanced clock to date and offer a powerful platform for exploring the interplay of gravity and quantum entanglement.

cavity quantum electrodynamics

Quantifying the impact of precision errors on quantum approximate optimization algorithms

The quantum approximate optimization algorithm (QAOA) is a hybrid quantum-classical algorithm that seeks to achieve approximate solutions to optimization problems by iteratively alternating between intervals of controlled quantum evolution. Here, we examine the effect of analog precision errors on QAOA performance from the perspective of both algorithmic training and performance guarantees. Leveraging cumulant expansions, we recast the faulty QAOA as a control problem in which precision errors are expressed as multiplicative control noise and derive bounds on the performance of QAOA. We show using both analytical techniques and numerical simulations that fixed precision implementations of QAOA circuits are subject to an exponential degradation in performance dependent upon the number of optimal QAOA layers and magnitude of the precision error. Despite this significant reduction, we show that it is possible to mitigate precision errors in QAOA via digitization of the variational parameters at the cost of increasing circuit depth.

quantum algorithms

Precision Agriculture using Networks of Degradable Analytical Sensors (PANDAS) (Final Technical Report)

Precision agriculture, where sensing of soil, environment and crop conditions are used to precisely synchronize inputs (such as water and fertilizer) to crop needs enhances input use efficiency. This can improve yields and farm profitability while mitigating environmental losses, improving soil carbon content and substantially decreasing energy use for food, feed and fuel crops. Unfortunately, farmers are not yet able to harness the full potential of these management technologies as there is a lack of available management information, and there is therefore a need for sensors that are able to economically measure spatio-temporal variability in soil and crop properties of extremely heterogeneous farm fields precisely at high resolution and at low cost. Real-time, in-situ monitoring of agricultural soil conditions is today carried out using devices that limit the total number of nodes that can be used economically to typically one per acre or less. Higher spatio-temporal resolution sensing would enable more precise agricultural input optimization, with significant benefits to the farmer and the environment. In order to address this issue, this project focused on developing additively manufactured, biodegradable, soil sensors with predicted costs of < $\$$1 per unit to monitor crop inputs (such as water and fertilizer) that predictably, harmlessly degrade away into the soil when no longer needed. These sensor nodes should be easy to place, accurately and continuously monitor soil and crop conditions for an entire season, be read remotely using existing farm equipment, require no ongoing maintenance, not impede farm operations and produce no persistent waste. This approach could enable a >100× increase in information density over current solutions for precision farming of row and other crops, and lead to significant reductions in input energy use and provide increased yield for biofuel crops. Over the course of this project the team at the University of Colorado Boulder, University of California Berkeley, and Colorado State University/Kansas State University investigated a wide range of printable biodegradable electronic materials and sensor designs for determining soil moisture and soil nitrate concentration. These efforts expanded the available materials set for printed soil degradable electronic materials, particularly for conductors, enabling high conductivity and stability. Printed soil moisture and nitrate sensors with suitable sensitivity and selectivity were developed and characterized. Low power and passive wireless electronic systems were integrated with the soil sensors, and testing was carried out with completed sensors to understand their functionality under agricultural conditions. Additionally, other sensor types enabled by the biodegradable materials set created during this project, such as soil microbial activity sensors, were also developed and demonstrated. Project outputs include 10 peer reviewed publications, 4 patent applications, 21 technical presentations, 3 PhD thesis, 10 media reports, 8 additional grants worth over $\$$6M, and the formation of 3 start-up companies.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Precision of ENDF and ENDL Formatted Data Files

The purpose is to ensure that today’s processing codes produced output to meet today’s accuracy needs. Since 1958 ENDL and about 1965 ENDF have each used a text format to define nuclear and atomic data in 11 columns for each data field. When these formats originated this was judged to be adequate to reproduce the accuracy of data at the time and to meet the needs of our applications. When these formats originated the dominant computer language of the day was FORTRAN and if written using an E11.4 format it would include only 4 or 5 digits of precision, e.g., 0.1234E-03 or 1.2345E-02, varying from one computer/system to another the result was not even unique. In the case of ENDF the 4 digit precision was not even adequate to uniquely define the atomic weight of the target, e.g., U238 = 92238 = 0.9224E+5 = WRONG! From its inceptions the ENDF format had a precision problem. One of my first tasks when in 1967 fresh out of graduate school I joined what later became the National Nuclear Data Center (NNDC), was to address this precision problem. By working with ENDF producers and users throughout the U.S. we verified, 1) E or D is not required to define FORTRAN readable numbers, e.g., E+4 or +4 are both o.k. 2) With ENDF energy eV and cross section in barns, 2 digit exponents are almost never required. 3) Since energy is never negative we could use the first of the 11 columns for a digit. Knowing this allowed us to produce ENDF/B-II to 6 or 7 digit precision, e.g., blank, decimal point, 2 or 3 digit exponent, e.g., ^1.23456-12 or ^1.234567-3. Below is an example of the actual ENDF/B-II data released. Note, the date 1970 and the atomic weight, ZA, uniquely defined to 6-digit accuracy, ^9.22350+ 4.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS

Low Precision and Efficient Programming Languages for Sustainable AI: Final Report for the Summer Project of 2024

This document contains all relevant material generated during the authors' summer internship at NREL in 2024. This report shows how to improve energy efficiency of a few code samples by using low-precision data types combined with mixed-precision algorithms. The main applications considered here are (i) linear system solvers using mixed precision, and (ii) neural networks using mixed precision. This report also discusses how programming languages affect energy consumption of algorithms, energy metrics for a code and tools, and the available current software and hardware infrastructure.

97 MATHEMATICS AND COMPUTING

The electroweak precision constraints of the 2HDM+S

The 2HDM+S is the singlet extension of the Two-Higgs-Doublets Model (2HDM). The singlet field and its mixing with the 2HDM Higgs sector lead to new contributions to the electroweak precision observables, in particular, the oblique parameters. In this paper, we performed a systematic study of the impacts of each mixing angle to the oblique parameters. We adopted the mixing angles and physical Higgs masses as our parameters, which allows a mapping when specific symmetry structure of the Higgs potential and various theoretical considerations are taken into account. We identify five benchmark cases, where at most one mixing angle is nonzero and analyze the 95\% C.L. allowed parameter space by the oblique parameters. In the alignment limit of the 2HDM, we find that other than the usual mass relations of $m_H\sim m_{H^\pm}$ or $m_A\sim m_{H^\pm}$, electroweak precision measurements also impose an upper limit on the neutral Higgs masses. In the cases with nonzero singlet mixing with the 2HDM Higgses $H$ or $A$, we find approximate mass relations of $c^2_{\alpha_{HS}} m_{H} + s^2_{\alpha_{HS}}m_{h_S} = m_{H^\pm}$ or $c^2_{\alpha_{AS}} m_{A} + s^2_{\alpha_{AS}}m_{A_S} = m_{H^\pm}$. Those relations are universal to the 2HDM+S models, with or without further symmetry assumption. We also study the non-alignment limit of the 2HDM+S, which typically has tighter constraints on the masses and mixing angles. At the end, we examine the complementarity between the electroweak precision analyses and the Higgs coupling precision measurements.

2HDM

Photocathode characterisation for robust PICOSEC Micromegas precise-timing detectors

The PICOSEC Micromegas detector is a precise-timing gaseous detector based on a Cherenkov radiator coupled with a semi-transparent photocathode and a Micromegas amplifying structure, targeting a time resolution of tens of picoseconds for minimum ionising particles. Initial single-pad prototypes have demonstrated a time resolution below σ = 25 ps, prompting ongoing developments to adapt the concept for High Energy Physics applications, where sub-nanosecond precision is essential for event separation, improved track reconstruction and particle identification. The achieved performance is being transferred to robust multi-channel detector modules suitable for large-area detection systems requiring excellent timing precision. To enhance the robustness and stability of the PICOSEC Micromegas detector, research on robust carbon-based photocathodes, including Diamond-Like Carbon (DLC) and Boron Carbide (B 4 C), is pursued. Results from prototypes equipped with DLC and B 4 C photocathodes exhibited a time resolution of σ ≈ 32 ps and σ ≈ 34.5 ps, respectively. Efforts dedicated to improve detector robustness and stability enhance the feasibility of the PICOSEC Micromegas concept for large experiments, ensuring sustained performance while maintaining excellent timing precision.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND

Understanding Mixed Precision GEMM with MPGemmFI: Insights into Fault Resilience

Emerging deep learning workloads urgently need fast general matrix multiplication (GEMM). Thus, one of the critical features of machine-learning-specific accelerators such as NVIDIA Tensor Cores, AMD Matrix Cores, and Google TPUs is the support of mixed-precision enabled GEMM. For DNN models, lower-precision FP data formats and computation offer acceptable correctness but significant performance, area, and memory footprint improvement. While promising, the mixed-precision computation on error resilience remains unexplored. To this end, we develop a fault injection framework that systematically injects fault into the mixed-precision computation results. We investigate how the faults affect the accuracy of machine learning applications. Based on the characteristics of error resilience, we offer lightweight error detection and correction solutions that significantly improve the overall model accuracy by 75% if the models experience hardware faults. The solutions can be efficiently integrated into the accelerator's pipelines.

Fang, Bo