Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “mixed precision”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

End-to-end codesign of Hessian-aware quantized neural networks for FPGAs

Here, we develop an end-to-end workflow for the training and implementation of co-designed neural networks (NNs) for efficient field-programmable gate array (FPGA) hardware. Our approach leverages Hessian-aware quantization of NNs, the Quantized Open Neural Network Exchange intermediate representation, and the hls4ml tool flow for transpiling NNs into FPGA firmware. This makes efficient NN implementations in hardware accessible to nonexperts in a single open sourced workflow that can be deployed for real-time machine-learning applications in a wide range of scientific and industrial settings. We demonstrate the workflow in a particle physics application involving trigger decisions that must operate at the 40-MHz collision rate of the CERN Large Hadron Collider (LHC). Given the high collision rate, all data processing must be implemented on FPGA hardware within the strict area and latency requirements. Based on these constraints, we implement an optimized mixed-precision NN classifier for high-momentum particle jets in simulated LHC proton-proton collisions.

47 OTHER INSTRUMENTATION↗

GenSLMs: Genome-scale language models reveal SARS-CoV-2 evolutionary dynamics

We seek to transform how new and emergent variants of pandemic-causing viruses, specifically SARS-CoV-2, are identified and classified. By adapting large language models (LLMs) for genomic data, we build genome-scale language models (GenSLMs) which can learn the evolutionary landscape of SARS-CoV-2 genomes. By pre-training on over 110 million prokaryotic gene sequences and fine-tuning a SARS-CoV-2-specific model on 1.5 million genomes, we show that GenSLMs can accurately and rapidly identify variants of concern. Thus, to our knowledge, GenSLMs represents one of the first whole-genome scale foundation models which can generalize to other prediction tasks. We demonstrate scaling of GenSLMs on GPU-based supercomputers and AI-hardware accelerators utilizing 1.63 Zettaflops in training runs with a sustained performance of 121 PFLOPS in mixed precision and peak of 850 PFLOPS. We present initial scientific insights from examining GenSLMs in tracking evolutionary dynamics of SARS-CoV-2, paving the path to realizing this on large biological data.

Zvyagin, Maxim↗

Matrix Product (GEMM) Performance Data from GPUs

Timing data for mixed precision GEMM matrix product operations on several GPU models, including NVIDIA V100 and A100, AMD MI100 and Intel P580. Also data from machine learning model training on this data using Scikit-learn.

97 MATHEMATICS AND COMPUTING↗

AmeriFlux FLUXNET-1F US-ARM ARM Southern Great Plains site- Lamont

This is the AmeriFlux Management Project (AMP) created FLUXNET-1F version of the carbon flux data for the site US-ARM ARM Southern Great Plains site- Lamont. This is the FLUXNET version of the carbon flux data for the site US-ARM ARM Southern Great Plains site- Lamont produced by applying the standard ONEFlux (1F) software. Site Description - Central facility tower crop field (winter wheat, corn, soy, alfalfa). The site also has continuous measurements of precise mixing ratios of CO2, CH4, CO, N2O, and isotopic ratio of 13CO2. In addition to 4 m system, there are sonic anemometers at 25 m and 60 m.

Biraud, Sebastien↗

Port and optimize the CEED software stack to Aurora/Frontier EA (ECP Milestone Report)

The goal of this milestone was to port the CEED software stack, including Nek, MFEM and libCEED to the Frontier and Aurora early access hardware, work on optimizing the performance on AMD and Intel GPUs, and demonstrate impact in CEED-enabled ECP applications. As part of this milestone, we also performed a number of other activities, including continued optimizations for CEED applications at scale on Summit, performance evaluation on non-ECP hardware, including the Fugaku’s A64FX chip and the NVIDIA A100 GPU architecture, exploring the use of just-in-time compilations in applications at scale and potential of mixed precision optimizations in the CEED discretization and solver algorithms, and more. During the milestone period, the CEED team released new versions of 6 of its packages, including major releases of MFEM, NekRS and OCCA. We also organized the fifth CEED Annual meeting (CEED5AM) which included nearly 100 researchers from national labs, universities and industry.

97 MATHEMATICS AND COMPUTING↗

User Manual - HydraGNN v5.0: Distributed Implementation of Multi-Tasking Graph Neural Networks

This document serves as the user manual for HydraGNN v5.0, a scalable graph neural network (GNN) architecture for simultaneous prediction of multiple target properties using multi-task learning (MTL). This version of HydraGNN has been developed primarily to support the development, training, and deployment of predictive graph-based deep learning (DL) models for atomistic materials modeling. HydraGNN is templated over 13 message-passing policies, including invariant models (GIN, PNA, PNAPlus, GAT, MFC, CGCNN, SAGE, SchNet, DimeNet) and equivariant models (EGNN, PNAEq, PAINN, MACE), and supports distributed training via distributed data parallelism (DDP), DeepSpeed, and Fully Sharded Data Parallelism (FSDP) on leadership-class supercomputers. Although HydraGNN can be applied to problems beyond atomistic materials modeling, its current use is confined to homogeneous graphs. Additional capabilities include machine-learned interatomic potentials with energy-conserving forces, General, Powerful, and Scalable Graph Transformer (GraphGPS) global attention, periodic boundary conditions, hyperparameter optimization, mixed-precision training, and uncertainty quantification.

97 MATHEMATICS AND COMPUTING↗

Fine-Grained Power and Energy Attribution on AMD GPU/APU-Based Exascale Nodes

Modern exascale GPU- and APU-based systems provide multiple power and energy sensors, but differences in scope, update rate, timing, and filtering complicate the attribution of short-lived accelerator activity. This paper presents a methodology to characterize and correct these effects on Cray EX systems with AMD Instinct MI250X GPUs (Frontier) and MI300A APUs (Portage). Using controlled square-wave workloads, we quantify update intervals, delay, aliasing, and variability across up to 512 GPUs and 480 APUs with on-chip (rocm-smi/amd-smi) and off-chip Cray Power Management sensors. We reconstruct power from cumulative energy counters to achieve faster response times, validate it against on-chip, off-chip, and node-level sensors, and integrate the resulting streams into a Score-P/PAPI-based tool for time-aligned, phase-level attribution. Applied to rocHPL, rocHPL-MxP, and HPG-MxP, the method separates energy savings due to reduced runtime from changes in power. Mixed precision reduces node energy on Frontier by 79% for rocHPL-MxP and 31% for HPG-MxP, with similar trends on Portage. These results provide portable guidance for sensor validation and power-aware optimization on current and future exascale systems.

Mcdaniel, Adam [ORNL] (ORCID:000000016926028X)↗

Energy-efficient, Large-scale Molecular Dynamics Simulations via Hardware- and Algorithm-level Optimization

This work aims to develop a framework for energy-efficient computing that will enable molecular dynamics (MD) simulations of large-scale phenomena with atomic precision and simultaneously remove computational bottlenecks limiting the speed of MD simulations. We seek to implement such an approach through the development of surrogate models for the interatomic force calculation combined with the use of mixed numerical precision formats. For a model system of neutral atoms (only pairwise interactions), significant force calculation efficiency improvements were achieved, without detrimental effects on atomic structures or average energies, using single precision, by developing a surrogate model (deep neural network), and by quantizing this surrogate model. For a model system of charged atoms, the reciprocal-space calculation of electrostatic interactions was identified as the main bottleneck, and the development of a surrogate model should be pursued to achieve an estimated one-order-of-magnitude additional speedup.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Compressed basis GMRES on high-performance graphics processing units

Krylov methods provide a fast and highly parallel numerical tool for the iterative solution of many large-scale sparse linear systems. To a large extent, the performance of practical realizations of these methods is constrained by the communication bandwidth in current computer architectures, motivating the investigation of sophisticated techniques to avoid, reduce, and/or hide the message-passing costs (in distributed platforms) and the memory accesses (in all architectures). This article leverages Ginkgo’s memory accessor in order to integrate a communication-reduction strategy into the (Krylov) GMRES solver that decouples the storage format (i.e., the data representation in memory) of the orthogonal basis from the arithmetic precision that is employed during the operations with that basis. Given that the execution time of the GMRES solver is largely determined by the memory accesses, the cost of the datatype transforms can be mostly hidden, resulting in the acceleration of the iterative step via a decrease in the volume of bits being retrieved from memory. Together with the special properties of the orthonormal basis (whose elements are all bounded by 1), this paves the road toward the aggressive customization of the storage format, which includes some floating-point as well as fixed-point formats with mild impact on the convergence of the iterative process. We develop a high-performance implementation of the “compressed basis GMRES” solver in the Ginkgo sparse linear algebra library using a large set of test problems from the SuiteSparse Matrix Collection. We demonstrate robustness and performance advantages on a modern NVIDIA V100 graphics processing unit (GPU) of up to 50% over the standard GMRES solver that stores all data in IEEE double-precision.

97 MATHEMATICS AND COMPUTING↗

Analysis of Vector Particle-In-Cell (VPIC) memory usage optimizations on cutting-edge computer architectures

Vector Particle-In-Cell (VPIC) is one of the fastest plasma simulation codes in the world, with particle numbers ranging from one trillion on the first petascale system, Roadrunner, to ten trillion particles on the more recent Blue Waters supercomputer. As supercomputers continue to grow rapidly in size, so too does the gap between computing capability and memory capability. Current memory systems limit VPIC simulations greatly as the maximum number of particles that can be simulated directly depends on the available memory. In this study, we present a suite of VPIC memory optimizations (i.e., particle weight, half-precision, and fixed-point optimizations) that enable a significant increase in the number of particles in VPIC simulations. Here, we assess the optimizations’ impact on memory and runtime performance for a suite of cutting-edge computer architectures such has the NVIDIA V100 GPU, the IBM Power9, and the Fujitsu A64FX architectures. Our optimizations enable a 31.25% reduction in memory usage and up to 40% increase in the number of particles. This paper extends our work on developing particle storage format optimizations Tan et al.

97 MATHEMATICS AND COMPUTING↗

DFT-FE 1.0: A massively parallel hybrid CPU-GPU density functional theory code using finite-element discretization

In this work, we present DFT-FE 1.0, building on DFT-FE 0.6 [Comput. Phys. Commun. 246, 106853 (2020)], to conduct fast and accurate large-scale density functional theory (DFT) calculations (reaching ~ 100,000 electrons) on both many-core CPU and hybrid CPU-GPU computing architectures. This work involves improvements in the real-space formulation—via an improved treatment of the electrostatic interactions that substantially enhances the computational efficiency—as well high-performance computing aspects, including the GPU acceleration of all the key compute kernels in DFT-FE. We demonstrate the accuracy by comparing the ground-state energies, ionic forces and cell stresses on a wide-range of benchmark systems against those obtained from widely used DFT codes. Further, we demonstrate the numerical efficiency of our implementation, which yields ~ 20× CPU-GPU speed-up by using GPU acceleration on hybrid CPU-GPU nodes. Notably, owing to the parallel-scaling of the GPU implementation, we obtain wall-times of 80–140 seconds for full ground-state calculations, with stringent accuracy, on benchmark systems containing ~ 6, 000 – 15,000 electrons.

pseudopotential↗

The St. Benedict Facility: Probing Fundamental Symmetries through Mixed Mirror $β$-Decays

Precise measurements of nuclear beta decays provide a unique insight into the Standard Model due to their connection to the electroweak interaction. These decays help constrain the unitarity or non-unitarity of the Cabibbo–Kobayashi–Maskawa (CKM) quark mixing matrix, and can uniquely probe the existence of exotic scalar or tensor currents. Of these decays, superallowed mixed mirror transitions have been the least well-studied, in part due to the absence of data on their Fermi to Gamow-Teller mixing ratios ($ρ$). At the Nuclear Science Laboratory (NSL) at the University of Notre Dame, the Superallowed Transition Beta-Neutrino Decay Ion Coincidence Trap (St. Benedict) is being constructed to determine the ρ for various mirror decays via a measurement of the beta–neutrino angular correlation parameter ($α_{βν}$) to a relative precision of 0.5%. In this work, we present an overview of the St. Benedict facility and the impact it will have on various Beyond the Standard Model studies, including an expanded sensitivity study of $ρ$ for various mirror nuclei accessible to the facility. A feasibility evaluation is also presented that indicates the measurement goals for many mirror nuclei, which are currently attainable in a week of radioactive beam delivery at the NSL.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Efficient Quantum Gibbs Samplers with Kubo–Martin–Schwinger Detailed Balance Condition

Lindblad dynamics and other open-system dynamics provide a promising path towards efficient Gibbs sampling on quantum computers. In these proposals, the Lindbladian is obtained via an algorithmic construction akin to designing an artificial thermostat in classical Monte Carlo or molecular dynamics methods, rather than being treated as an approximation to weakly coupled system-bath unitary dynamics. Recently, Chen, Kastoryano, and Gilyén (arXiv:2311.09207) introduced the first efficiently implementable Lindbladian satisfying the Kubo–Martin–Schwinger (KMS) detailed balance condition, which ensures that the Gibbs state is a fixed point of the dynamics and is applicable to non-commuting Hamiltonians. This Gibbs sampler uses a continuously parameterized set of jump operators, and the energy resolution required for implementing each jump operator depends only logarithmically on the precision and the mixing time. In this work, we build upon the structural characterization of KMS detailed balanced Lindbladians by Fagnola and Umanità, and develop a family of efficient quantum Gibbs samplers using a finite set of jump operators (the number can be as few as one), akin to the classical Markov chain-based sampling algorithm. Compared to the existing works, our quantum Gibbs samplers have a comparable quantum simulation cost but with greater design flexibility and a much simpler implementation and error analysis. Moreover, it encompasses the construction of Chen, Kastoryano, and Gilyén as a special instance.

97 MATHEMATICS AND COMPUTING↗

Combined First-Principles and Experimental Investigation into the Reactivity of Codeposited Chromium–Carbon under Pressure

High-pressure synthesis in the diamond anvil cell suffers from the lack of a general approach for the control of precursor stoichiometry and homogeneity. Here, we present results from a new method we have developed that uses magnetron cosputtering to prepare stoichiometrically precise and atomically mixed amorphous films of Cr:C. Laser-heated diamond anvil cell experiments carried out on a flake of this sample at pressures between 13.5 and 24.3 GPa lead to the observation of Cr 3 C (Pnma) over the entire pressure range–in good agreement with our in-house theoretical predictions–but also reveal two other metastable phases that were not expected: a novel monoclinic chromium carbide phase and the NaCl-type CrC (Fm3̅m) phase. The unexpected stability of CrC is investigated by using first-principles methods, revealing a large stabilizing effect tied to substoichiometry at the carbon site. These results offer an important case study into the current limitations of crystal structure prediction methods with regard to phase complexity and bolster the growing need for advanced theoretical approaches that can more completely survey experimentally unexplored phase space.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Engineering advancements in microfluidic systems for enhanced mixing at low Reynolds numbers

Mixing within micro- and millichannels is a pivotal element across various applications, ranging from chemical synthesis to biomedical diagnostics and environmental monitoring. The inherent low Reynolds number flow in these channels often results in a parabolic velocity profile, leading to a broad residence time distribution. Achieving efficient mixing at such small scales presents unique challenges and opportunities. This review encompasses various techniques and strategies to evaluate and enhance mixing efficiency in these confined environments. It explores the significance of mixing in micro- and millichannels, highlighting its relevance for enhanced reaction kinetics, homogeneity in mixed fluids, and analytical accuracy. We discuss various mixing methodologies that have been employed to get a narrower residence time distribution. The role of channel geometry, flow conditions, and mixing mechanisms in influencing the mixing performance are also discussed. Various emerging technologies and advancements in microfluidic devices and tools specifically designed to enhance mixing efficiency are highlighted. We emphasize the potential applications of micro- and millichannels in fields of nanoparticle synthesis, which can be utilized for biological applications. Additionally, the prospects of machine learning and artificial intelligence are offered toward incorporating better mixing to achieve precise control over nanoparticle synthesis, ultimately enhancing the potential for applications in these miniature fluidic systems.

Biochemistry & Molecular Biology↗

Participation in Intensity Frontier Neutrino Physics (closeout report)

The flagship currently-running experiment at Fermilab is the NuMI Off-axis νe Appearance (NOvA) experiment. Together with the (complementary) T2K experiment in Japan, it provides the opportunity for improved measurement precision on neutrino mixing parameters and addresses topics of strong interest: whether neutrino masses follow a normal (NH) or inverted (IH) hierarchy; whether the ν 3 state contains a symmetric mixture of ν µ and ν τ (“maximal mixing,” pointing to a possible new symmetry of nature), or, if not, what is the octant of θ 23 ; and whether neutrino mixing violates CP symmetry. Recent results from both experiments suggest we may be on the verge of important discoveries. Towards the end of the decade, these experiments will be superseded by DUNE in the U.S. and HyperK in Japan.

2x2↗

Addressing Low-Cost Methane Sensor Calibration Shortcomings with Machine Learning

Quantifying methane emissions is essential for meeting near-term climate goals and is typically carried out using methane concentrations measured downwind of the source. One major source of methane that is important to observe and promptly remediate is fugitive emissions from oil and gas production sites but installing methane sensors at the thousands of sites within a production basin is expensive. In recent years, relatively inexpensive metal oxide sensors have been used to measure methane concentrations at production sites. Current methods used to calibrate metal oxide sensors have been shown to have significant shortcomings, resulting in limited confidence in methane concentrations generated by these sensors. To address this, we investigate using machine learning (ML) to generate a model that converts metal oxide sensor output to methane mixing ratios. To generate test data, two metal oxide sensors, TGS2600 and TGS2611, were collocated with a trace methane analyzer downwind of controlled methane releases. Over the duration of the measurements, the trace gas analyzer’s average methane mixing ratio was 2.40 ppm with a maximum of 147.6 ppm. The average calculated methane mixing ratios for the TGS2600 and TGS2611 using the ML algorithm were 2.42 ppm and 2.40 ppm, with maximum values of 117.5 ppm and 106.3 ppm, respectively. A comparison of histograms generated using the analyzer and metal oxide sensors mixing ratios shows overlap coefficients of 0.95 and 0.94 for the TGS2600 and TGS2611, respectively. Overall, our results showed there was a good agreement between the ML-derived metal oxide sensors’ mixing ratios and those generated using the more accurate trace gas analyzer. This suggests that the response of lower-cost sensors calibrated using ML could be used to generate mixing ratios with precision and accuracy comparable to higher priced trace methane analyzers. This would improve confidence in low-cost sensors’ response, reduce the cost of sensor deployment, and allow for timely and accurate tracking of methane emissions.

03 NATURAL GAS↗

SoLAr: Solar Neutrinos in Liquid Argon

SoLAr is a new concept for a liquid-argon neutrino detector technology to extend the sensitivities of these devices to the MeV energy range - expanding the physics reach of these next-generation detectors to include solar neutrinos. We propose this novel concept to significantly improve the precision on solar neutrino mixing parameters and to observe the "hep branch" of the proton-proton fusion chain. The SoLAr detector will achieve flavour-tagging of solar neutrinos in liquid argon. The SoLAr technology will be based on the concept of monolithic light-charge pixel-based readout which addresses the main requirements for such a detector: a low energy threshold with excellent energy resolution (approximately 7%) and background rejection through pulse-shape discrimination. The SoLAr concept is also timely as a possible technology choice for the DUNE "Module of Opportunity", which could serve as a next-generation multi-purpose observatory for neutrinos from the MeV to the GeV range. The goal of SoLAr is to observe solar neutrinos in a 10 ton-scale detector and to demonstrate that the required background suppression and energy resolution can be achieved. SoLAr will pave the way for a precise measurement of the 8-B flux, an improved precision on solar neutrino mixing parameters, and ultimately lead to the first observation of hep neutrinos in the DUNE Module of Opportunity.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗