Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “mixed precision”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

End-to-end codesign of Hessian-aware quantized neural networks for FPGAs

Here, we develop an end-to-end workflow for the training and implementation of co-designed neural networks (NNs) for efficient field-programmable gate array (FPGA) hardware. Our approach leverages Hessian-aware quantization of NNs, the Quantized Open Neural Network Exchange intermediate representation, and the hls4ml tool flow for transpiling NNs into FPGA firmware. This makes efficient NN implementations in hardware accessible to nonexperts in a single open sourced workflow that can be deployed for real-time machine-learning applications in a wide range of scientific and industrial settings. We demonstrate the workflow in a particle physics application involving trigger decisions that must operate at the 40-MHz collision rate of the CERN Large Hadron Collider (LHC). Given the high collision rate, all data processing must be implemented on FPGA hardware within the strict area and latency requirements. Based on these constraints, we implement an optimized mixed-precision NN classifier for high-momentum particle jets in simulated LHC proton-proton collisions.

47 OTHER INSTRUMENTATION↗

GenSLMs: Genome-scale language models reveal SARS-CoV-2 evolutionary dynamics

We seek to transform how new and emergent variants of pandemic-causing viruses, specifically SARS-CoV-2, are identified and classified. By adapting large language models (LLMs) for genomic data, we build genome-scale language models (GenSLMs) which can learn the evolutionary landscape of SARS-CoV-2 genomes. By pre-training on over 110 million prokaryotic gene sequences and fine-tuning a SARS-CoV-2-specific model on 1.5 million genomes, we show that GenSLMs can accurately and rapidly identify variants of concern. Thus, to our knowledge, GenSLMs represents one of the first whole-genome scale foundation models which can generalize to other prediction tasks. We demonstrate scaling of GenSLMs on GPU-based supercomputers and AI-hardware accelerators utilizing 1.63 Zettaflops in training runs with a sustained performance of 121 PFLOPS in mixed precision and peak of 850 PFLOPS. We present initial scientific insights from examining GenSLMs in tracking evolutionary dynamics of SARS-CoV-2, paving the path to realizing this on large biological data.

Zvyagin, Maxim↗

AmeriFlux FLUXNET-1F US-ARM ARM Southern Great Plains site- Lamont

This is the AmeriFlux Management Project (AMP) created FLUXNET-1F version of the carbon flux data for the site US-ARM ARM Southern Great Plains site- Lamont. This is the FLUXNET version of the carbon flux data for the site US-ARM ARM Southern Great Plains site- Lamont produced by applying the standard ONEFlux (1F) software. Site Description - Central facility tower crop field (winter wheat, corn, soy, alfalfa). The site also has continuous measurements of precise mixing ratios of CO2, CH4, CO, N2O, and isotopic ratio of 13CO2. In addition to 4 m system, there are sonic anemometers at 25 m and 60 m.

Biraud, Sebastien↗

User Manual - HydraGNN v5.0: Distributed Implementation of Multi-Tasking Graph Neural Networks

This document serves as the user manual for HydraGNN v5.0, a scalable graph neural network (GNN) architecture for simultaneous prediction of multiple target properties using multi-task learning (MTL). This version of HydraGNN has been developed primarily to support the development, training, and deployment of predictive graph-based deep learning (DL) models for atomistic materials modeling. HydraGNN is templated over 13 message-passing policies, including invariant models (GIN, PNA, PNAPlus, GAT, MFC, CGCNN, SAGE, SchNet, DimeNet) and equivariant models (EGNN, PNAEq, PAINN, MACE), and supports distributed training via distributed data parallelism (DDP), DeepSpeed, and Fully Sharded Data Parallelism (FSDP) on leadership-class supercomputers. Although HydraGNN can be applied to problems beyond atomistic materials modeling, its current use is confined to homogeneous graphs. Additional capabilities include machine-learned interatomic potentials with energy-conserving forces, General, Powerful, and Scalable Graph Transformer (GraphGPS) global attention, periodic boundary conditions, hyperparameter optimization, mixed-precision training, and uncertainty quantification.

97 MATHEMATICS AND COMPUTING↗

Fine-Grained Power and Energy Attribution on AMD GPU/APU-Based Exascale Nodes

Modern exascale GPU- and APU-based systems provide multiple power and energy sensors, but differences in scope, update rate, timing, and filtering complicate the attribution of short-lived accelerator activity. This paper presents a methodology to characterize and correct these effects on Cray EX systems with AMD Instinct MI250X GPUs (Frontier) and MI300A APUs (Portage). Using controlled square-wave workloads, we quantify update intervals, delay, aliasing, and variability across up to 512 GPUs and 480 APUs with on-chip (rocm-smi/amd-smi) and off-chip Cray Power Management sensors. We reconstruct power from cumulative energy counters to achieve faster response times, validate it against on-chip, off-chip, and node-level sensors, and integrate the resulting streams into a Score-P/PAPI-based tool for time-aligned, phase-level attribution. Applied to rocHPL, rocHPL-MxP, and HPG-MxP, the method separates energy savings due to reduced runtime from changes in power. Mixed precision reduces node energy on Frontier by 79% for rocHPL-MxP and 31% for HPG-MxP, with similar trends on Portage. These results provide portable guidance for sensor validation and power-aware optimization on current and future exascale systems.

Mcdaniel, Adam [ORNL] (ORCID:000000016926028X)↗

The First Swift Intensive AGN Accretion Disk Reverberation Mapping Survey

Swift intensive accretion disk reverberation mapping of four AGN yielded light curves sampled ∼200–350 times in 0.3–10 keV X-ray and six UV/optical bands. Uniform reduction and cross-correlation analysis of these data sets yields three main results: (1) The X-ray/UV correlations are much weaker than those within the UV/optical, posing severe problems for the lamp-post reprocessing model in which variations in a central X-ray corona drive and power those in the surrounding accretion disk. (2) The UV/optical interband lags are generally consistent with t μ l4 3 as predicted by the centrally illuminated thin accretion disk model. While the average interband lags are somewhat larger than predicted, these results alone are not inconsistent with the thin disk model given the large systematic uncertainties involved. (3) The one exception is the U band lags, which are on average a factor of ∼2.2 larger than predicted from the surrounding band data and fits. This excess appears to be due to diffuse continuum emission from the broad-line region (BLR). The precise mixing of disk and BLR components cannot be determined from these data alone. The lags in different AGN appear to scale with mass or luminosity. We also find that there are systematic differences between the uncertainties derived by JAVELIN versus more standard lag measurement techniques, with JAVELIN reporting smaller uncertainties by a factor of 2.5 on average. In order to be conservative only standard techniques were used in the analyses reported herein.

R. Edelson↗

Performance Optimization Methods for a Memory-Bound, Unstructured-Grid CFD Application on Massively Parallel GPU Platforms

Computational performance of the FUN3D unstructured-grid computational fluid dynamics (CFD) application on massively parallel GPU environments is memory-bound and highly dependent upon efficient reads from and atomic updates to the irregular cell-, edge-, and node-based data structures. In this talk, we present recent efforts into optimizing select performance-critical kernels on NVIDIA Tesla V100 and A100 GPUs and AMD CDNA MI100 GPUs. A novel use of L2 cache residency controls and asynchronous loads into on-chip shared memory are explored on the A100 GPU for the sparse iterative solver, which is dominated by mixed-precision, sparse matrix vector multiplication. Demonstrations show that these methods improve global memory bandwidth utilization by 13.5% on the A100 GPU. Several techniques are also presented that use registers and/or shared memory to facilitate array transposition and aggregation which combine to reduce the frequency and increase the cache efficiency of floating-point atomic updates to the irregular data structures. These methods are demonstrated to improve the kernel throughput by nearly 500% on select kernels on the AMD MI100 over atomic updates directly to global memory. Overall, both V100 and A100 GPUs outperformed the MI100 GPU on kernels dominated by double-precision atomic updates; however, the techniques demonstrated here reduced the performance gap and improved the MI100 performance.

GPU CPU unstructured CFD memory↗

Results of Propellant Mixing Variable Study Using Precise Pressure-Based Burn Rate Calculations

A designed experiment was conducted in which three mix processing variables (pre-curative addition mix temperature, pre-curative addition mixing time, and mixer speed) were varied to estimate their effects on within-mix propellant burn rate variability. The chosen discriminator for the experiment was the 2-inch diameter by 4-inch long (2x4) Center-Perforated (CP) ballistic evaluation motor. Motor nozzle throat diameters were sized to produce a common targeted chamber pressure. Initial data analysis did not show a statistically significant effect. Because propellant burn rate must be directly related to chamber pressure, a method was developed that showed statistically significant effects on chamber pressure (either maximum or average) by adjustments to the process settings. Burn rates were calculated from chamber pressures and these were then normalized to a common pressure for comparative purposes. The pressure-based method of burn rate determination showed significant reduction in error when compared to results obtained from the Brooks' modification of the propellant web-bisector burn rate determination method. Analysis of effects using burn rates calculated by the pressure-based method showed a significant correlation of within-mix burn rate dispersion to mixing duration and the quadratic of mixing duration. The findings were confirmed in a series of mixes that examined the effects of mixing time on burn rate variation, which yielded the same results.

Stefanski, Philip L.↗

Energy-efficient, Large-scale Molecular Dynamics Simulations via Hardware- and Algorithm-level Optimization

This work aims to develop a framework for energy-efficient computing that will enable molecular dynamics (MD) simulations of large-scale phenomena with atomic precision and simultaneously remove computational bottlenecks limiting the speed of MD simulations. We seek to implement such an approach through the development of surrogate models for the interatomic force calculation combined with the use of mixed numerical precision formats. For a model system of neutral atoms (only pairwise interactions), significant force calculation efficiency improvements were achieved, without detrimental effects on atomic structures or average energies, using single precision, by developing a surrogate model (deep neural network), and by quantizing this surrogate model. For a model system of charged atoms, the reciprocal-space calculation of electrostatic interactions was identified as the main bottleneck, and the development of a surrogate model should be pursued to achieve an estimated one-order-of-magnitude additional speedup.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Compressed basis GMRES on high-performance graphics processing units

Krylov methods provide a fast and highly parallel numerical tool for the iterative solution of many large-scale sparse linear systems. To a large extent, the performance of practical realizations of these methods is constrained by the communication bandwidth in current computer architectures, motivating the investigation of sophisticated techniques to avoid, reduce, and/or hide the message-passing costs (in distributed platforms) and the memory accesses (in all architectures). This article leverages Ginkgo’s memory accessor in order to integrate a communication-reduction strategy into the (Krylov) GMRES solver that decouples the storage format (i.e., the data representation in memory) of the orthogonal basis from the arithmetic precision that is employed during the operations with that basis. Given that the execution time of the GMRES solver is largely determined by the memory accesses, the cost of the datatype transforms can be mostly hidden, resulting in the acceleration of the iterative step via a decrease in the volume of bits being retrieved from memory. Together with the special properties of the orthonormal basis (whose elements are all bounded by 1), this paves the road toward the aggressive customization of the storage format, which includes some floating-point as well as fixed-point formats with mild impact on the convergence of the iterative process. We develop a high-performance implementation of the “compressed basis GMRES” solver in the Ginkgo sparse linear algebra library using a large set of test problems from the SuiteSparse Matrix Collection. We demonstrate robustness and performance advantages on a modern NVIDIA V100 graphics processing unit (GPU) of up to 50% over the standard GMRES solver that stores all data in IEEE double-precision.

97 MATHEMATICS AND COMPUTING↗

Analysis of Vector Particle-In-Cell (VPIC) memory usage optimizations on cutting-edge computer architectures

Vector Particle-In-Cell (VPIC) is one of the fastest plasma simulation codes in the world, with particle numbers ranging from one trillion on the first petascale system, Roadrunner, to ten trillion particles on the more recent Blue Waters supercomputer. As supercomputers continue to grow rapidly in size, so too does the gap between computing capability and memory capability. Current memory systems limit VPIC simulations greatly as the maximum number of particles that can be simulated directly depends on the available memory. In this study, we present a suite of VPIC memory optimizations (i.e., particle weight, half-precision, and fixed-point optimizations) that enable a significant increase in the number of particles in VPIC simulations. Here, we assess the optimizations’ impact on memory and runtime performance for a suite of cutting-edge computer architectures such has the NVIDIA V100 GPU, the IBM Power9, and the Fujitsu A64FX architectures. Our optimizations enable a 31.25% reduction in memory usage and up to 40% increase in the number of particles. This paper extends our work on developing particle storage format optimizations Tan et al.

97 MATHEMATICS AND COMPUTING↗

DFT-FE 1.0: A massively parallel hybrid CPU-GPU density functional theory code using finite-element discretization

In this work, we present DFT-FE 1.0, building on DFT-FE 0.6 [Comput. Phys. Commun. 246, 106853 (2020)], to conduct fast and accurate large-scale density functional theory (DFT) calculations (reaching ~ 100,000 electrons) on both many-core CPU and hybrid CPU-GPU computing architectures. This work involves improvements in the real-space formulation—via an improved treatment of the electrostatic interactions that substantially enhances the computational efficiency—as well high-performance computing aspects, including the GPU acceleration of all the key compute kernels in DFT-FE. We demonstrate the accuracy by comparing the ground-state energies, ionic forces and cell stresses on a wide-range of benchmark systems against those obtained from widely used DFT codes. Further, we demonstrate the numerical efficiency of our implementation, which yields ~ 20× CPU-GPU speed-up by using GPU acceleration on hybrid CPU-GPU nodes. Notably, owing to the parallel-scaling of the GPU implementation, we obtain wall-times of 80–140 seconds for full ground-state calculations, with stringent accuracy, on benchmark systems containing ~ 6, 000 – 15,000 electrons.

pseudopotential↗

Studies of long-life pulsed CO2 laser with Pt/SnO2 catalyst

Closed-cycle CO2 laser testing with and without a catalyst and with and without CO addition indicate that a catalyst is necessary for long-term operation. Initial results indicate that CO addition with a catalyst may prove optimal, but a precise gas mix has not yet been determined. A long-term run of 10 to the 6th power pulses using 1.3% added CO and a 2% Pt on SnO2 catalyst yields an efficiency of about 95% of open-cycle steady-state power. A simple mathematical analysis yields results which may be sufficient for determining optimum running conditions. Future plans call for testing various catalysts in the laser and longer tests, 10 to the 7th power pulses. A Gas Chromatograph will be installed to measure gas species concentration and the analysis will be slightly modified to include neglected but possibly important parameters.

Sidney, Barry D.↗

The St. Benedict Facility: Probing Fundamental Symmetries through Mixed Mirror $β$-Decays

Precise measurements of nuclear beta decays provide a unique insight into the Standard Model due to their connection to the electroweak interaction. These decays help constrain the unitarity or non-unitarity of the Cabibbo–Kobayashi–Maskawa (CKM) quark mixing matrix, and can uniquely probe the existence of exotic scalar or tensor currents. Of these decays, superallowed mixed mirror transitions have been the least well-studied, in part due to the absence of data on their Fermi to Gamow-Teller mixing ratios ($ρ$). At the Nuclear Science Laboratory (NSL) at the University of Notre Dame, the Superallowed Transition Beta-Neutrino Decay Ion Coincidence Trap (St. Benedict) is being constructed to determine the ρ for various mirror decays via a measurement of the beta–neutrino angular correlation parameter ($α_{βν}$) to a relative precision of 0.5%. In this work, we present an overview of the St. Benedict facility and the impact it will have on various Beyond the Standard Model studies, including an expanded sensitivity study of $ρ$ for various mirror nuclei accessible to the facility. A feasibility evaluation is also presented that indicates the measurement goals for many mirror nuclei, which are currently attainable in a week of radioactive beam delivery at the NSL.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

GEER Status Update

History: The Glenn Extreme Environments Rig (GEER) first became operational in the early part of 2015. Since that time GEER has completed a number of scientific tests and has undergone improvements in the chemical delivery system and analytics following a year of operations experience. Recent Updates: In June 2016, the GEER process system was rebuilt to provide a more robust system, higher accuracy and new capabilities. New insulation was installed on the exterior of the pressure vessel and gas lines. The newly revamped GEER plumbing system can provide extremely precise custom gas mixtures using any gas desired by the investigator in any combination. GEER can heat the resulting mixture up to 500 deg C and 1500 psia. The 304 stainless steel vessel walls were polished to reduce corrosion rate and reduce unwanted chemical reactions. The process lines were replaced with high purity Sulfinert coated tubing. The GEER team added the ability to individually boost specialty gases to GEER thus allowing operators to make very precise changes to the gas chemistry inside of GEER during a test while at high temperature and pressure. High accuracy mass flow meters were added to further improve gas mixing accuracy and precision. An in-line, integrated Inficon MicroGC Fusion was added for real time gas analysis along with a high purity gas sampling system, providing fully automated, real time analysis of the gas chemistry inside of GEER in minutes. This complements a co-located mass spectrometer and both are used for regular monitoring of the vessel chemistry. All internal vessel components were replaced with polished 304SS equivalents. Hot vent down capability was increased. Finally, an automated liquid injection system was added and is rated for max vessel operating conditions of (1500 psia, 500 C). Recent Results and Publications: In May 2016, GEER completed a test that exposed high temperature electronics to Venus surface conditions for 21.5 days. This demonstrated the potential for operating robotic spacecraft in the Venus environment without the need for thermal or environmental protection. Results from this test were published in December 2016 and received national media attention. In April 2017 GEER implemented an 80 day test at Venus surface conditions to simulate chemical weathering of expected Venus minerals. This test supported a ROSES award to a team led by Prof. Ralph Harvey of Case Western Reserve University. The test concluded in July 2017 and nearly doubled previous operation record of 42 days at Venus surface conditions. Preliminary results of these and previous experiments were presented at the recent Venus Modeling Workshop. In June 2017, NASA TM2017-219437 "Chemical and Microstructural Changes in Metallic and Ceramic Materials Exposed to Venusian Surface Conditions" was published. This report provides an extensive and valuable resource detailing the behavior of a variety of engineering materials at Venus surface conditions. Community Involvement: An external science advisory panel has been formed.

Kremic, Tibor↗

Quick mixing of epoxy components

Two materials are mixed quickly, thoroughly, and in precise proportion by disposable cartridge. Cartridge mixes components of fast-curing epoxy resins, with no mess, just before they are used. It could also be used in industry and home for caulking, sealing, and patching. Materials to be mixed are initially isolated by cylinder wall within cartridge. Cylinder has vanes, with holes in them, at one end and handle at opposite end. When handle is pulled, grooves on shaft rotate cylinder so that vanes rotate to extrude material A uniformly into material B.

Dunlap, D. E., Jr.↗

Atmospheric CO2 Column Measurements with an Airborne Intensity-Modulated Continuous-Wave 1.57-micron Fiber Laser Lidar

The 2007 National Research Council (NRC) Decadal Survey on Earth Science and Applications from Space recommended Active Sensing of CO2 Emissions over Nights, Days, and Seasons (ASCENDS) as a mid-term, Tier II, NASA space mission. ITT Exelis, formerly ITT Corp., and NASA Langley Research Center have been working together since 2004 to develop and demonstrate a prototype Laser Absorption Spectrometer for making high-precision, column CO2 mixing ratio measurements needed for the ASCENDS mission. This instrument, called the Multifunctional Fiber Laser Lidar (MFLL), operates in an intensity-modulated, continuous-wave mode in the 1.57- micron CO2 absorption band. Flight experiments have been conducted with the MFLL on a Lear-25, UC-12, and DC-8 aircraft over a variety of different surfaces and under a wide range of atmospheric conditions. Very high-precision CO2 column measurements resulting from high signal-to-noise (great than 1300) column optical depth measurements for a 10-s (approximately 1 km) averaging interval have been achieved. In situ measurements of atmospheric CO2 profiles were used to derive the expected CO2 column values, and when compared to the MFLL measurements over desert and vegetated surfaces, the MFLL measurements were found to agree with the in situ-derived CO2 columns to within an average of 0.17% or approximately 0.65 ppmv with a standard deviation of 0.44% or approximately 1.7 ppmv. Initial results demonstrating ranging capability using a swept modulation technique are also presented.

Dobler, Jeremy T.↗