Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “hardware efficiency”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

On the Marriage of Asynchronous Many Task Runtimes and Big Data: A Glance

The rise of the accelerator-based architectures and reconfigurable computing have showcased the weakness of software stack toolchains that still maintain a static view of the hardware instead of relying on a symbiotic relationship between static (e.g., compilers) and dynamic tools (e.g., runtimes). In the past decades, this need has given rise to adaptive runtimes with increasingly finer computational tasks. These finer tasks help to take advantage of the hardware by switching out when a long latency operation is encountered (because of the deeper memory hierarchies and new memory technologies that might target streaming instead of random access), thus trading off idle time for unrelated work. Examples of these finer task runtimes are Asynchronous Many Task (AMT) runtimes, in which highly efficient computational graphs run on a variety of hardware. Due to its inherent latency tolerant characteristics, Latency-sensitive applications, such as Graph Analytics and Big Data can effectively use these runtimes. This paper aims to present an example of how the careful design of an AMT can exploit the hardware substrate when faced with high latency applications such as the ones given in the Big Data domain. Moreover, with its introspection and adaptive capabilities, we aim to show the power of these runtimes when facing the changing requirements of the application workloads. We use the Performance Open Community Runtime (P-OCR) as our vehicle to demonstrate the concepts presented here.

adaptive runtime, big data analysis↗

Collective neutrino oscillations on a quantum computer with hybrid quantum-classical algorithm

We simulate the time evolution of collective neutrino oscillations in two-flavor settings on a quantum computer. We explore the generalization of Trotter-Suzuki approximation to time-dependent Hamiltonian dynamics. The trotterization steps are further optimized using the Cartan decomposition of two-qubit unitary gates U ϵ SU(4) in the minimum number of controlled-NOT (CNOT) gates making the algorithm more resilient to the hardware noise. As a result, a more efficient hybrid quantum-classical algorithm is also explored to solve the problem on noisy intermediate-scale quantum devices.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

A Conceptual Framework for HPC Operational Data Analytics

This paper provides a broad framework for under- standing trends in Operational Data Analytics (ODA) for High- Performance Computing (HPC) facilities. The goal of ODA is to allow for the continuous monitoring, archiving, and analysis of near real-time performance data, providing immediately actionable information for multiple operational uses. In this work, we combine two models to provide a comprehensive HPC ODA framework: one is an evolutionary model of analytics capabilities that consists of four types, which are descriptive, diagnostic, predictive and prescriptive, while the other is a four- pillar model for energy-efficient HPC operations that covers facility, system hardware, system software, and applications. This new framework is then overlaid with a description of current development and production deployments of ODA within leading- edge HPC facilities. Finally, we perform a comprehensive survey of ODA works and classify them according to our framework, in order to demonstrate its effectiveness.

Netti, Alessio↗

A Unifying Framework to Enable Artificial Intelligence in High-Performance Computing Workflows

Current trends point to a future where large-scale scientific applications are tightly coupled high-performance computing/artificial intelligence (HPC/AI) hybrids. Hence, we urgently need to invest in creating a seamless, scalable framework where HPC and AI/machine learning can efficiently work together and adapt to novel hardware and vendor libraries without starting from scratch every few years. Finally, the current ecosystem and sparsely connected community are not sufficient to tackle these challenges, and we require a breakthrough catalyst for science similar to what PyTorch enabled for AI.

high-performance computing↗

How efficiently can AI recognize Wireless Devices?

This poster presents a hardware benchmarking methodology for a 3-layer CNN waveform classifier deployed using ONNX Runtime on an NVIDIA Jetson AGX Orin. The dataset consist of 9 signal types, -30 to +30 dB SNR with 5dB increments. Benchmarking on the Jetson AGX Orin gave an accuracy of 91.9% and GPU throughput of 107,120 predictions/sec (23× faster than CPU). The Jetson GPU reached approximately 27M samples/sec with stable performance but fell below the 40 MHz rate needed for real-time radio feeds. Sustained testing of 5 minutes confirmed stable performance with no memory leaks, establishing a reproducible benchmarking baseline for future edge-deployment optimization.

99 - GENERAL AND MISCELLANEOUS↗

Efficient, Compact, and Smooth Variable Propulsion Motor (Final Report)

In this project, a new architecture of highly efficient hydraulic motor was developed for the propulsion of off-highway vehicles. The motor uses an adjustable linkage driving a cam to vary the displacement of the piston, resulting in a Variable Displacement Linkage Motor (VDLM). The motor uses low friction rolling element bearings to significantly reduce mechanical friction, especially in the demanding low-speed high-torque conditions experienced by off-highway vehicles. The VDLM has high torque capabilities for its size due to the radial piston packaging and use of a multi-lobe cam. A VDLM is very smooth due to the ability to tune the torque ripple through the design of the cam profile. The project was divided into three periods. During the first period, a dynamic model was constructed of the motor to predict the performance of the motor and the vehicle. During the second period, a single-cylinder learning prototype was designed, built, and tested to validate the models constructed in the first period. In the third period, a multi-cylinder prototype motor was optimized, designed, fabricated, and tested. The motor demonstrated excellent mechanical efficiency (above 92.5% across the range of displacements), but the experimentally measure volumetric efficiency was lower than expected due to higher leakage rates created by the poor tolerance control on the prototype. To validate the dynamic models developed in the first period and better understand design trade-offs. In the third period a multi-cylinder concept demonstration prototype will be designed, fabricated, and tested. The final prototype will be tested on a motor dynamometer and will be utilized in hardware-in-the-loop testing to demonstrate its efficiency and performance impacts on the overall drive train. The experimental results were used in a drive train simulation of a compact track loader operating through a drive cycle. Using the VDLM in a hydrostatic circuit yielded 17.1% reduction in fuel consumption and 36.5% reduction in a series hybrid transmission.

99 GENERAL AND MISCELLANEOUS↗

Efficient Mixed-Precision Matrix Factorization of the Inverse Overlap Matrix in Electronic Structure Calculations with AI-Hardware and GPUs

In recent years, a new kind of accelerated hardware has gained popularity in the artificial intelligence (AI) community which enables extremely high-performance tensor contractions in reduced precision for deep neural network calculations. In this article, we exploit Nvidia Tensor cores, a prototypical example of such AI-hardware, to develop a mixed precision approach for computing a dense matrix factorization of the inverse overlap matrix in electronic structure theory, S –1 . This factorization of S –1 , written as ZZT = S –1 , is used to transform the general matrix eigenvalue problem into a standard matrix eigenvalue problem. Here we present a mixed precision iterative refinement algorithm where Z is given recursively using matrix–matrix multiplications and can be computed with high performance on Tensor cores. To understand the performance and accuracy of Tensor cores, comparisons are made to GPU-only implementations in single and double precision. Additionally, we propose a nonparametric stopping criteria which is robust in the face of lower precision floating point operations. The algorithm is particularly useful when we have a good initial guess to Z, for example, from previous time steps in quantum-mechanical molecular dynamics simulations or from a previous iteration in a geometry optimization.

36 MATERIALS SCIENCE↗

Zero and Finite Temperature Quantum Simulations Powered by Quantum Magic

We introduce a quantum information theory-inspired method to improve the characterization of many-body Hamiltonians on near-term quantum devices. We design a new class of similarity transformations that, when applied as a preprocessing step, can substantially simplify a Hamiltonian for subsequent analysis on quantum hardware. By design, these transformations can be identified and applied efficiently using purely classical resources. In practice, these transformations allow us to shorten requisite physical circuit-depths, overcoming constraints imposed by imperfect near-term hardware. Importantly, the quality of our transformations is t u n a b l e : we define a 'ladder' of transformations that yields increasingly simple Hamiltonians at the cost of more classical computation. Using quantum chemistry as a benchmark application, we demonstrate that our protocol leads to significant performance improvements for zero and finite temperature free energy calculations on both digital and analog quantum hardware. Specifically, our energy estimates not only outperform traditional Hartree-Fock solutions, but this performance gap also consistently widens as we tune up the quality of our transformations. In short, our quantum information-based approach opens promising new pathways to realizing useful and feasible quantum chemistry algorithms on near-term hardware.

Physics↗

OpenABLext: An automatic code generation framework for agent-based simulations on CPU-GPU-FPGA heterogeneous platforms

The execution of agent-based simulations (ABSs) on hardware accelerator devices such as graphics processing units (GPUs) has been shown to offer great performance potentials. However, in heterogeneous hardware environments, it can become increasingly difficult to find viable partitions of the simulation and provide implementations for different hardware devices. To automate this process, we present OpenABLext, an extension to OpenABL, a model specification language for ABSs. By providing a device-aware OpenCL backend, OpenABLext enables the co-execution of ABS on heterogeneous hardware platforms consisting of central processing units, GPUs, and field programmable gate arrays (FPGAs).We present a novel online dispatching method that efficiently profiles partitions of the simulation during run-time to optimize the hardware assignment while using the profiling results to advance the simulation itself. In addition, OpenABLext features automated conflict resolution based on user-specified rules, supports graph-based simulation spaces, and utilizes an efficient neighbor search algorithm. We show the improved performance of OpenABLext and demonstrate the potential of FPGAs in the context of ABS. We illustrate how co-execution can be used to further lower execution times. OpenABLext can be seen as an enabler to tap the computing power of heterogeneous hardware platforms for ABS.

97 MATHEMATICS AND COMPUTING↗

Ps and Qs: Quantization-Aware Pruning for Efficient Low Latency Neural Network Inference

Efficient machine learning implementations optimized for inference in hardware have wide-ranging benefits, depending on the application, from lower inference latency to higher data throughput and reduced energy consumption. Two popular techniques for reducing computation in neural networks are pruning, removing insignificant synapses, and quantization, reducing the precision of the calculations. In this work, we explore the interplay between pruning and quantization during the training of neural networks for ultra low latency applications targeting high energy physics use cases. Techniques developed for this study have potential applications across many other domains. We study various configurations of pruning during quantization-aware training, which we term quantization-aware pruning, and the effect of techniques like regularization, batch normalization, and different pruning schemes on performance, computational complexity, and information content metrics. We find that quantization-aware pruning yields more computationally efficient models than either pruning or quantization alone for our task. Further, quantization-aware pruning typically performs similar to or better in terms of computational efficiency compared to other neural architecture search techniques like Bayesian optimization. Surprisingly, while networks with different training configurations can have similar performance for the benchmark application, the information content in the network can vary significantly, affecting its generalizability.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

GaN‐based split phase transformer‐less PV inverter with auxiliary ZVT circuit

This paper explores performance enhancement of the common ground dynamic dc‐link (CGDL) inverter for single phase photovoltaic (PV) applications by a combination of gallium nitride (GaN) devices, split phase topology, coupled inductors, and zero voltage transition (ZVT) scheme. The CGDL inverter has the inherent advantage of minimised dc‐link capacitance and negligible leakage current due to the common ground configuration, but its reported efficiency was usually lower because of the higher dc‐link voltage used for the reduction of decoupling capacitance to a great extent. To solve the efficiency problem, in this study, a soft switching circuit is proposed for the first stage, while a coupled inductor integrated magnetics is incorporated in the second stage to reduce inductor loss, volume, and cost. Both of these topological improvements combined with the use of GaN devices facilitate in achieving high efficiency without compromising converter power density. Extensive experimental results are provided from a GaN based 1 kVA hardware prototype to demonstrate the superior performance of the CGDL inverter attaining a peak efficiency of 98.7% and a California Energy Commission efficiency of 98.5% at 75/50 kHz switching frequency.

Xia, Yinglai↗

Benchmark Tracking System for Performance Monitoring

Benchmarking is essential for high-performance software development, particularly for monitoring performance across code iterations. This project focused on enhancing the benchmarking process for Lamellar, an asynchronous runtime for High-Performance Computing (HPC) systems developed at Pacific Northwest National Laboratory. Prior to this work, benchmark results were difficult to track and compare across code versions, presenting significant challenges in identifying performance regressions and long-term trends. The primary objective was to establish a systematic, reproducible approach for measuring performance and detecting regressions following code commits. Our methodology involved three key components: standardizing benchmark outputs, implementing data versioning, and developing analysis tools. We standardized the benchmark output format to JSON Line records containing specific fields (execution time, hardware specifications, and environmental variables). To address data management challenges, we evaluated several options and eventually chose a git repository dedicated to benchmark data. We developed a suite of Python tools that processed benchmark results, enriched them with metadata, and facilitated search in the repository. The resulting system enables more efficient filtering and comparison of performance metrics across commit histories, hardware configurations, and benchmark variants through a unified query interface. Our implementation reduces computational overhead by first checking for existing results through configuration matching before initiating new benchmark runs, thereby conserving resources. The system has been validated by Lamellar developers. It organizes results by benchmark type and build configurations for efficient retrieval. Future developments include a planned Large Language Model interface for predicting benchmark performance, incorporating the criterion package for statistical analysis, which will enable automated detection of statistically significant performance changes, and integration with continuous integration pipelines. Despite these enhancements being reserved for future work, this project has successfully provided the Lamellar development team with a framework for maintaining consistent performance standards and identifying optimization opportunities across workloads and hardware environments.

97 MATHEMATICS AND COMPUTING↗

Citation network datasets for benchmarking spiking graph neural networks on experimental neuromorphic hardware

Spiking neural networks (SNNs) running on neuromorphic computers offer an energy-efficient alternative for AI tasks. Recently, spiking graph neural networks (S-GNNs) have been shown to produce encouraging results on benchmark citation network datasets such as Cora, CiteSeer, and PubMed for node classification tasks. These S-GNNs were run on SNN simulators only because they contain up to tens of thousands of neurons and up to millions of synapses, translating poorly to neuromorphic hardware. Therefore, in this paper, we create a suite of benchmark datasets from the CiteSeer dataset that can be accommodated on current neuromorphic hardware platforms. Our contribution consists of a collection of three datasets. First, we have an induced subgraph of CiteSeer, which we call MiniSeer, containing 2110 papers, 3604 binary features, and 6 topics. Second, MicroSeer is a very small dataset consisting of 84 papers, 1227 features, and 6 topics. Lastly, BiteSeer is a collection of 15 binary classification datasets. We present creation of these datasets along with accuracies, running times, and spike counts when simulated. We believe that our results in this paper will be used by the neuromorphic community to benchmark, test, and develop neuromorphic hardware and simulators.

Zhu, Kevin [George Mason University, Virginia]↗

Printing polymer blends through in situ active mixing during fused filament fabrication

Fused filament fabrication (FFF) enables production of 3D objects over a range of material compositions at low-cost relative to traditional manufacturing approaches. To date, a limited but growing number of materials are able to be used with FFF, however many applications exist where specific mechanical, thermal, or chemical properties are needed that cannot currently be met with the available feedstock selection. Therefore, a need exists to tune these materials for specific chemical or mechanical properties. One common formulation strategy to address these demanding design parameters is to develop composites or polymer blend filaments. Typically, this is a time-consuming and costly optimization process. Here, we have developed hardware for reproducibly mixing two filaments of similar or dissimilar compositions at the time of printing within individual printed layers. This mixing occurs via software-controlled rotating hardware in the chamber of an extruder’s hot-end. The efficiency of mixing within the printed layers has been characterized in detail as a function of the rotational speed and geometry of the blending hardware. These parameters were exploited to program the ratio and distribution of thermoplastic-based filaments blended within printed extrudate. Example printed specimens were produced with thermoplastic polyurethane (TPU) elastomer blended with rigid polylactic acid (PLA) and Nylon blended with PLA. In addition, a conductive carbon nanotube (CNT)-PLA composite was blended as a function of mixer geometry and input feed ratios with non-conductive PLA and resistance values were measured across the resulting printed specimens.

36 MATERIALS SCIENCE↗

True random number generation using the spin crossover in LaCoO 3

While digital computers rely on software-generated pseudo-random number generators, hardware-based true random number generators (TRNGs), which employ the natural physics of the underlying hardware, provide true stochasticity, and power and area efficiency. Research into TRNGs has extensively relied on the unpredictability in phase transitions, but such phase transitions are difficult to control given their often abrupt and narrow parameter ranges (e.g., occurring in a small temperature window). Here we demonstrate a TRNG based on self-oscillations in LaCoO 3 that is electrically biased within its spin crossover regime. The LaCoO 3 TRNG passes all standard tests of true stochasticity and uses only half the number of components compared to prior TRNGs. Assisted by phase field modeling, we show how spin crossovers are fundamentally better in producing true stochasticity compared to traditional phase transitions. As a validation, by probabilistically solving the NP-hard max-cut problem in a memristor crossbar array using our TRNG as a source of the required stochasticity, we demonstrate solution quality exceeding that using software-generated randomness.

97 MATHEMATICS AND COMPUTING↗

Towards Sustainable Post-Exascale Leadership Computing

As computing systems approach the limits of traditional silicon technology, the diminishing returns in performance per watt present a significant barrier to sustaining growth in HPC. From a large-scale scientific supercomputing facility point of view, we propose a multifaceted strategy toward specialized hardware and architectures that are optimized for energy efficiency in specific applications. We also emphasize the need for integrating energy-aware practices across all levels of HPC, from system design and software development to operational policies. We discuss strategic opportunities such as the adoption of application-specific accelerators, the development of energy-efficient algorithms, and the implementation of data-driven operational analytics. Our goal is to develop a comprehensive roadmap ensuring that future leadership systems at OLCF can meet scientific demands while operating within stringent energy budgets, thereby supporting sustainable computing growth.

Shin, Woong↗

Functional Lithography for Advanced Additive Manufacturing

The viability of 3D printing as a manufacturing technology is limited by the material properties of the printed parts. polySpectra successfully developed a scalable, efficient, and economic print process using commercially available hardware in order to 3D print production parts using novel functional lithography resins. The printed parts exhibited unmatched physical and chemical properties. The combination of the novel resin with unique properties and a facile, economic print process will help elevate additive manufacturing from a prototyping tool to a viable manufacturing technology.

42 ENGINEERING↗

Quantum-classical tradeoffs and multi-controlled quantum gate decompositions in variational algorithms

The computational capabilities of near-term quantum computers are limited by the noisy execution of gate operations and a limited number of physical qubits. Hybrid variational algorithms are well-suited to near-term quantum devices because they allow for a wide range of tradeoffs between the amount of quantum and classical resources used to solve a problem. This paper investigates tradeoffs available at both the algorithmic and hardware levels by studying a specific case – applying the Quantum Approximate Optimization Algorithm (QAOA) to instances of the Maximum Independent Set (MIS) problem. We consider three variants of the QAOA which offer different tradeoffs at the algorithmic level in terms of their required number of classical parameters, quantum gates, and iterations of classical optimization needed. Since MIS is a constrained combinatorial optimization problem, the QAOA must respect the problem constraints. This can be accomplished by using many multi-controlled gate operations which must be decomposed into gates executable by the target hardware. We study the tradeoffs available at this hardware level, combining the gate fidelities and decomposition efficiencies of different native gate sets into a single metric called the gate decomposition cost .

Tomesh, Teague↗