Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “hardware efficiency”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

CEAZ: Accelerating Parallel I/O Via Hardware-Algorithm Co-Designed Adaptive Lossy Compression

As supercomputers continue to grow to exa-scale, the amount of data that needs to be saved or transmitted is exploding. To this end, many previous works have studied using error-bounded lossy compressors to reduce the data size and improve the I/O performance. However, little work has been done for effectively offloading lossy compression onto FPGA-based SmartNICs to reduce the compression overhead. In this paper, we propose a hardware-algorithm co-design of efficient and adaptive lossy compressor for scientific data on FPGAs (called CEAZ) to accelerate parallel I/O. Our contribution is fourfold: (1) We propose an efficient Huffman coding approach that can adaptively update Huffman codewords online based on codewords generated offline (from a variety of representative scientific datasets). (2) We derive a theoretical analysis to support a precise control of compression ratio under an error-bounded compression mode, enabling accurate offline Huffman codewords generation. This also help us create a fixed-ratio compression mode for consistent throughput. (3) We develop an efficient compression pipeline by adopting cuSZ’s dual-quantization algorithm to our hardware use case. (4) We evaluate CEAC on five real-world datasets with both a single FPGA board and 256 nodes from Bridges2 supercomputer. Experiments show that CEAZ outperforms the second-best FPGA-based lossy compressor by 2× of throughput and 9.6× of compression ratio. It also improves MPI_File_write and MPI_Gather throughputs by up to 32.7× and 31.4×, respectively.

Zhang, Chengming↗

A GPU accelerated mixed-precision Smoothed Particle Hydrodynamics framework with cell-based relative coordinates

Smoothed Particle Hydrodynamics (SPH) is essential for modeling complex large-deformation problems across various applications, requiring significant computational power. A major portion of SPH computation time is dedicated to the Nearest Neighboring Particle Search (NNPS) process. While advanced NNPS algorithms have been developed to enhance SPH efficiency, the potential efficiency gains from modern computation hardware remain underexplored. Here, this study investigates the impact of GPU parallel architecture, low-precision computing on GPUs, and GPU memory management on NNPS efficiency. Our approach employs a GPU-accelerated mixed-precision SPH framework, utilizing low precision float-point 16 (FP16) for NNPS while maintaining high precision for other components. To ensure FP16 accuracy in NNPS, we introduce a Relative Coordinated-based Link List (RCLL) algorithm, storing FP16 relative coordinates of particles within background cells. Our testing results show three significant speedup rounds for CPU-based NNPS algorithms. The first comes from parallel GPU computations, with up to a 1000x efficiency gain. The second is achieved through low-precision GPU computing, where the proposed FP16-based RCLL algorithm offers a 1.5x efficiency improvement over the FP64-based approach on GPUs. By optimizing GPU memory bandwidth utilization, the efficiency of the FP16 RCLL algorithm can be further boosted by 2.7x, as demonstrated in an example with 1 million particles. Our code is released at https://github.com/pnnl/lpNNPS4SPH.

97 MATHEMATICS AND COMPUTING↗

A Memory Efficient Lock-Free Circular Queue

Hardware queues are import in many applications, such as data transfer, synchronization of concurrent modules with the need of mutual exclusion constructs. State of the art bounded (of a fixed size) lock free circular queues are implemented either by read/write atomic operations, or barrier conditions, or by separating dequeue and enqueue operations. However, these queues always require an unused element at all the times to safe-guard the front and rear pointers of the queue, so as to avoid data race conditions, which leads to the waste of memory. The waste of memory is especially disadvantageous in applications such as I/O data transfer, and image transfer between processing filters, when large element size is needed, We propose a lock-free solution of the bounded circular queue through read/write atomic operations, but without the need of an extra element in the queue. The proposed solution is implemented and verified in both Verilog and ’C’ languages. We also demonstrate its effectiveness by comparing its area and delay metrics with the implementations of other existing designs of queue.

Miniskar, Narasinga Rao↗

End-to-end codesign of Hessian-aware quantized neural networks for FPGAs

Here, we develop an end-to-end workflow for the training and implementation of co-designed neural networks (NNs) for efficient field-programmable gate array (FPGA) hardware. Our approach leverages Hessian-aware quantization of NNs, the Quantized Open Neural Network Exchange intermediate representation, and the hls4ml tool flow for transpiling NNs into FPGA firmware. This makes efficient NN implementations in hardware accessible to nonexperts in a single open sourced workflow that can be deployed for real-time machine-learning applications in a wide range of scientific and industrial settings. We demonstrate the workflow in a particle physics application involving trigger decisions that must operate at the 40-MHz collision rate of the CERN Large Hadron Collider (LHC). Given the high collision rate, all data processing must be implemented on FPGA hardware within the strict area and latency requirements. Based on these constraints, we implement an optimized mixed-precision NN classifier for high-momentum particle jets in simulated LHC proton-proton collisions.

47 OTHER INSTRUMENTATION↗

QuComm: Optimizing Collective Communication for Distributed Quantum Computing

Distributed quantum computing (DQC) is a scalable way to build a large-scale quantum computing system while the error-prone nonlocal communication between DQC nodes may heavily degrade the fidelity of the distributed quantum program and thus demands specific compiler optimizations. Previous compilers on DQC communication optimization either assumes unlimited communication resource or a few communication qubits due to the hardware limitation. The former compilers may not be efficient when interfacing with communication-resource-constrained DQC hardware while the latter compilers lose the opportunities of optimizing collective communication and routing concurrent communication as they unnecessarily couple limited communication qubits with the implementation of expensive inter-node operations. In this paper, we invent the communication buffer, a communication facility consisting of idle qubits in each compute node, to decouple the execution of inter-node quantum operations from communication qubits: communication qubits are devoted to generating inter-node entanglement while internode operations are conducted in the communication buffer. The communication buffer provides an intermediate layer for inter-node communication and paves the way for collective communication optimization. We then propose QuComm, a buffer-based compiler framework that first performs smart buffer allocation according to communication characteristics of the distributed quantum program and then optimizes and collectively routes inter-node quantum operations. Experimental results on a hierarchical DQC system show that the proposed QuComm can reduce the most expensive inter-node communication request and the latency of various distributed quantum programs by 50.4% and 47.6% on average, respectively.

Wu, Anbang↗

Quantum Computing Strategy 2026

Quantum computing (QC) is a rapidly maturing technology with the potential for revolutionary impacts on stockpile stewardship science and national security. Recent developments in fault-tolerant architectures have compressed vendor roadmaps, and predictions of a production-ready quantum computer by the mid-2030s are becoming increasingly credible. This strategy provides a roadmap for integrating QC into the Advanced Simulation and Computing (ASC) program by investing in four strategic focus areas: 1. Develop Capabilities in Mission-Relevant Quantum Applications: ASC will prioritize developing quantum-ready applications in mission areas that have shown significant promise for quantum advantage, including simulations of materials in extreme environments, nuclear dynamics, solving linear and nonlinear partial differential equations, and uncertainty quantification. These applications directly support stockpile stewardship science and modernization objectives. 2. Conduct R&D in Algorithms, Software, and Hardware: Sustained research into quantum algorithms, robust software tools, and quantum hardware is essential. ASC will develop efficient quantum algorithms; invest in quantum compilers, debuggers, and performance tools; and explore specialized quantum hardware tailored to NNSA’s unique requirements. 3. Engage with Vendors and Partners: Early and active collaboration with commercial quantum hardware vendors and academic partners is critical. Through testbeds, co-design agreements, and quantum demonstration facilities, ASC will influence hardware design, gain early access to emerging technologies, and ensure that quantum platforms evolve to meet mission needs. 4. Build Knowledge, Experience, and Workforce: Expanding and upskilling the quantum-trained workforce is essential to long-term success. This includes hiring, internal training, university outreach, and postdoctoral support to ensure ASC maintains the expertise required to operate, program, and integrate quantum systems as they become available. While quantum computing will never replace classical computing, it has the potential to solve certain problems with speed and accuracy that would be unachievable using any conceivable classical high-performance computing (HPC) system. By investing strategically in QC, ASC will help propel the emergent QC industry, maintain U.S. technological leadership, ensure mission readiness, and position itself to rapidly adopt quantum technologies as they mature.

97 MATHEMATICS AND COMPUTING↗

Matrix-vector multiplication using digital partitioning for more accurate optical computing

Digital partitioning offers a flexible means of increasing the accuracy of an optical matrix-vector processor. This algorithm can be implemented with the same architecture required for a purely analog processor, which gives optical matrix-vector processors the ability to perform high-accuracy calculations at speeds comparable with or greater than electronic computers as well as the ability to perform analog operations at a much greater speed. Digital partitioning is compared with digital multiplication by analog convolution, residue number systems, and redundant number representation in terms of the size and the speed required for an equivalent throughput as well as in terms of the hardware requirements. Digital partitioning and digital multiplication by analog convolution are found to be the most efficient alogrithms if coding time and hardware are considered, and the architecture for digital partitioning permits the use of analog computations to provide the greatest throughput for a single processor.

Gary, C. K.↗

A Generative Control Capability for a Model-based Executive

This paper describes Burton, a core element of a new generation of goal-directed model-based autonomous executives. This executive makes extensive use of component-based declarative models to analyze novel situations and generate novel control actions both at the goal and hardware levels. It uses an extremely efficient online propositional inference engine to efficiently determine likely states consistent with current observations and optimal target states that achieve high level goals. It incorporates a flexible generative control sequencing algorithm within the reactive loop to bridge the gap between current and target states. The system is able to detect and avoid damaging and irreversible situations, After every control action it uses its model and sensors to detect anomalous situations and immediately take corrective action. Efficiency is achieved through a series of model compilation and online policy construction methods, and by exploiting general conventions of hardware design that permit a divide and conquer approach to planning. The paper presents a formal characterization of Burton's capability, develops efficient algorithms, and reports on experience with the implementation in the domain of spacecraft autonomy. Burton is being incorporated as one of the key elements of the Remote Agent core autonomy architecture for Deep Space One, the first spacecraft for NASA's New Millenium program.

Williams, Brian C.↗

ArborX: A Performance Portable Geometric Search Library

Searching for geometric objects that are close in space is a fundamental component of many applications. The performance of search algorithms comes to the forefront as the size of a problem increases both in terms of total object count as well as in the total number of search queries performed. Scientific applications requiring modern leadership-class supercomputers also pose an additional requirement of performance portability, i.e., being able to efficiently utilize a variety of hardware architectures. In this article, we introduce a new open-source C++ search library, ArborX, which we have designed for modern supercomputing architectures. Herein, we examine scalable search algorithms with a focus on performance, including a highly efficient parallel bounding volume hierarchy implementation, and propose a flexible interface making it easy to integrate with existing applications. We demonstrate the performance portability of ArborX on multi-core CPUs and GPUs and compare it to the state-of-the-art libraries such as Boost.Geometry.Index and nanoflann.

97 MATHEMATICS AND COMPUTING↗

The Ejectable Data Recorder: A Lean, Risk-Informed Approach for Hardware Development

NASA is developing the Orion spacecraft to transport crew from the Earth to the Moon as part of the Artemis series of missions. To provide a crew escape capability from pre-launch through ascent, the Orion vehicle is equipped with a Launch Abort System (LAS), built by Lockheed Martin, which pulls the capsule away from the launch vehicle in the event of an abort scenario. The Ascent Abort 2 (AA-2) test flight occurred on July 2, 2019,and tested a production version of the LAS to ensure that it can operate as intended, and to collect a large data set from hundreds of sensors on the vehicle to support Orion flight certification. In the original AA-2 architecture, a single-string set of communications antennas on the LAS would downlink all of the in-flight test data to ground stations. However, that communications architecture was predicted to have data dropouts during abort and jettison of the LAS, and would not support data transmission at all after LAS jettison. As a result, a comprehensive trade study was completed, yielding the addition of antennas on the crew module (CM), a buffer/rebroadcast capability for key portions of the flight, and an ejectable data recorder (EDR) subsystem. This EDR subsystem would serve as a backup to the radio frequency (RF) communications system, and would be non-flight critical, providing a unique capability that enabled management to take a different approach with the hardware and software development. The Crew Module and Separation Ring were developed as “Class 1”Flight Hardware, albeit with some tailoring approaches to enable efficiencies. The Class 1 designation requires full rigor for flight hardware and software, documenting everything that happens to a piece of hardware from procurement through disposal, requiring a full spectrum of acceptance tests, and the highest rigor of quality assurance processes. At the other end of the spectrum, Class 3hardware is controlled, but not intended for flight, and leaves the level of rigor up to the project manager. This classification is often used for research and development projects. Similarly,Class-1E has been recently defined at NASA for ISS payloads and technology development projects that are not flight critical and do not need the full rigor of Class 1 to be successful. The EDR subsystem was challenged at commencement to adopt a skunkworks and agile-like approach to hardware development, allowing for a different risk posture than the rest of the AA-2 hardware. After initially pursuing Class 1 processes, the EDR subsystem design evolved to incorporating numerous commercial components, leading to re-designation as a Class-1E subsystem. The resulting EDR subsystem was fully successful in meeting all flight system requirements, and achieved 100% retrieval of flight test data. This paper will discuss the risk posture of the EDR subsystem and the subsequent tailoring that was enacted as part of its Class-1E status.

EDR↗

Batched Sparse Linear Algebra (Final Report for Subcontract B648960)

This report finalizes design specifications for developing batched kernels for small tensor operations for unassembled matrix-free iterative solvers, batched solvers for partially assembled operators, and batched solvers with support for various sparse formats. The outcome of the project milestones is a set of interfaces to Batched Sparse LA solvers running on hardware accelerators for use in ECP Libraries and Applications. It is part of the development of sparse batched kernels, solvers/preconditioners as well as creating interoperability in xSDK libraries with sparse and dense batched functions to benefit ECP applications. The participants included representatives from ECP libraries (not limited to the xSDK project), applications, and vendors (AMD, Intel, and NVIDIA). Batched sparse linear algebra solvers form the new frontier for algorithmic development and performance engineering. Many applications (ECP and non-ECP alike) require simultaneous solutions of small linear systems of equations that are structurally sparse. To move towards high hardware utilization, it is important to provide these applications with appropriate interfaces to efficient batched sparse solvers running on modern hardware accelerators. We present interface designs in use by HPC software libraries supporting batched sparse linear algebra and the development of sparse batched kernel codes for solvers and preconditioners. We also address the potential interoperability opportunities to keep the software portable between the major hardware accelerators from AMD, Intel, and NVIDIA. The presented interface specifications includes batched band, sparse iterative, and sparse direct solvers. This report summarizes progress in Kokkos Kernels and the xSDK libraries MAGMA, Ginkgo, hypre, SUNDIALS, and SuperLU_dist.

97 MATHEMATICS AND COMPUTING↗

Using Reinforcement Learning to Optimize Quantum Circuits in thePresence of Noise

Many quantum computing frameworks currently use noise aware algorithms for implementing quantum circuits which do not scale efficiently as the size of the hardware architecture increases. As we move towards devices which utilize more qubits, it becomes increasing more important to map quantum circuits in a way that uses resources efficiently as well as maximizes the reliability of the results of that circuit. However, as the hardware increases to the point where Quantum supremacy is attainable, it will infeasible for a brute-force algorithm to find the most optimal circuit layout for circuits of medium to large depth sizes. To this end, we will to rely on reinforcement learning (RL) as a method of building quantum circuits based on observations of the noise characteristics in its environment. In this work, we create a working reinforcement learning environment in which an agent is able to make action which will build the class of circuits which creates the GHZ state. In addition to this, we also get preliminary results of the performance of a Deep Q Neural Network, which initially does not perform as well as we believe it can. In the future, we want to improve the performance of the agent and potentially generalize this environment to more classes of circuits.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Vind: A Blockchain-Enabled Supply Chain Provenance Framework for Energy Delivery Systems

Enterprise-level energy delivery systems (EDSs) depend on different software or hardware vendors to achieve operational efficiency. Critical components of these systems are typically manufactured and integrated by overseas suppliers, which expands the attack surface to adversaries with additional opportunities to infiltrate into EDSs. Due to this reason, the risk management of the EDS supply chain is crucial to ensure that we are knowledgeable about the vulnerabilities in software and hardware components that comprise any critical part, quantifiable risk metrics to assess the severity and exploitability of the attack, and provide remediation solutions that can influence a prioritized mitigation plan. There is a need to realize cyber supply chain risk management for industrial control systems’ hardware, software, and computing and networking services associated with bulk electric system (BES) operations. This article proposes a blockchain-based cyber supply chain provenance platform (“Vind”) for EDSs to realize data provenance in a cyber supply chain ecosystem.

Bandara, Eranga↗

Computer architecture for efficient algorithmic executions in real-time systems: New technology for avionics systems and advanced space vehicles

Improvements and advances in the development of computer architecture now provide innovative technology for the recasting of traditional sequential solutions into high-performance, low-cost, parallel system to increase system performance. Research conducted in development of specialized computer architecture for the algorithmic execution of an avionics system, guidance and control problem in real time is described. A comprehensive treatment of both the hardware and software structures of a customized computer which performs real-time computation of guidance commands with updated estimates of target motion and time-to-go is presented. An optimal, real-time allocation algorithm was developed which maps the algorithmic tasks onto the processing elements. This allocation is based on the critical path analysis. The final stage is the design and development of the hardware structures suitable for the efficient execution of the allocated task graph. The processing element is designed for rapid execution of the allocated tasks. Fault tolerance is a key feature of the overall architecture. Parallel numerical integration techniques, tasks definitions, and allocation algorithms are discussed. The parallel implementation is analytically verified and the experimental results are presented. The design of the data-driven computer architecture, customized for the execution of the particular algorithm, is discussed.

Carroll, Chester C.↗

Solar power from satellites

Microwave beaming of satellite-collected solar energy to earth for conversion to useful industrial power is evaluated for feasibility, with attention given to system efficiencies and costs, ecological impact, hardware to be employed, available options for energy conversion and transmission, and orbiting and assembly. Advantages of such a power generation and conversion system are listed, plausible techniques for conversion of solar energy (thermionic, thermal electric, photovoltaic) and transmission to earth (lasers, arrays of mirrors, microwave beams) are compared. Structural fatigue likely to result from brief daily eclipses, 55% system efficiency at the present state of the art, present projections of system costs, and projected economic implications of the technology are assessed. Two-stage orbiting and assembly plans are described.

Glaser, P. E.↗

Assessing Relay Communications for Mars Sample Return Surface Mission Concepts

The Mars Sample Return (MSR) Campaign is a 3-mission campaign concept supported by NASA and ESA to return samples from the Mars surface. MSR will, for the firsttime ever, present a need to communicate with multiple surfaceassets that are co-located on Mars in a coordinated effort toaccomplish the unified objective of fetching, transporting, andreturning samples from Mars. Currently, Mars surface assetsrelay data to and from Earth using a number of orbiters inwhat’s known as the Mars Relay Network (MRN). This networkis characterized by a small number of surface assets distributedacross the Martian globe and a larger number of orbiters toprovide relay services. As of June 2020, there are two surfaceassets for which five orbiters are providing relay. During theMSR Campaign, there will be two rovers and a lander that allwill require relay communication from a small number of Marsorbiters to meet the aggressive MSR timeline. The inversion ofthe current MRN paradigm, a system of many surface assetsrequiring relay and few orbiters to provide relay, necessitatesthe unique challenge of optimally allocating relay passes tomaximize the operational capability of all assets. The allocationmust consider a large number of trade variables includingMars asset operational requirements and Earth ground systemconstraints, including staffing schedules, operations planningacross time zones, and more. To address these telecommunicationchallenges, the Mars Asset Relay Mission Link AllocationDesign Environment (MARMLADE) tool was developed. Itis a MATLAB-based tool to assign orbiter passes or Direct-From-Earth (DFE) links to each of the three surface assets andquantify the operational efficiency of each surface asset.MARMLADE uses a data set of simulated Mars relay orbitergeometry and telecommunication capabilities provided by JPL’sTelecom Orbit Analysis and Simulation Tool (TOAST) softwareto compute which asset should get each pass based on a seriesof heuristics and predictions of all assets’ states. WithinMARMLADE, the user can provide inputs including the optionfor time-based pass splitting, fixed FWD data rate capabilities,DFE communication capabilities, and link parameters allowingfor the assessment of complex operations and hardware tradesusing surface mission operational efficiency as a primary figureof merit. As the MSR mission concepts continue to mature,MARMLADE is being used to assess ability of all MSR elementsto meet the surface mission timeline requirements and to provide relay link allocations to each of the MSR surface assets.

Lee, Charles↗