Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “hardware algorithms”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

A Roadmap for Reaching the Potential of Brain-Derived Computing

Neuromorphic computing is a critical future technology for the computing industry, but it has yet to achieve its promise and has struggled to establish a cohesive research community. A large part of the challenge is that full realization of the potential of brain inspiration requires advances in both device hardware, computing architectures, and algorithms. This simultaneous development across technology scales is unprecedented in the computing field. This article presents a strategy, framed by market and policy pressures, for moving past these current technological and cultural hurdles to realize its full impact across technology. Achieving the full potential of brain-derived algorithms as well as post-complementary metal-oxide-semiconductor (CMOS) scaling neuromorphic hardware requires appropriately balancing the near-term opportunities of deep learning applications with the long-term potential of less understood opportunities in neural computing.

97 MATHEMATICS AND COMPUTING↗

Using Grover's search algorithm to characterize the Rigetti Quantum Computing Platform

Performance of the Rigetti quantum computing platform was tested during the period of time from 12/16/2019 to 05/18/2020. The Grover's search (GS) algorithm was used as a testing tool, in particular, 3- and 4-level versions of the algorithm. As a result, a number of hardware issues were revealed, so the algorithm was split to smaller blocks and individual gates, and all of them were tested separately. The fidelity decay was found to be due to both decoherent processes and coherent errors of the native hardware gates. These errors were estimated from RX gate benchmarks and were in a form of extra rotation of the quantum state. As a consequence of that, performance of the algorithm was shown to be strongly dependent on the native gate decomposition of the program. Suggestions and possible improvements of the future Rigetti runs are made based on the obtained observations.

97 MATHEMATICS AND COMPUTING↗

ARQUIN: Architectures for Multinode Superconducting Quantum Computers

Many proposals to scale quantum technology rely on modular or distributed designs wherein individual quantum processors, called nodes, are linked together to form one large multinode quantum computer (MNQC). One scalable method to construct an MNQC is using superconducting quantum systems with optical interconnects. However, internode gates in these systems may be two to three orders of magnitude noisier and slower than local operations. Surmounting the limitations of internode gates will require improvements in entanglement generation, use of entanglement distillation, and optimized software and compilers. Still, it remains unclear what performance is possible with current hardware and what performance algorithms require. In this article, we employ a systems analysis approach to quantify overall MNQC performance in terms of hardware models of internode links, entanglement distillation, and local architecture. We show how to navigate tradeoffs in entanglement generation and distillation in the context of algorithm performance, lay out how compilers and software should balance between local and internode gates, and discuss when noisy quantum internode links have an advantage over purely classical links. Here, we find that a factor of 10–100× better link performance is required and introduce a research roadmap for the co-design of hardware and software towards the realization of early MNQCs. While we focus on superconducting devices with optical interconnects, our approach is general across MNQC implementations.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Development of a Hardware-in-The-Loop Testbed for a Decentralized, Data-Driven Electric Vehicle Charging Control Algorithm

This study presents the design of an electric vehicle (EV)-grid integration (EVGI) hardware test-bed to implement smart EV charging algorithms. Here, the proposed test-bed also allows to create different grid events via flexible integration of other power hardware (e.g., controllable loads and battery energy storage systems) and test their impacts on EV charging. The design uses a real-time digital simulator to realize a complex distribution grid model with primary and secondary networks. A grid simulator physically realizes the selected nodes of the simulated grid to power an actual EV, forming a hardware-in-the-loop (HIL) test setup. The EV-grid integration is demonstrated based on the custom hardware and software implementation of the J1772 charging protocol using dSPACE MicroLabBox, operating as a custom EV Supply Equipment (EVSE). The HIL test-bed features a novel testing platform for accurate implementation and analysis of scalable charging algorithms. To this end, a data-driven, decentralized, model-free charging controller based on the Additive Increase and Multiplicative Decrease (AIMD) algorithm is presented and validated on an EV using the HIL test-bed. We tested the proposed algorithm under various case studies, and presented a comparison study with an existing droop-based, decentralized charging solution. The results showed that the EV successfully performed charging commands generated by the EVSE and regulated its charging power to effectively reduce the system loading caused by high EV penetration.

33 ADVANCED PROPULSION SYSTEMS↗

Scaling quantum approximate optimization on near-term hardware

The quantum approximate optimization algorithm (QAOA) is an approach for near-term quantum computers to potentially demonstrate computational advantage in solving combinatorial optimization problems. However, the viability of the QAOA depends on how its performance and resource requirements scale with problem size and complexity for realistic hardware implementations. Here, we quantify scaling of the expected resource requirements by synthesizing optimized circuits for hardware architectures with varying levels of connectivity. Assuming noisy gate operations, we estimate the number of measurements needed to sample the output of the idealized QAOA circuit with high probability. We show the number of measurements, and hence total time to solution, grows exponentially in problem size and problem graph degree as well as depth of the QAOA ansatz, gate infidelities, and inverse hardware graph degree. These problems may be alleviated by increasing hardware connectivity or by recently proposed modifications to the QAOA that achieve higher performance with fewer circuit layers.

97 MATHEMATICS AND COMPUTING↗

Reconfigurable neuromorphic components and algorithms for next-generation artificial intelligence

Digital transistor-based general-purpose hardware (e.g., central processing units) is the dominant solution to support both traditional computing (logic, arithmetic, etc.) as well as modern artificial intelligence. State-of-the-art research has shown feasibility of post-digital physics-based neuromorphic hardware, which is hypothesized to support artificial intelligence algorithms with orders-of-magnitude improved time/energy efficiencies. But such research has not been widely deployed mainly because of such novel hardware’s extreme application-specificity, and the dominance of low-cost general-purpose (but inefficient) digital hardware. To make use of the novel algorithms and the superlative performance of physics-based hardware, we need to identify scientific principles that can enable generality in physics-based hardware. This work resulted in two important broad outcomes – first, we demonstrate fully reconfigurable neuromorphic components, and second, we demonstrate a viable artificial intelligence learning algorithm that can exploit the functioning of neuromorphic hardware. We demonstrate up to five orders of magnitude improvement in energy efficiency compared to the best general-purpose digital hardware.

97 MATHEMATICS AND COMPUTING↗

Parallel sorting algorithm classification: is manual instrumentation necessary?

Understanding parallel algorithms is crucial for accelerating scientific simulations on complex, distributed memory, high-performance computers. Modern algorithm classification approaches learn semantics directly from source code to differentiate between algorithms, however, accessing source code is not always possible. We can learn about parallel algorithms from observing their performance, as programs running the same algorithms and using the same hardware should exhibit similar performance characteristics. We present an approach to learn algorithm classes from parallel performance data directly in order to classify algorithms without access to the source code. We extend previous work to enable classifying parallel sorting algorithms using automatic instrumentation instead of requiring manual region annotations in the source code. In this work, we design and demonstrate a study for classification of parallel sorting algorithms using parallel performance data collected from automatic instrumentation, and evaluate the performance of our new methodology on classification. We leverage Caliper to collect the performance data, Thicket for our exploratory data analysis (EDA), and PyTorch and Scikit-learn to evaluate the effectiveness of random forests, support vector machines (SVMs), decision trees, neural networks, and logistic regressions on parallel performance data. Additionally, we study noise in parallel performance data, whether the removal of noise and pre-processing of the data is necessary to accurately classify parallel sorting algorithms, and determine the effectiveness of features created from performance data. In conclusion, we demonstrate classification accuracy for these five different models of up to 97.7% across four different parallel algorithm classes.

Algorithm Classification↗

Collective neutrino oscillations on a quantum computer with hybrid quantum-classical algorithm

We simulate the time evolution of collective neutrino oscillations in two-flavor settings on a quantum computer. We explore the generalization of Trotter-Suzuki approximation to time-dependent Hamiltonian dynamics. The trotterization steps are further optimized using the Cartan decomposition of two-qubit unitary gates U ϵ SU(4) in the minimum number of controlled-NOT (CNOT) gates making the algorithm more resilient to the hardware noise. As a result, a more efficient hybrid quantum-classical algorithm is also explored to solve the problem on noisy intermediate-scale quantum devices.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Geometry Optimization of Cable-Based Actuation for Small-Scale Model Testing of a Floating Marine Turbine: Preprint

This research aims to apply combined wave and tidal current loads to a small-scale floating marine current turbine in a wave tank, where an actuation system applies hydrodynamic and mooring forces on the hardware based on results from a simulation. We use a real-time hybrid test setup with physical wave forcing from the wave tank and simulated current and mooring forces implemented through a tensioned cable array. For this paper, we developed an optimization algorithm that adjusts the hardware geometry of the cable array to achieve more efficient tension allocation across the cables for the loads that need to be actuated. By adjusting the points where the cables are attached to the floating platform and the angles between the platform and the winches, an optimal cable geometry can be found to minimize tension variations in the lines, maintain the desired pretension, and prevent excessive tensions. We present the optimization problem formulation, the actuation system evaluation approach, and optimization results that show effective cable actuation setups that are being considered for implementation in the wave tank tests.

cable-based actuation↗

Constrained quantum optimization for extractive summarization on a trapped-ion quantum computer

Abstract Realizing the potential of near-term quantum computers to solve industry-relevant constrained-optimization problems is a promising path to quantum advantage. In this work, we consider the extractive summarization constrained-optimization problem and demonstrate the largest-to-date execution of a quantum optimization algorithm that natively preserves constraints on quantum hardware. We report results with the Quantum Alternating Operator Ansatz algorithm with a Hamming-weight-preserving XY mixer (XY-QAOA) on trapped-ion quantum computer. We successfully execute XY-QAOA circuits that restrict the quantum evolution to the in-constraint subspace, using up to 20 qubits and a two-qubit gate depth of up to 159. We demonstrate the necessity of directly encoding the constraints into the quantum circuit by showing the trade-off between the in-constraint probability and the quality of the solution that is implicit if unconstrained quantum optimization methods are used. We show that this trade-off makes choosing good parameters difficult in general. We compare XY-QAOA to the Layer Variational Quantum Eigensolver algorithm, which has a highly expressive constant-depth circuit, and the Quantum Approximate Optimization Algorithm. We discuss the respective trade-offs of the algorithms and implications for their execution on near-term quantum hardware.

97 MATHEMATICS AND COMPUTING↗

A Benthic Habitat Monitoring Approach for Marine and Hydrokinetic Sites (Final Technical Report)

This final technical report summarizes the work completed as part of the Standardized and Cost-Effective Benthic Habitat Mapping and Monitoring Tools for MHK Environmental Assessments project funded by U.S. Department of Energy (DOE) under contract DE-EE007826 to Integral Consulting Inc. The overall goal of this project was to demonstrate a consistent, repeatable, and semi-automated seafloor and sediment mapping approach for rapidly characterizing benthic physical and biological/habitat conditions to support marine environmental assessments for marine and hydrokinetic energy sites. The approach evaluated combines sediment profile imaging and plan view (SPI/PV) technology with multibeam echosounder surveys as an effective and low-cost benthic habitat mapping protocol. A key innovation was the development of a semi-automated computer vision system that standardizes the extraction of data from the SPI/PV images. Other elements of the project were SPI camera hardware modifications, including the design and fabrication of a prototype “power” SPI camera to improve camera prism penetration in firm substrates, and outreach to agency regulators and other stakeholders on this habitat mapping approach. This technical report consists of five main subsections that summarize: 1) the benthic mapping approach; 2) benthic mapping results from the three areas’ surveys; 3) the development and performance of the image processing algorithms; 4) SPI camera hardware improvements and prototype testing; and 5) the regulatory outreach efforts.

16 TIDAL AND WAVE POWER↗

K-Spin Hamiltonian for Quantum-Resolvable Markov Decision Processes

The Markov decision process is the mathematical formalization underlying the modern field of reinforcement learning when transition and reward functions are unknown. We derive a pseudo-Boolean cost function that is equivalent to a K-spin Hamiltonian representation of the discrete, finite, discounted Markov decision process with infinite horizon. This K-spin Hamiltonian furnishes a starting point from which to solve for an optimal policy using heuristic quantum algorithms such as adiabatic quantum annealing and the quantum approximate optimization algorithm on near-term quantum hardware. In arguing that the variational minimization of our Hamiltonian is approximately equivalent to the Bellman optimality condition for a prevalent class of environments we establish an interesting analogy with classical field theory. Along with proof-of-concept calculations to corroborate our formulation by simulated and quantum annealing against classical Q-Learning, we analyze the scaling of physical resources required to solve our Hamiltonian on quantum hardware.

Hamiltonian↗

Exploring the scaling limitations of the variational quantum eigensolver with the bond dissociation of hydride diatomic molecules

Abstract Materials simulations involving strongly correlated electrons pose fundamental challenges to state‐of‐the‐art electronic structure methods but are hypothesized to be the ideal use case for quantum computing algorithms. To date, no quantum computer has simulated a molecule of a size and complexity relevant to real‐world applications, despite the fact that the variational quantum eigensolver (VQE) algorithm can predict chemically accurate total energies. Nevertheless, because of the many applications of moderately sized, strongly correlated systems, such as molecular catalysts, the successful use of the VQE stands as an important waypoint in the advancement toward useful chemical modeling on near‐term quantum processors. In this paper, we take a significant step in this direction. We lay out the steps, write, and run parallel code for an (emulated) quantum computer to compute the bond dissociation curves of the TiH, LiH, NaH, and KH diatomic hydride molecules using the VQE. TiH was chosen as a relatively simple chemical system that incorporates d orbitals and strong electron correlation. Because current VQE implementations on existing quantum hardware are limited by qubit error rates, the number of qubits available, and the allowable gate depth, recent studies using it have focused on chemical systems involving s and p block elements. Through VQE + UCCSD calculations of TiH, we evaluate the near‐term feasibility of modeling a molecule with d‐orbitals on real quantum hardware. We demonstrate that the inclusion of d‐orbitals and the use of the UCCSD ansatz, which are both necessary to capture the correct TiH physics, dramatically increase the cost of this problem. We estimate the approximate error rates necessary to model TiH on current quantum computing hardware using VQE + UCCSD and show them to likely be prohibitive until significant improvements in hardware and error correction algorithms are available.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

MRT 7365: Power flow physics and key physics phenomena

The Z accelerator at Sandia National Laboratories conducts z-pinch experiments at 26 MA in support of DOE missions in stockpile stewardship, dynamic materials, fusion, and other basic sciences. Increasing the current delivered to the z-pinch would extend our reach in each of these disciplines. To achieve increases in current and accelerator efficiency, a fraction of Z’s shots are set aside for research into transmission-line power flow. These shots, with supporting simulations and theory, are incorporated into this Advanced Diagnostics milestone report. The efficiency of Z is reduced as some portion of the total current is shunted across the transmission-line gaps prior to the load. This is referred to as “current loss”. Electrode plasmas have long been implicated in this process, so the bulk of dedicated power-flow experiments are designed to measure the plasma environment. The experimental analyses are enhanced by simulations conducted using realistic hardware and Z voltage pulses. In the same way that diagnostics are continually being improved for sensitivity and resolution, the modeling capability is continually being improved to provide faster and more realistic simulations. The specifics of the experimental hardware, diagnostics, simulations, and algorithm developments are provided in this report. The combined analysis of simulation and data confirms that electrode plasmas have the most detrimental impact on current delivery. Experiments over the last three years have tested the theoretical current-loss mechanisms of enhanced ion current, plasma gap closure, and Hall-related current. These mechanisms are not mutually exclusive and may be coincident in the final feed as well as in upstream transmission lines. The final-feed geometries tested here, however, observe lower-density plasmas without dominant ion currents which is consistent with a Hall-related current. The picture of plasma formation and transport formed from experiment and simulation is informing hardware designs being fielded on Z now and being proposed for the Next-Generation Pulsed Power (NGPP) facility. In this picture, the strong magnetic fields that heat the electrodes above particle emission thresholds also confine the charged particles near the surface. Some portion of the plasmas thus formed is transported into the transmission-line gap under the force of the electric field, with aid from plasma instabilities. The gap plasmas are then transported towards the load by a cross-field drift, where they accumulate and contribute to a likely Hall-related cross-gap current. The achievements in experimental execution, model validation, and physical analysis presented in this report set the stage for continued progress in power flow and load diagnostics on Z. The planned shot schedule for Z and Mykonos will provide data for extrapolation to higher current to ensure the predicted performance and efficiency of a NGPP facility.

43 PARTICLE ACCELERATORS↗

tih_vqe [SWR-23-32]

This software supports the paper, "Exploring the scaling limitations of the variational quantum eigensolver with the bond dissociation of hydride diatomic molecules," published in the International Journal of Quantum Chemistry, whose abstract is as follows: Materials simulations involving strongly correlated electrons pose fundamental challenges to state-of-the-art electronic structure methods but are hypothesized to be the ideal use case for quantum computing. To date, no quantum computer has simulated a molecule of a size and complexity relevant to real-world applications, despite the fact that the variational quantum eigensolver (VQE) algorithm can predict chemically accurate total energies. Nevertheless, because of the many applications of moderately-sized, strongly correlated systems, such as molecular catalysts, the successful use of the VQE stands as an important waypoint in the advancement toward useful chemical modeling on near-term quantum processors. In this paper, we take a significant step in this direction. We lay out the steps, write, and run parallel code for an (emulated) quantum computer to compute the bond dissociation curves of the TiH, LiH, NaH, and KH diatomic hydride molecules using VQE. TiH was chosen as a relatively simple chemical system that incorporates d orbitals and strong electron correlation. Because current VQE implementations on existing quantum hardware are limited by qubit error rates, the number of qubits available, and the allowable gate depth, recent studies have focused on chemical systems involving s and p block elements. Through VQE + UCCSD calculations of TiH, we evaluate the near-term feasibility of modeling a molecule with d-orbitals on real quantum hardware. We demonstrate that the inclusion of d-orbitals and the use of the UCCSD ansatz, which are both necessary to capture the correct TiH physics, dramatically increase the cost of this problem. We estimate the approximate error rates necessary to model TiH on current quantum computing hardware using VQE+UCCSD and show them to likely be prohibitive until significant improvements in hardware and error correction algorithms are available.

Graf, Peter↗

Performance Evaluation of Intelligent Solar Control Software Through Hardware-in-the-Loop (CRADA Final Report)

Recent research has highlighted the potential for solar to act as a zero-marginal-cost and zero-emission flexibility resource on the bulk power system when operated with advanced control systems. To increase the performance of these systems, leading technologies, including machine learning (ML) and hierarchical inverter set point allocation, have been developed by Latimer Controls, Inc. to estimate the headroom of large PV plants for grid operation and control; however, these technologies lack comprehensive validation under real-world application scenarios. Latimer Controls, Inc. received two voucher awards for research at a national laboratory from the Department of Energy American Made Solar Prize Round 6. The National Renewable Energy Laboratory (NREL) was selected to collaborate with Latimer staff to conduct a performance evaluation of Latimer PV control software. The NREL team will develop a hardware-in-the-loop (HIL) testbed to perform testing and validation of the Latimer PV control technology in a de-risked yet realistic testbed environment. Latimer and NREL worked together to analyze the test data, draw conclusions from the results, and disseminate the resulting scientific findings. In this CRADA work, we propose to test and validate the real-world application of the Latimer Control solution in an HIL environment. We evaluate the performance of different flexible solar technologies in responding to automatic generation control signals in a closed-loop fashion. In particular, a data-driven potential high limit (PHL) estimation is developed for large solar plants to accurately estimate their headroom so that they have fast and short-time regulation and control capability to participate in grid services and respond to grid signals in real time (e.g., AGC). This PHL estimation algorithm is embedded in a hardware power plant controller (PPC) and tested with an IEEE-39 bus system model developed in RTDS. To account for the varying cloud conditions and diverse inverter dispatches, we developed a 135-MW PV plant with detailed modeling of 27 individual PV modules and inverters using RTDS. The real-world communications used in such big plants, such as ModBus TCP/IP for inverter level and DNP3 for plant level, were developed to emulate the real-world applications in big PV plants. The ML-based PHL estimation method is tested under nine separate weather scenarios against the ‘reference-control’ solution, hereafter referred to as the baseline solution. The baseline method reserves a subset of inverters (reference group) to operate at their PHL at all times and dispatches only the remaining inverters (control group) at curtailed levels to fulfill the flexibility need. Despite being successfully piloted by NREL in California in 2017 and Chile in 2020, there exist two gaps in the state of the art to fully unlock the flexibility of PV plants: a. There is a trade-off between the PHL estimation accuracy and the flexibility range. b. There lacks granularity in the PHL estimation to capture the variation across inverters. The Latimer solution seeks to address these gaps by applying machine learning methods to improve PHL estimation accuracy while accounting for variability at every inverter. Performance metrics were taken from the 2023 Georgia Power CARES utility-scale RFP. The results demonstrate that the ML-based approach outperforms the traditional baseline method in PHL estimation accuracy for 7 of 9 scenarios. The average PHL error across the nine scenarios was 7.40% for the ML-based method, 2.06% less than the 9.46% PHL error average across scenarios that was exhibited by the baseline method. Additionally, the PHL error was below 5% for at least 95% of the testing interval for 3 of 9 tested intervals with the ML approach, whereas it did not achieve this metric for any of the baseline tests. Overall, simulation results indicate the superior performance of an ML-based approach compared to the conventional baseline reference-control approach, showcasing its potential to support grid stability and operational efficiency. This laboratory HIL testing using real PPC, representative power system simulation models in real-time with detailed PV plant and inverter models, and real-world communication protocols gives us confidence that this machine learning based PHL estimation algorithm works well in the hardware PPC and therefore de-risks future field commissioning. The end goal of this project is to advance grid technology to address the grid operation challenges brought by solar plant’s variability and uncertainties in power generation.

14 SOLAR ENERGY↗

Flexible silicon photonic architecture for accelerating distributed deep learning

The increasing size and complexity of deep learning (DL) models have led to the wide adoption of distributed training methods in datacenters (DCs) and high-performance computing (HPC) systems. However, communication among distributed computing units (CUs) has emerged as a major bottleneck in the training process. In this study, we propose Flex-SiPAC, a flexible silicon photonic accelerated compute cluster designed to accelerate multi-tenant distributed DL training workloads. Flex-SiPAC takes a co-design approach that combines a silicon photonic hardware platform with a tailored collective algorithm, optimized to leverage the unique physical properties of the architecture. The hardware platform integrates a novel wavelength-reconfigurable transceiver design and a micro-resonator-based wavelength-reconfigurable switch, enabling the system to achieve flexible bandwidth steering in the wavelength domain. The collective algorithm is designed to support reconfigurable topologies, enabling efficient all-reduce communications that are commonly used in DL training. The feasibility of the Flex-SiPAC architecture is demonstrated through two testbed experiments. First, an optical testbed experiment demonstrates the flexible routing of wavelengths by shuffling an array of input wavelengths using a custom-designed spatial-wavelength selective switch. Second, a four-GPU testbed running two DL workloads shows a 23% improvement in job completion time compared to a similarly sized leaf-spine topology. We further evaluate Flex-SiPAC using large-scale simulations, which show that Flex-SiPAC is able to reduce the communication time by 26% to 29% compared to state-of-the-art compute clusters under representative collective operations.

Wu, Zhenguo (ORCID:0000000322847985)↗