Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “hardware algorithms”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

SpecFIDLER User Manual (Software V.2.6.0)

The Spectroscopic Field Instrument for Detection of Low Energy Radiation (SpecFIDLER) allows response teams to detect and quantify plutonium contamination on the ground. Notional scenarios include dispersion from a weapon accident, or the launch failure of a space probe containing a radioisotope thermoelectric generator. Unlike other instruments, the thin-window sodium iodide detector is sensitive to the low-energy gamma rays emitted by plutonium isotopes. The system supports both mobile survey as well as stationary sampling. This manual provides information about installing, maintaining, and troubleshooting the SpecFIDLER. The scope of this document includes the physical hardware, software for data acquisition, and algorithms for data analysis. Recent changes to the software and algorithms aim to streamline the operation of the system.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

SpecFIDLER User Manual (Software V.2.3.0)

The Spectroscopic Field Instrument for Detection of Low Energy Radiation (SpecFIDLER) allows response teams to detect and quantify plutonium contamination on the ground. Notional scenarios include dispersion from a weapon accident, or the launch failure of a space probe containing a radioisotope thermoelectric generator. Unlike other instruments, the thin-window sodium iodide detector is sensitive to the low-energy gamma rays emitted by plutonium isotopes. The system supports both mobile survey as well as stationary sampling. This manual provides information about installing, maintaining, and troubleshooting the SpecFIDLER. The scope of this document includes the physical hardware, software for data acquisition, and algorithms for data analysis. Recent changes to the software and algorithms aim to streamline the operation of the system.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Demonstrating autonomous controls on hardware test beds is a necessity for successful missions to Mars and beyond

NASA and the Department of Defense are planning for a mission to Mars in the 2030s–2040s using nuclear thermal propulsion (NTP). NTP uses a nuclear reactor to heat flowing hydrogen and create thrust. A serious concern for crewed and uncrewed missions to Mars is the loss of reactor control. The reactor startup and initial rocket impulse are initiated in cislunar or near-earth orbital regions; therefore, radio communications between ground control and the NTP engine should occur in real time. However, radio communications can take more than 20 min, depending on planet positions, to reach Mars orbiters from ground control. To address this delay, local autonomous controls are implemented onboard the NTP engine to ensure acceptable operation. However, autonomous controls have not been demonstrated or implemented in research or power reactor contexts because of safety and reliability concerns. To enable autonomous controls development, demonstration, and validation, Oak Ridge National Laboratory has created a nonnuclear hardware-in-the-loop test bed. Sensors throughout the test bed relay system status and hardware response to the user control algorithm, including measurements of temperature, flow, pressure of a loop, control drum position, and drum speed. This paper discusses the development of this facility and user accessibility.

33 ADVANCED PROPULSION SYSTEMS↗

Nuclear Thermal Rocket Emulator for a Hardware-in-the-Loop Test Bed

To support NASA’s mission to use nuclear thermal rockets for future Mars missions, an instrumentation and control test bed has been built at Oak Ridge National Laboratory. The system is designed as a hardware-in-the-loop test bed for testing control elements and autonomous control algorithms for nuclear thermal propulsion rockets. The mock reactor system consists of a modular and scalable framework, using inexpensive components and open-source software. The hardware system consists of a two-phase flow loop and a mock reactor with six control drums. A single-board computer (NVIDIA Jetson) handles reactor core emulation and hosts a message queuing telemetry transport broker that allows user-deployed control algorithms to interact with the system hardware. The reactor emulator receives sensor data from the hardware and provides the simulated performance of the reactor under steady-state, transient, and fault conditions. The emulator uses a reactivity lookup table and the point kinetics equations to solve for the reactor dynamics in real time. Emulated reactor dynamics and sensor input inform the autonomous control algorithm’s decision-making in a closed-loop manner. The current system is capable of operating at 10 Hz, but faster cycle rates are an area of ongoing research. This test bed will enable NASA and other space vendors to rigorously test their autonomous control systems for NTP rockets under transient (reactor startup and shutdown), steady-state, and fault conditions to reduce development time and risk for autonomous control systems in future missions.

autonomous control↗

Graph decomposition techniques for solving combinatorial optimization problems with variational quantum algorithms

The quantum approximate optimization algorithm (QAOA) has the potential to approximately solve complex combinatorial optimization problems in polynomial time. However, current noisy quantum devices cannot solve large problems due to hardware constraints. In this work, we develop an algorithm that decomposes the QAOA input problem graph into a smaller problem and solves MaxCut using QAOA on the reduced graph. The algorithm requires a subroutine that can be classical or quantum—in this work, we implement the algorithm twice on each graph. One implementation uses the classical solver Gurobi in the subroutine and the other uses QAOA. We solve these reduced problems with QAOA. On average, the reduced problems require only approximately 1/10 of the number of vertices than the original MaxCut instances. Furthermore, the average approximation ratio of the original MaxCut problems is 0.75, while the approximation ratios of the decomposed graphs are on average of 0.96 for both Gurobi and QAOA. With this decomposition, we are able to measure optimal solutions for ten 100-vertex graphs by running single-layer QAOA circuits on the Quantinuum trapped-ion quantum computer H1-1, sampling each circuit only 500 times. This approach is best suited for sparse, particularly k-regular graphs, as k-regular graphs on n vertices can be decomposed into a graph with at most $\frac{nk}{k+1}$ vertices in polynomial time. Further reductions can be obtained with a potential trade-off in computational time. In conclusion, while this paper applies the decomposition method to the MaxCut problem, it can be applied to more general classes of combinatorial optimization problems.

97 MATHEMATICS AND COMPUTING↗

Noise-induced transition in optimal solutions of variational quantum algorithms

Variational quantum algorithms are promising candidates for delivering practical quantum advantage on noisy intermediate-scale quantum (NISQ) hardware. However, optimizing the noisy cost functions associated with these algorithms is challenging for system sizes relevant to quantum advantage. In this work, we investigate the effect of noise on optimization by studying a variational quantum eigensolver (VQE) algorithm calculating the ground state of a spin chain model, and we observe an abrupt transition induced by noise to the optimal solutions. We will present numerical simulations, a demonstration using an IBM quantum processor unit (QPU), and a theoretical analysis indicating the origin of this transition. Our findings suggest that careful analysis is crucial to avoid misinterpreting the noise-induced features as genuine algorithm results.

Li, Andy C.Y.↗

A Software Framework for Comparing Training Approaches for Spiking Neuromorphic Systems

There are a wide variety of training approaches for spiking neural networks for neuromorphic deployment. However, it is often not clear how these training algorithms perform or compare when applied across multiple neuromorphic hardware platforms and multiple datasets. In this work, we present a software framework for comparing performance across four neuromorphic training algorithms across three neuromorphic simulators and four simple classification tasks. We introduce an approach for training a spiking neural network using a decision tree, and we compare this approach to training algorithms based on evolutionary algorithms, back-propagation, and reservoir computing. We present a hyperparameter optimization approach to tune the hyperparameters of the algorithm, and show that these optimized hyperparameters depend on the processor, algorithm, and classification task. Finally, we compare the performance of the optimized algorithms across multiple metrics, including accuracy, training time, and resulting network size, and we show that there is not one best training algorithm across all datasets and performance metrics.

Schuman, Catherine↗

Stochastic noise can be helpful for variational quantum algorithms

Saddle points constitute a crucial challenge for first-order gradient descent algorithms. In notions of classical machine learning, they are avoided, for example, by means of stochastic gradient descent methods. In this work, we provide evidence that the saddle-points problem can be naturally avoided in variational quantum algorithms by exploiting the presence of stochasticity. We prove convergence guarantees and present practical examples in numerical simulations and on quantum hardware. We argue that the natural stochasticity of variational algorithms can be beneficial for avoiding strict saddle points, i.e., those saddle points with at least one negative Hessian eigenvalue. This insight that some levels of shot noise could help is expected to add a new perspective to notions of near-term variational quantum algorithms. Published by the American Physical Society 2025

Liu, Junyu↗

The JSNS 2 data acquisition system

The JSNS 2 (J-PARC Sterile Neutrino Search at J-PARC Spallation Neutron Source) experiment aims to search for neutrino oscillations over a 24 m short baseline at J-PARC. The JSNS 2 inner detector is filled with 17 tons of gadolinium(Gd)-loaded liquid scintillator (LS) with an additional 31 tons of unloaded LS in the intermediate γ-catcher and an optically separated outer veto volumes. A total of 120 10-inch photomultiplier tubes observe the scintillating optical photons and each analog waveform is stored with flash analog-to-digital converters. Here, we present details of the data acquisition, processing, and data quality monitoring system. We also present two different trigger logics which are developed for the beam and self-trigger.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Hls4ml Synthesis Testing

HLS4ml (high level synthesis for machine learning) Is a Python package used to translate commonly used open-source machine learning models into HLS. This is useful in machine learning applications on FPGAs. Machine learning algorithms are only as fast as the hardware that they are used on, and some applications require high speed without sacrificing accuracy. In these situations, an FPGA is a good choice since it is faster than a CPU or a GPU, but programming an FPGA is difficult. This is where HLS4ml can be used to simplify the process, as a well-known learning model can be converted to HLS and more easily deployed onto an FPGA. There are many use cases for a machine learning algorithm running on an FPGA. For example, detectors in a particle accelerator cannot keep every event that they detect, and so a computer must decide which events to keep and which to discard. Using an FPGA with a machine learning algorithm would be a good way to keep as many events as possible.

Swanson, Caiden↗

hls4ml

hls4ml (high level synthesis for machine learning) Is a Python package used to translate commonly used open-source machine learning models into HLS. This is useful in machine learning applications on FPGAs. Machine learning algorithms are only as fast as the hardware that they are used on, and some applications require high speed without sacrificing accuracy. In these situations, an FPGA is a good choice since it is faster than a CPU or a GPU, but programming an FPGA is difficult. This is where hls4ml can be used to simplify the process, as a well-known learning model can be converted to HLS and more easily deployed onto an FPGA. There are many use cases for a machine learning algorithm running on an FPGA. For example, detectors in a particle accelerator cannot keep every event that they detect, and so a computer must decide which events to keep and which to discard. Using an FPGA with a machine learning algorithm would be a good way to keep as many events as possible.

Swanson, Caiden↗

Enabling Scalable VQE Simulation on Leading HPC Systems

Large-scale simulations of quantum circuits pose significant challenges, especially in the context of quantum chemistry, due to the number of qubits, circuit depth, and the number of circuits needed per problem. High-performance computing (HPC) systems offer massive computational capabilities that could help overcome these obstacles. We developed a high-performance quantum circuit simulator, called NWQ-Sim, and demonstrate its capability to simulate large quantum chemistry problems on NERSC's Perlmutter supercomputer. Integrating NWQ-Sim with XACC, we have executed QPE and VQE algorithms for downfolded quantum chemistry systems at unprecedented scales. Our work demonstrates the potential of leveraging HPC resources to advance quantum chemistry and other applications of near-term quantum devices.

Wang, Meng↗

Even Higher-Level Synthesis: An Exploration of AI Hardware Accelerators using HLS4ML

With the rise of artificial intelligence, the popularization of deep learning, and a constantly evolving industry, the demand for flexible and efficient tools has never been greater. As algorithms grow more complex, their runtime and energy consumption increase exponentially. Customized hardware accelerators, long used for specific mathematical operations, remain essential for managing modern applications' computational and power demands. Hardware accelerators can speed up complex computations by orders of magnitude, but their manual design and verification processes are often challenging and time-consuming. High-Level Synthesis (HLS) provides a solution by transforming high-level algorithm descriptions, typically written in C++ or SystemC, into synthesizable RTL suitable for hardware implementation. This approach reduces development time for RTL engineers while offering flexibility beyond what traditional handwritten RTL can provide. We extended this capability to the machine-learning domain with the open-source framework hls4ml, which allows neural networks trained in Python frameworks like Tensorflow or PyTorch to be synthesized into efficient hardware representations for the traditional FPGA and ASIC flows. This breakthrough addresses the growing need for reduced design turnaround and easy verification of ML hardware accelerators with low latency and power efficiency constraints. During this tutorial, we will demonstrate how Python complements HLS by simplifying the ML design process, bridging the gap between software and hardware development. Attendees will explore how we translate neural networks modeled in Python into fixed-point C++ models suitable for HLS workflows. We will dive into strategies like Value-Range Analysis and Quantization-Aware Training, which optimize these designs for deployment and evaluate their accuracy, power consumption, and energy efficiency. To exemplify these concepts, experts from Fermilab will share their experiences applying this technology to high-energy physics experiments, where real-time, low-latency processing is critical. Over the years, Fermilab engineers have demonstrated how deep neural networks, optimized for hardware using hls4ml, can meet the stringent requirements of trigger systems at the CERN Large Hadron Collider. These systems rely on rapid decision-making to process immense data volumes while retaining only the most relevant events for further analysis. The application of hls4ml has also been extended to innovative technologies like smart pixel arrays. These smart pixels integrate ML inference capabilities directly into sensor devices, enabling localized data processing at the pixel level. This approach drastically reduces the need to transmit raw data to external processing units, significantly decreasing power consumption and latency. By embedding neural networks within the pixel architecture, the smart pixels can identify and prioritize relevant data in real time, providing a highly efficient solution for edge computing in scenarios such as particle detectors and imaging systems. Fermilab's work highlights the potential of hardware-accelerated ML in scenarios where both speed and power efficiency are mission-critical. Through this tutorial, attendees will gain valuable insights into the challenges and solutions of deploying ML in hardware. Understanding how HLS and hls4ml streamline the development of neural network-based hardware accelerators is fundamental for the industry's future. Participants will learn how these technologies are shaping the future of AI and scientific computing.

Di Guglielmo, Giuseppe [Fermilab]↗

Performance Evaluation of Distributed Energy Resource Management Algorithm in Large Distribution Networks

This paper presents performance evaluation of hierarchical optimization and control for distributed energy resource management system (DERMS) in large distribution networks via an advanced hardware-in-the-loop (HIL) platform. The HIL platform provides realistic testing in a laboratory envi-ronment, including the accurate modeling of a full-scale dis-tribution system of 11,000 nodes, the DERMS software con-troller, and 90 power hardware photovoltaics (PVs) and bat-tery inverters. The applied DERMS algorithm is designed based on a real-time optimal power flow algorithm and im-plemented with acceleration design that performs fast dis-patch of simulated PVs and real physical hardware DER devices every 4 seconds.

distributed energy management system (DERMS)↗

Determining quantum phase diagrams of topological Kitaev-inspired models on NISQ quantum hardware

Topological protection is employed in fault-tolerant error correction and in developing quantum algorithms with topological qubits. But, topological protection intrinsic to models being simulated , also robustly protects calculations, even on NISQ hardware. We leverage it by simulating Kitaev-inspired models on IBM quantum computers and accurately determining their phase diagrams. This requires constructing conventional quantum circuits for Majorana braiding to prepare the ground states of Kitaev-inspired models. The entanglement entropy is then measured to calculate the quantum phase boundaries. We show how maintaining particle-hole symmetry when sampling through the Brillouin zone is critical to obtaining high accuracy. This work illustrates how topological protection intrinsic to a quantum model can be employed to perform robust calculations on NISQ hardware, when one measures the appropriate protected quantum properties. It opens the door for further simulation of topological quantum models on quantum hardware available today.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Particle flow reconstruction for the CMS Phase-II Level-1 Trigger

The upgrade of the CMS detector for the high-luminosity LHC will include trackfinding for the first time in the Level-1 trigger, enabling Particle Flow reconstruction of every event in addition to comprehensive pileup mitigation. The Correlator trigger will reconstruct isolated leptons and photons, hadronic jets, and energy sums, assisted in many cases by machine learning to benefit from the complete particle-level event record. Here, we present the logic of these algorithms, possible implementations using large FPGAs and their demonstration in prototype hardware, in addition to the expected physics performance.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Quantum optimization of maximum independent set using Rydberg atom arrays

Realizing quantum speedup for practically relevant, computationally hard problems is a central challenge in quantum information science. Using Rydberg atom arrays with up to 289 qubits in two spatial dimensions, we experimentally investigate quantum algorithms for solving the maximum independent set problem. We use a hardware-efficient encoding associated with Rydberg blockade, realize closed-loop optimization to test several variational algorithms, and subsequently apply them to systematically explore a class of graphs with programmable connectivity. We find that the problem hardness is controlled by the solution degeneracy and number of local minima, and we experimentally benchmark the quantum algorithm’s performance against classical simulated annealing. On the hardest graphs, we observe a superlinear quantum speedup in finding exact solutions in the deep circuit regime and analyze its origins.

Science & Technology - Other Topics↗