Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “hardware algorithms”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 451 records · Page 25

Performance Evaluation of Distributed Energy Resource Management Algorithm in Large Distribution Networks

This paper presents performance evaluation of hierarchical optimization and control for distributed energy resource management system (DERMS) in large distribution networks via an advanced hardware-in-the-loop (HIL) platform. The HIL platform provides realistic testing in a laboratory environment, including the accurate modeling of a full-scale distribution system of 11,000 nodes, the DERMS software controller, and 90 power hardware photovoltaics (PVs) and battery inverters. The applied DERMS algorithm is designed based on a realtime optimal power flow algorithm and implemented with acceleration design that performs fast dispatch of simulated PVs and real physical hardware DER devices every 4 seconds.

DERMS↗

A hybrid six degree-of-freedom tracking system

An optical/digital/mechanical six degree of freedom tracking system using an optical correlator for image processing is being constructed at NASA's Johnson Space Center. The degrees of freedom are expressed in sensor coordinates as azimuth, elevation, range, line-of-sight rotation, and the two out-of-plane object rotation angles. Hardware for an initial configuration has been assembled and various tracking algorithms and filtering techniques are being implemented and evaluated. The current correlator hardware is based on LCTV SLMs from a commercial television projector. Correlation peak detection and measurement are made using commercially available digital image processing boards. Out-of-plane object rotation, range, and line-of-sight rotation are tracked by various correlation filter techniques. Performance of the current system is presented, as are plans for future configurations.

Monroe, Stanley E., Jr.↗

Union: A Unified HW-SW Co-Design Ecosystem in MLIR for Evaluating Tensor Operationson Spatial Accelerators

To meet the extreme compute demands for deep learning across commercial and scientific applications, dataflow accelerators are becoming increasingly popular. While these“domain-specific” accelerators are not fully programmable like CPUs and GPUs, they retain varying levels of flexibility with respect to data orchestration, i.e., dataflow and tiling optimizations to enhance efficiency. There are several challenges when designing new algorithms and mapping approaches to execute the algorithms for a target problem on new hardware. Previous works have addressed these challenges individually. To address this challenge as a whole, in this work, we present an HW-SW co-design ecosystem for spatial accelerators called Union within the popular MLIR compiler infrastructure. Our framework allows exploring different algorithms and their mappings on several accelerator cost models. Union also includes a plug-and-play library of accelerator cost models and mappers which can easily be extended. The algorithms and accelerator cost models are connected via a novel mapping abstraction that captures the map space of spatial accelerators which can be systematically pruned based on constraints from the hardware, workload, and mapper. We demonstrate the value of Union for the community with several case studies which examine offloading different tensor operations (CONV/GEMM/Tensor Contraction) on diverse accelerator architectures using different mapping schemes.

Jeong, Geonhwa↗

State Dependent Optimization with Quantum Circuit Cutting

Quantum circuits can be reduced through optimization to better fit the constraints of quantum hardware. One such method, initial-state dependent optimization (ISDO), reduces gate count by leveraging knowledge of the input quantum states. Surprisingly, we found that ISDO is broadly applicable to the downstream circuits produced by circuit cutting. Circuit cutting also requires measuring upstream qubits and has some flexibility of selection observables to do reconstruction. Therefore, we propose a state-dependent optimization (SDO) framework that incorporates ISDO, our newly proposed measure-state dependent optimization (MSDO), and a biased observable selection strategy. Building on the strengths of the SDO framework and recognizing the scalability challenges of circuit cutting, we propose nonseparate circuit cutting-a more flexible approach that enables optimizing gates without fully separating them. We validate our methods on noisy simulations of QAOA, QFT, and BV circuits. Results show that our approach consistently mitigates noise and improves overall circuit performance, demonstrating its promise for enhancing quantum algorithm execution on near-term hardware.

Li, Xinpeng↗

A survey on the design of multiprocessing systems for artificial intelligence applications

Some issues in designing computers for artificial intelligence (AI) processing are discussed. These issues are divided into three levels: the representation level, the control level, and the processor level. The representation level deals with the knowledge and methods used to solve the problem and the means to represent it. The control level is concerned with the detection of dependencies and parallelism in the algorithmic and program representations of the problem, and with the synchronization and sheduling of concurrent tasks. The processor level addresses the hardware and architectural components needed to evaluate the algorithmic and program representations. Solutions for the problems of each level are illustrated by a number of representative systems. Design decisions in existing projects on AI computers are classed into top-down, bottom-up, and middle-out approaches.

Wah, Benjamin W.↗

Bio-Inspired Neural Model for Learning Dynamic Models

A neural-network mathematical model that, relative to prior such models, places greater emphasis on some of the temporal aspects of real neural physical processes, has been proposed as a basis for massively parallel, distributed algorithms that learn dynamic models of possibly complex external processes by means of learning rules that are local in space and time. The algorithms could be made to perform such functions as recognition and prediction of words in speech and of objects depicted in video images. The approach embodied in this model is said to be "hardware-friendly" in the following sense: The algorithms would be amenable to execution by special-purpose computers implemented as very-large-scale integrated (VLSI) circuits that would operate at relatively high speeds and low power demands.

Duong, Tuan↗

ECP ST Capability Assessment Report (CAR) for VTK-m (FY20)

The ECP/VTK-m project is providing the core capabilities to perform scientific visualization on Exascale architectures. The ECP/VTK-m project fills the critical feature gap of performing visualization and analysis on processors like graphics-based processors. The results of this project will be delivered in tools like ParaView, Vislt, and Ascent as well as in stand-alone form. Moreover, these projects are depending on this ECP effort to be able to make effective use of ECP architectures. One of the biggest recent changes in high-performance computing is the increasing use of accelerators. Accelerators contain processing cores that independently are inferior to a core in a typical CPU, but these cores are replicated and grouped such that their aggregate execution provides a very high computation rate at a much lower power. Current and future CPU processors also require much more explicit parallelism. Each successive version of the hardware packs more cores into each processor, and technologies like hyper threading and vector operations require even more parallel processing to leverage each core's full potential. VTK-m is a toolkit of scientific visualization algorithms for emerging processor architectures. VTK-m supports the fine-grained concurrency for data analysis and visualization algorithms required to drive extreme scale computing by providing abstract models for data and execution that can be applied to a variety of algorithms across many different processor architectures. The ECP/VTK-m project is building up the VTK-m codebase with the necessary visualization algorithm implementations that run across the varied hardware platforms to be leveraged at the Exascale. We will be working with other ECP projects, such as ALPINE, to integrate the new VTK-m code into production software to enable visualization on our HPC systems.

97 MATHEMATICS AND COMPUTING↗

Magellan SAR processing algorithm and H/W design

The SAR (synthetic-aperture radar) data-processing algorithm to be used for the Magellan mission is described. Radar system design, SAR data characteristics, and hardware (H/W) constraints, which are critical to the processing algorithm design, are highlighted. Data flow and the H/W architecture are given to show the real-time data processing capability. Simulation results obtained from processing the synthetic point-target echos are presented to demonstrate the performance of the processing algorithm.

Chen, M.↗

High performance compression of science data

In the future, NASA expects to gather over a tera-byte per day of data requiring space for levels of archival storage. Data compression will be a key component in systems that store this data (e.g., optical disk and tape) as well as in communications systems (both between space and Earth and between scientific locations on Earth). We propose to develop algorithms that can be a basis for software and hardware systems that compress a wide variety of scientific data with different criteria for fidelity/bandwidth tradeoffs. The algorithmic approaches we consider are specially targeted for parallel computation where data rates of over 1 billion bits per second are achievable with current technology.

Storer, James A.↗

High performance compression of science data

In the future, NASA expects to gather over a tera-byte per day of data requiring space for levels of archival storage. Data compression will be a key component in systems that store this data (e.g., optical disk and tape) as well as in communications systems (both between space and Earth and between scientific locations on Earth). We propose to develop algorithms that can be a basis for software and hardware systems that compress a wide variety of scientific data with different criteria for fidelity/bandwidth tradeoffs. The algorithmic approaches we consider are specially targeted for parallel computation where data rates of over 1 billion bits per second are achievable with current technology.

Storer, James A.↗

Injecting Errors for Testing Built-In Test Software

Two algorithms have been conceived to enable automated, thorough testing of Built-in test (BIT) software. The first algorithm applies to BIT routines that define pass/fail criteria based on values of data read from such hardware devices as memories, input ports, or registers. This algorithm simulates effects of errors in a device under test by (1) intercepting data from the device and (2) performing AND operations between the data and the data mask specific to the device. This operation yields values not expected by the BIT routine. This algorithm entails very small, permanent instrumentation of the software under test (SUT) for performing the AND operations. The second algorithm applies to BIT programs that provide services to users application programs via commands or callable interfaces and requires a capability for test-driver software to read and write the memory used in execution of the SUT. This algorithm identifies all SUT code execution addresses where errors are to be injected, then temporarily replaces the code at those addresses with small test code sequences to inject latent severe errors, then determines whether, as desired, the SUT detects the errors and recovers

Gender, Thomas K.↗

Inferring the Dynamics of the State Evolution During Quantum Annealing

To solve an optimization problem using a commercial quantum annealer, one has to represent the problem of interest as an Ising or a quadratic unconstrained binary optimization (QUBO) problem and submit its coefficients to the annealer, which then returns a user-specified number of low-energy solutions. It would be useful to know what happens in the quantum processor during the anneal process so that one could design better algorithms or suggest improvements to the hardware. However, existing quantum annealers are not able to directly extract such information from the processor. Hence, in this work we propose to use advanced features of D-Wave 2000Q to indirectly infer information about the dynamics of the state evolution during the anneal process. Specifically, D-Wave 2000Q allows the user to customize the anneal schedule, that is, the schedule with which the anneal fraction is changed from the start to the end of the anneal. Furthermore, using this feature, we design a set of modified anneal schedules whose outputs can be used to generate information about the states of the system at user-defined time points during a standard anneal. With this process, called "slicing", we obtain approximate distributions of lowest-energy anneal solutions as the anneal time evolves. We use our technique to obtain a variety of insights into the annealer, such as the state evolution during annealing, when individual bits in an evolving solution flip during the anneal process and when they stabilize, and we introduce a technique to estimate the freeze-out point of both the system as well as of individual qubits.

42 ENGINEERING↗

Multiple Uplinks Per Antenna (MUPA) Signal Acquisition Schemes

The Deep Space Network (DSN) currently makes use of the technique of Multiple Spacecraft per Antenna (MSPA) where a single antenna is used to track multiple spacecraft downlinks within its beam, such as in the case of multiple spacecraft orbiting Mars at 8.4 GHz (X-band). It is desired to extend this technique to the uplink where a single station is used to send a signal to multiple spacecraft in order to make more efficient use of ground resources. This would be applicable to numerous smallsat constellations being considered for future missions or to future spacecraft at Venus, Mars, or more distant destinations that are all within the half-power beamwidth of a single 34-m diameter antenna. In one scheme, each spacecraft’s command sequences would be time multiplexed onto a single uplink frequency. Each spacecraft would lock onto the uplink signal and would accept only commands intended for it via special identifier codes. Each spacecraft would also emit a downlink signal to the ground that is coherent with the uplink signal but would have its own allocated frequency channel and identifier information. A couple of key challenges associated with using this technique need to be addressed. Because of the single uplink frequency, coherent turnaround for two-way Doppler and ranging would not conform to established ratios, thus the radios employed by the spacecraft would need to be capable of variable turnaround ratios. In addition, because of the different orbits or spacecraft trajectories, the relative Doppler shifts and rates can be large with respect to the common uplink signal whose frequency would lie at the centroid of the frequencies of the expected received signals of the constellation. This would be problematic with standard analog spacecraft radios whose acquisition bandwidths are relatively small (~1.7 kHz) relative to the large frequency offsets (~100 kHz) expected using the single frequency uplink technique. With the advent of software defined radios (SDRs), signal frequency search algorithms can be utilized within the flight software and/or programmable hardware (e.g., FPGAs) that can easily acquire and track signals with large frequency offsets and varying dynamics. Such techniques could include FFT search algorithms, step-and-sweep search algorithms, or onboard frequency steering making use of trajectory vectors uplinked to each member spacecraft. Other challenges include mitigation of potential interference between received signals. We have identified several software defined radios that are in different stages of development and whose key parameters have been tabulated. We have examined each radio’s capabilities with respect to acquiring and tracking signals with large frequency offsets. Such analyses made use of previous studies supplemented with specially designed tests using both simulation tools and/or existing testbeds. We have compared signal acquisition times computed from provided algorithms along with measured values derived from tests using existing hardware and simulation tools for the purpose of conducting tradeoff studies between the various radio designs and software/firmware programming approaches.

Abraham, Douglas S.↗

Quantum Reinforcement Learning for Volt-VAR Control in Power Distribution Systems

Volt-VAR control (VVC) is crucial in active distribution networks for optimizing voltage profiles and minimizing network losses. While traditional deep reinforcement learning (DRL) algorithms exhibit promise for VVC, they often require extensive computational resources to handle such a high-dimensional problem. As a potential solution, quantum reinforcement learning (QRL) algorithms integrate the computational capabilities of quantum computing into the DRL framework. However, existing QRL algorithms struggle with complex VVC problems due to the limitations of current quantum hardware. To bridge this gap, this paper proposes an innovative QRL algorithm featuring an end-to-end architecture that integrates a classical autoencoder, variational quantum circuits (VQCs), and classical post-processing layers. This design efficiently compresses high-dimensional grid states, enabling VQCs to leverage quantum advantages while producing multiple control device outputs tailored for VVC tasks. Numerical studies on three representative distribution systems verify the effectiveness and scalability of the proposed QRL algorithm, and demonstrate its enhanced performance over classical approaches with only approximately 1% of the parameters. Additionally, the robustness of our developed algorithm is validated through noisy quantum environments.

97 MATHEMATICS AND COMPUTING↗

Addressing the Real-World Challenges in the Development of Propulsion IVHM Technology Experiment (PITEX)

The Propulsion IVHM Technology Experiment (PITEX) has been an on-going research effort conducted over several years. PITEX has developed and applied a model-based diagnostic system for the main propulsion system of the X-34 reusable launch vehicle, a space-launch technology demonstrator. The application was simulation-based using detailed models of the propulsion subsystem to generate nominal and failure scenarios during captive carry, which is the most safety-critical portion of the X-34 flight. Since no system-level testing of the X-34 Main Propulsion System (MPS) was performed, these simulated data were used to verify and validate the software system. Advanced diagnostic and signal processing algorithms were developed and tested in real-time on flight-like hardware. In an attempt to expose potential performance problems, these PITEX algorithms were subject to numerous real-world effects in the simulated data including noise, sensor resolution, command/valve talkback information, and nominal build variations. The current research has demonstrated the potential benefits of model-based diagnostics, defined the performance metrics required to evaluate the diagnostic system, and studied the impact of real-world challenges encountered when monitoring propulsion subsystems.

Maul, William A.↗

Accelerators for Classical Molecular Dynamics Simulations of Biomolecules

Atomistic Molecular Dynamics (MD) simulations provide researchers the ability to model biomolecular structures such as proteins and their interactions with drug-like small molecules with greater spatiotemporal resolution than is otherwise possible using experimental methods. MD simulations are notoriously expensive computational endeavors that have traditionally required massive investment in specialized hardware to access biologically relevant spatiotemporal scales. Our goal is to summarize the fundamental algorithms that are employed in the literature to then highlight the challenges that have affected accelerator implementations in practice. We consider three broad categories of accelerators: Graphics Processing Units (GPUs), Field-Programmable Gate Arrays (FPGAs), and Application Specific Integrated Circuits (ASICs). These categories are comparatively studied to facilitate discussion of their relative trade-offs and to gain context for the current state of the art. We conclude by providing insights into the potential of emerging hardware platforms and algorithms for MD.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

MIDAS, prototype Multivariate Interactive Digital Analysis System, phase 1. Volume 1: System description

The MIDAS System is described as a third-generation fast multispectral recognition system able to keep pace with the large quantity and high rates of data acquisition from present and projected sensors. A principal objective of the MIDAS program is to provide a system well interfaced with the human operator and thus to obtain large overall reductions in turnaround time and significant gains in throughput. The hardware and software are described. The system contains a mini-computer to control the various high-speed processing elements in the data path, and a classifier which implements an all-digital prototype multivariate-Gaussian maximum likelihood decision algorithm operating at 200,000 pixels/sec. Sufficient hardware was developed to perform signature extraction from computer-compatible tapes, compute classifier coefficients, control the classifier operation, and diagnose operation.

Kriegler, F. J.↗

Microsecond-latency feedback at a particle accelerator by online reinforcement learning on hardware

The commissioning and operation of future large-scale scientific experiments will challenge current tuning and control methods. Reinforcement learning (RL) algorithms are a promising solution due to their ability to dynamically adapt to changing environments and consider delayed consequences. In many real-world applications, RL policies must produce actions in real time, often within microseconds to milliseconds, imposing significant constraints on system latency and computational overhead that conventional machine learning libraries are not designed to handle. To control phenomena in real time at these timescales, RL needs to be deployed on-the-edge, namely on dedicated hardware located near the system it controls, without relying on a host CPU or cloud-based inference. In this work we present the design and deployment of an experience accumulator system in a particle accelerator. In this system, deep-RL algorithms run using hardware acceleration and act within a few microseconds, enabling the use of RL for control of phenomena like beam instabilities. The training uses the collected data offline to reduce the number of operations carried out on the acceleration hardware. The proposed architecture was tested in real experimental conditions at the Karlsruhe research accelerator, a synchrotron light source, where the system was used to control artificially induced horizontal betatron oscillations in real-time, with a control loop period of just 2.7 μs. The results showed a performance comparable to the commercial feedback system available at the accelerator, demonstrating the viability and potential of this approach. Due to the self-learning and reconfiguration capability of this implementation, a seamless application to other control problems is possible. Applications range from particle accelerators to large-scale research and industrial facilities.

FPGA↗