Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “hardware algorithms”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

An Overview of Production and Validation of the SMAP Passive Soil Moisture Product

The Soil Moisture Active Passive (SMAP) mission is an L-band mission scheduled for launch in Jan. 2015. The SMAP instruments consist of a radar and a radiometer to obtain complementary information from space for soil moisture and freeze/thaw state research and applications. By utilizing novel designs in antenna construction, retrieval algorithms, and acquisition hardware, SMAP provides a capability for global mapping of soil moisture and freeze/thaw state with unprecedented accuracy, resolution, and coverage. This improvement in hydrosphere state measurement is expected to advance our understanding of the processes that link the terrestrial water, energy and carbon cycles, improve our capability in flood prediction and drought monitoring, and enhance our skills in weather and climate forecast. For swath-based soil moisture measurement, SMAP generates three operational geophysical data products: (1) the radiometer-only soil moisture product (L2_SM_P) posted at 36-kilometer resolution, (2) the radar-only soil moisture product (L2_SM_A) posted at 3-kilometers resolution, and (3) the radar-radiometer combined soil moisture product (L2_SM_AP) posted at 9-kilometers resolution. Each product draws on the strengths of the underlying sensor(s) and plays a unique role in hydroclimatological and hydrometeorological applications. A full suite of SMAP data products is given in Table 1.

passive microwave↗

Dynamical Decoupling for Measuring and Suppressing Crosstalk

Dynamical decoupling (DD) is a noise-mitigating strategy in which sequences of pulses are applied to single qubits to average out their interaction with the environment. DD has been extensively studied and demonstrated for suppressing single-qubit decoherence and can be tailored for different noise spectrum. We report another important adaptation of DD where crosstalk between qubits are suppressed. We demonstrate the efficiency of this procedure on quantum circuits on superconducting transmon-based quantum devices. We designed a family of syncopated DD sequences that effectively suppress ZZ coupling between qubit pairs, which is the dominating crosstalk form on the device. We insert DD to a quantum circuit whenever single qubits are idle (often during two-qubits gates on other qubits). While standard periodic DD suppress crosstalk between these qubits and their neighbors, the syncopated DD further decouples crosstalk between these qubits. We further designed short sequences that maximize the application of DD without adding time to the quantum circuit execution. Such DD sequences yield significant improvement of the performance of the algorithm on the hardware. The performance is further boosted by combining DD with another mitigation strategy, randomized compilation. Our work demonstrated that syncopated DD is effective and practical way to suppress crosstalk in quantum circuits and serves as a great probe to characterize the crosstalk and inform hardware design.

Quantum Computing↗

Automated tracking of prefabricated components for areal-time evaluator to optimize and automateinstallation

Trade associations for prefabricated construction estimate that about 50% of prefabricated wall projects have alignment problems that lead to defects and rework. Additionally, component installation times average between 30 and 60 minutes per component. To address these issues, a real-time evaluator (RTE) system was introduced to decrease cost and automate prefabricated component installation by reducing the installation time, decreasing rework, and enhancing energy performance through higher installation quality. The RTE uses commonly available hardware and software to perform autonomous tracking to measure the real-time location and orientation of components as they are crane-lifted and installed. The hardware, software, and algorithms that allow the autonomous tracking of components are detailed. An algorithm to automate the initial search for a component with three attached retroreflectors is proposed. Algorithms to automate the measurement of component position and orientation are also proposed. Simple lab-scale proof-of-concept experiments were conducted to assess the algorithms for automation of component searching, measurement of real-time movement, and measurement of component orientation. With additional development, the system can be used as a tool to generate the commands for autonomous crane operation or single-task construction robots.

Hayes, Nolan↗

A Roadmap for Reaching the Potential of Brain-Derived Computing

Neuromorphic computing is a critical future technology for the computing industry, but it has yet to achieve its promise and has struggled to establish a cohesive research community. A large part of the challenge is that full realization of the potential of brain inspiration requires advances in both device hardware, computing architectures, and algorithms. This simultaneous development across technology scales is unprecedented in the computing field. This article presents a strategy, framed by market and policy pressures, for moving past these current technological and cultural hurdles to realize its full impact across technology. Achieving the full potential of brain-derived algorithms as well as post-complementary metal-oxide-semiconductor (CMOS) scaling neuromorphic hardware requires appropriately balancing the near-term opportunities of deep learning applications with the long-term potential of less understood opportunities in neural computing.

97 MATHEMATICS AND COMPUTING↗

Using Grover's search algorithm to characterize the Rigetti Quantum Computing Platform

Performance of the Rigetti quantum computing platform was tested during the period of time from 12/16/2019 to 05/18/2020. The Grover's search (GS) algorithm was used as a testing tool, in particular, 3- and 4-level versions of the algorithm. As a result, a number of hardware issues were revealed, so the algorithm was split to smaller blocks and individual gates, and all of them were tested separately. The fidelity decay was found to be due to both decoherent processes and coherent errors of the native hardware gates. These errors were estimated from RX gate benchmarks and were in a form of extra rotation of the quantum state. As a consequence of that, performance of the algorithm was shown to be strongly dependent on the native gate decomposition of the program. Suggestions and possible improvements of the future Rigetti runs are made based on the obtained observations.

97 MATHEMATICS AND COMPUTING↗

Trellises and Trellis-Based Decoding Algorithms for Linear Block Codes: A Recursive Maximum Likelihood Decoding - Part 3

The Viterbi algorithm is indeed a very simple and efficient method of implementing the maximum likelihood decoding. However, if we take advantage of the structural properties in a trellis section, other efficient trellis-based decoding algorithms can be devised. Recently, an efficient trellis-based recursive maximum likelihood decoding (RMLD) algorithm for linear block codes has been proposed. This algorithm is more efficient than the conventional Viterbi algorithm in both computation and hardware requirements. Most importantly, the implementation of this algorithm does not require the construction of the entire code trellis, only some special one-section trellises of relatively small state and branch complexities are needed for constructing path (or branch) metric tables recursively. At the end, there is only one table which contains only the most likely code-word and its metric for a given received sequence r = (r(sub 1), r(sub 2),...,r(sub n)). This algorithm basically uses the divide and conquer strategy. Furthermore, it allows parallel/pipeline processing of received sequences to speed up decoding.

Lin, Shu↗

Real-Time Adaptive Lossless Hyperspectral Image Compression using CCSDS on Parallel GPGPU and Multicore Processor Systems

The proposed CCSDS (Consultative Committee for Space Data Systems) Lossless Hyperspectral Image Compression Algorithm was designed to facilitate a fast hardware implementation. This paper analyses that algorithm with regard to available parallelism and describes fast parallel implementations in software for GPGPU and Multicore CPU architectures. We show that careful software implementation, using hardware acceleration in the form of GPGPUs or even just multicore processors, can exceed the performance of existing hardware and software implementations by up to 11x and break the real-time barrier for the first time for a typical test application.

realtime↗

ARQUIN: Architectures for Multinode Superconducting Quantum Computers

Many proposals to scale quantum technology rely on modular or distributed designs wherein individual quantum processors, called nodes, are linked together to form one large multinode quantum computer (MNQC). One scalable method to construct an MNQC is using superconducting quantum systems with optical interconnects. However, internode gates in these systems may be two to three orders of magnitude noisier and slower than local operations. Surmounting the limitations of internode gates will require improvements in entanglement generation, use of entanglement distillation, and optimized software and compilers. Still, it remains unclear what performance is possible with current hardware and what performance algorithms require. In this article, we employ a systems analysis approach to quantify overall MNQC performance in terms of hardware models of internode links, entanglement distillation, and local architecture. We show how to navigate tradeoffs in entanglement generation and distillation in the context of algorithm performance, lay out how compilers and software should balance between local and internode gates, and discuss when noisy quantum internode links have an advantage over purely classical links. Here, we find that a factor of 10–100× better link performance is required and introduce a research roadmap for the co-design of hardware and software towards the realization of early MNQCs. While we focus on superconducting devices with optical interconnects, our approach is general across MNQC implementations.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Computation of Earth Science Products on Spaceborne Platforms

Spaceborne sensors like NASA's Hyperion hyperspectral imager generate huge data volumes, and several near-term trends indicate that data volumes will only increase. Next-generation hyperspectral missions, such as NASA's Hyperspectral Infrared Imager (HyspIRI), will operate at higher duty cycles and higher data rates, and their users will expect products to be generated from the data in near real time [1]. Barring a sudden advance in satellite downlink capacity, these trends point to a need to process data and generate products onboard the spacecraft. Rather than downlink an entire hyperspectral image cube, onboard processing enables satellites to downlink partial or completed scientific data products, which are often one to two orders of magnitude smaller than the original image. In addition, a satellite with onboard data processing resources and direct broadcast transmission equipment could send data products directly to first responders, research scientists or other users on the ground. Next-generation space-capable data processors will have a combination of reconfigurable gate arrays, digital signal processors and general-purpose CPUs. Correctly programmed and configured, these resources are sufficient to run sophisticated data analysis programs, including hyperspectral image processing algorithms that commonly run on desktop computers [2]. This paper describes how we implemented one such program, the HSEG hierarchical image segmentation algorithm, software commonly used on desktop and parallel processors, on a hardware platform designed to mimic a next-generation space-capable data processor [3]. We also describe our approach to porting the algorithm to and optimizing it for the new platform, and determine the expected performance gains enabled by our design. This extended abstract will describe the HSEG algorithm and hardware platform in greater detail, provide an analysis of the key function within the algorithm that required hardware acceleration, and describe our implementation of that function in hardware.

Fisher, Kevin↗

Development of a Hardware-in-The-Loop Testbed for a Decentralized, Data-Driven Electric Vehicle Charging Control Algorithm

This study presents the design of an electric vehicle (EV)-grid integration (EVGI) hardware test-bed to implement smart EV charging algorithms. Here, the proposed test-bed also allows to create different grid events via flexible integration of other power hardware (e.g., controllable loads and battery energy storage systems) and test their impacts on EV charging. The design uses a real-time digital simulator to realize a complex distribution grid model with primary and secondary networks. A grid simulator physically realizes the selected nodes of the simulated grid to power an actual EV, forming a hardware-in-the-loop (HIL) test setup. The EV-grid integration is demonstrated based on the custom hardware and software implementation of the J1772 charging protocol using dSPACE MicroLabBox, operating as a custom EV Supply Equipment (EVSE). The HIL test-bed features a novel testing platform for accurate implementation and analysis of scalable charging algorithms. To this end, a data-driven, decentralized, model-free charging controller based on the Additive Increase and Multiplicative Decrease (AIMD) algorithm is presented and validated on an EV using the HIL test-bed. We tested the proposed algorithm under various case studies, and presented a comparison study with an existing droop-based, decentralized charging solution. The results showed that the EV successfully performed charging commands generated by the EVSE and regulated its charging power to effectively reduce the system loading caused by high EV penetration.

33 ADVANCED PROPULSION SYSTEMS↗

Efficient and Optimal Attitude Determination Using Recursive Global Positioning System Signal Operations

In this paper, a new and efficient algorithm is developed for attitude determination from Global Positioning System signals. The new algorithm is derived from a generalized nonlinear predictive filter for nonlinear systems. This uses a one time-step ahead approach to propagate a simple kinematics model for attitude determination. The advantages of the new algorithm over previously developed methods include: it provides optimal attitudes even for coplanar baseline configurations; it guarantees convergence even for poor initial conditions; it is a non-iterative algorithm; and it is computationally efficient. These advantages clearly make the new algorithm well suited to on-board applications. The performance of the new algorithm is tested on a dynamic hardware simulator. Results indicate that the new algorithm accurately estimates the attitude of a moving vehicle, and provides robust attitude estimates even when other methods, such as a linearized least-squares approach, fail due to poor initial starting conditions.

Crassidis, John L.↗

Scaling quantum approximate optimization on near-term hardware

The quantum approximate optimization algorithm (QAOA) is an approach for near-term quantum computers to potentially demonstrate computational advantage in solving combinatorial optimization problems. However, the viability of the QAOA depends on how its performance and resource requirements scale with problem size and complexity for realistic hardware implementations. Here, we quantify scaling of the expected resource requirements by synthesizing optimized circuits for hardware architectures with varying levels of connectivity. Assuming noisy gate operations, we estimate the number of measurements needed to sample the output of the idealized QAOA circuit with high probability. We show the number of measurements, and hence total time to solution, grows exponentially in problem size and problem graph degree as well as depth of the QAOA ansatz, gate infidelities, and inverse hardware graph degree. These problems may be alleviated by increasing hardware connectivity or by recently proposed modifications to the QAOA that achieve higher performance with fewer circuit layers.

97 MATHEMATICS AND COMPUTING↗

Preliminary results from a subsonic high-angle-of-attack flush airdata sensing (HI-FADS) system - Design, calibration, algorithm development, and flight test evaluation

A nonintrusive high angle-of-attack flush airdata sensing (HI-FADS) system was installed and flight-tested on the F-18 high alpha research vehicle. This paper discusses the airdata algorithm development and composite results expressed as airdata parameter estimates and describes the HI-FADS system hardware, calibration techniques, and algorithm development. An independent empirical verification was performed over a large portion of the subsonic flight envelope. Test points were obtained for Mach numbers from 0.15 to 0.94 and angles of attack from -8.0 to 55.0 deg. Angles of sideslip ranged from -15.0 to 15.0 deg, and test altitudes ranged from 18,000 to 40,000 ft. The HI-FADS system gave excellent results over the entire subsonic Mach number range up to 55 deg angle of attack. The internal pneumatic frequency response of the system is accurate to beyond 10 Hz.

Whitmore, Stephen A.↗

Efficient Implementation for Unitary Coupled Cluster State Preparation for Near-Term Quantum Computers

Unitary coupled cluster theory (UCC) is a common wave function ansatz for quantum simulation of molecular electronic structure using the variational quantum eigenvalue solver (VQE). Even for small molecules using a double-ζ basis, the number of variational parameters required to minimize the electronic energy (i.e., optimize the circuit) is large and beyond the reach of current quantum computers. For example, a circuit simulating C2 using the UCCSD ansatz and the cc-pVDZ basis set with frozen-core will require over 10,000 variational parameters and a Hilbert space of over 10^8 determinants. To make progress on simulating such molecular systems on near-term quantum computers, we explore how much of the optimization can be approximately prepared with classical simulation while reducing the number of optimization steps performed on a quantum device. Recently, Chen, Cheng, and Freericks [J. Chem. Theory Comput. 2021, 17, 841-847] presented an algorithm for the factorized form of the UCC ansatz that allows for efficient UCC optimizations on classical hardware. We flip the algorithm around and use it to prepare approximate quantum circuits for systems that require a large number of qubits to represent. We will present results from our implementation and discuss strategies for incorporating this implementation for algorithms involving near-term quantum computers.

J Wayne Mullinax↗

Efficient Implementation for Unitary Coupled Cluster State Preparation for Near-Term Quantum Computers

Unitary coupled cluster theory (UCC) is a common wave function ansatz for quantum simulation of molecular electronic structure using the variational quantum eigenvalue solver (VQE). Even for small molecules using a double-ζ basis, the number of variational parameters required to minimize the electronic energy (i.e., optimize the circuit) is large and beyond the reach of current quantum computers. For example, a circuit simulating C2 using the UCCSD ansatz and the cc-pVDZ basis set with frozen-core will require over 10,000 variational parameters and a Hilbert space of over 10^(8) determinants. To make progress on simulating such molecular systems on near-term quantum computers, we explore how much of the optimization can be approximately prepared with classical simulation while reducing the number of optimization steps performed on a quantum device. Recently, Chen, Cheng, and Freericks [J. Chem. Theory Comput. 2021, 17, 841-847] presented an algorithm for the factorized form of the UCC ansatz that allows for efficient UCC optimizations on classical hardware. We flip the algorithm around and use it to prepare approximate quantum circuits for systems that require a large number of qubits to represent. We will present results from our implementation and discuss strategies for incorporating this implementation for algorithms involving near-term quantum computers.

Quantum Computing↗

Efficient Implementation for Unitary Coupled Cluster State Preparation for Near-Term Quantum Computers

Unitary coupled cluster theory (UCC) is a common wave function ansatz for quantum simulation of molecular electronic structure using the variational quantum eigenvalue solver (VQE). Even for small molecules using a double-ζ basis, the number of variational parameters required to minimize the electronic energy (i.e., optimize the circuit) is large and beyond the reach of current quantum computers. For example, a circuit simulating C2 using the UCCSD ansatz and the cc-pVDZ basis set with frozen-core will require over 10,000 variational parameters and a Hilbert space of over 10^(8) determinants. To make progress on simulating such molecular systems on near-term quantum computers, we explore how much of the optimization can be approximately prepared with classical simulation while reducing the number of optimization steps performed on a quantum device. Recently, Chen, Cheng, and Freericks [J. Chem. Theory Comput. 2021, 17, 841-847] presented an algorithm for the factorized form of the UCC ansatz that allows for efficient UCC optimizations on classical hardware. We flip the algorithm around and use it to prepare approximate quantum circuits for systems that require a large number of qubits to represent. We will present results from our implementation and discuss strategies for incorporating this implementation for algorithms involving near-term quantum computers.

Quantum Computing↗

Reconfigurable neuromorphic components and algorithms for next-generation artificial intelligence

Digital transistor-based general-purpose hardware (e.g., central processing units) is the dominant solution to support both traditional computing (logic, arithmetic, etc.) as well as modern artificial intelligence. State-of-the-art research has shown feasibility of post-digital physics-based neuromorphic hardware, which is hypothesized to support artificial intelligence algorithms with orders-of-magnitude improved time/energy efficiencies. But such research has not been widely deployed mainly because of such novel hardware’s extreme application-specificity, and the dominance of low-cost general-purpose (but inefficient) digital hardware. To make use of the novel algorithms and the superlative performance of physics-based hardware, we need to identify scientific principles that can enable generality in physics-based hardware. This work resulted in two important broad outcomes – first, we demonstrate fully reconfigurable neuromorphic components, and second, we demonstrate a viable artificial intelligence learning algorithm that can exploit the functioning of neuromorphic hardware. We demonstrate up to five orders of magnitude improvement in energy efficiency compared to the best general-purpose digital hardware.

97 MATHEMATICS AND COMPUTING↗

Framework for a space shuttle main engine health monitoring system

A framework developed for a health management system (HMS) which is directed at improving the safety of operation of the Space Shuttle Main Engine (SSME) is summarized. An emphasis was placed on near term technology through requirements to use existing SSME instrumentation and to demonstrate the HMS during SSME ground tests within five years. The HMS framework was developed through an analysis of SSME failure modes, fault detection algorithms, sensor technologies, and hardware architectures. A key feature of the HMS framework design is that a clear path from the ground test system to a flight HMS was maintained. Fault detection techniques based on time series, nonlinear regression, and clustering algorithms were developed and demonstrated on data from SSME ground test failures. The fault detection algorithms exhibited 100 percent detection of faults, had an extremely low false alarm rate, and were robust to sensor loss. These algorithms were incorporated into a hierarchical decision making strategy for overall assessment of SSME health. A preliminary design for a hardware architecture capable of supporting real time operation of the HMS functions was developed. Utilizing modular, commercial off-the-shelf components produced a reliable low cost design with the flexibility to incorporate advances in algorithm and sensor technology as they become available.

Hawman, Michael W.↗