Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “hardware algorithms”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Assessing and advancing the potential of quantum computing: A NASA case study

Quantum computing is one of the most enticing computational paradigms with the potential to revolutionize diverse areas of future-generation computational systems. While quantum computing hardware has advanced rapidly, from tiny laboratory experiments to quantum chips that can outperform even the largest supercomputers on specialized computational tasks, these noisy-intermediate scale quantum (NISQ) processors are still too small and non-robust to be directly useful for any real-world applications. In this paper, we describe NASA’s work in assessing and advancing the potential of quantum computing. We discuss advances in algorithms, both near- and longer-term, and the results of our explorations on current hardware as well as with simulations, including illustrating the benefits of algorithm-hardware co-design in the NISQ era. This work also includes physics-inspired classical algorithms that can be used at application scale today. We discuss innovative tools supporting the assessment and advancement of quantum computing and describe improved methods for simulating quantum systems of various types on high-performance computing systems that incorporate realistic error models. We provide an overview of recent methods for benchmarking, evaluating, and characterizing quantum hardware for error mitigation, as well as insights into fundamental quantum physics that can be harnessed for computational purposes.

Rieffel, Eleanor G.↗

On the practical usefulness of the Hardware Efficient Ansatz

Variational Quantum Algorithms (VQAs) and Quantum Machine Learning (QML) models train a parametrized quantum circuit to solve a given learning task. The success of these algorithms greatly hinges on appropriately choosing an ansatz for the quantum circuit. Perhaps one of the most famous ansatzes is the one-dimensional layered Hardware Efficient Ansatz (HEA), which seeks to minimize the effect of hardware noise by using native gates and connectives. The use of this HEA has generated a certain ambivalence arising from the fact that while it suffers from barren plateaus at long depths, it can also avoid them at shallow ones. In this work, we attempt to determine whether one should, or should not, use a HEA. We rigorously identify scenarios where shallow HEAs should likely be avoided (e.g., VQA or QML tasks with data satisfying a volume law of entanglement). More importantly, we identify a Goldilocks scenario where shallow HEAs could achieve a quantum speedup: QML tasks with data satisfying an area law of entanglement. We provide examples for such scenario (such as Gaussian diagonal ensemble random Hamiltonian discrimination), and we show that in these cases a shallow HEA is always trainable and that there exists an anti-concentration of loss function values. Our work highlights the crucial role that input states play in the trainability of a parametrized quantum circuit, a phenomenon that is verified in our numerics.

97 MATHEMATICS AND COMPUTING↗

Grid Application and Controls Development for Medium-Voltage SiC-Based Grid Interconnects

This paper studies, the grid-interconnection requirements of directly connected medium-voltage (MV) Silicon Carbide-based (SiC-based) converters and details the necessary control implementations. These power electronic converters follow the interconnection standard IEEE 1547-2018, but can also have additional controls supported by the microgrid standards in IEEE 2030.7, which is demonstrated in this work. Additionally, this paper defines the grid applications realized by these directly connected medium-voltage power electronic converters at the distribution scale and demonstrates results. The controls for the studied grid applications are then developed and implemented on the controller. The paper also includes details of the control algorithm development and validations of the control algorithm through hardware-in-the-loop results.

28 EE - Advanced Manufacturing Office (EE-5A)↗

Contextual subspace variational quantum eigensolver calculation of the dissociation curve of molecular nitrogen on a superconducting quantum computer

Abstract We present an experimental demonstration of the Contextual Subspace Variational Quantum Eigensolver on superconducting hardware. Calculating the potential energy curve of molecular nitrogen proves challenging for many conventional quantum chemistry techniques, since static correlation dominates in the dissociation limit. Our quantum simulations retain good agreement with the Full Configuration Interaction energy, outperforming all benchmarked single-reference wavefunction techniques in capturing the bond-breaking appropriately. Moreover, our methodology is competitive with multiconfigurational approaches but at a saving of quantum resource, meaning larger active spaces can be treated for a fixed qubit allowance. To achieve this result, we deploy an error mitigation/suppression strategy comprised of Dynamical Decoupling, Measurement-Error Mitigation and Zero-Noise Extrapolation. Circuit parallelization also provides passive noise-averaging and improves the effective shot yield to reduce the measurement overhead. Furthermore, we introduce a modified adaptive ansatz construction algorithm that incorporates hardware awareness into our variational circuits, minimizing the transpilation cost for the target qubit topology.

Physics↗

Variational approaches to constructing the many-body nuclear ground state for quantum computing

Here, we explore the preparation of specific nuclear states on gate-based quantum hardware using variational algorithms. Large-scale classical diagonalizations of the nuclear shell model have reached sizes of 10 9 –10 10 basis states but are still severely limited by computational resources. Quantum computing can, in principle, solve such systems exactly with exponentially fewer resources than classical computing. Exact solutions for large systems require many qubits and large gate depth, but variational approaches can effectively limit the required gate depth. We use the unitary coupled cluster approach to construct approximations of the ground-state vectors, later to be used in dynamics calculations. The testing ground is the phenomenological shell model space, which allows us to mimic the complexity of the internucleon interactions. We find that often one needs to minimize over a large number of parameters, using a large number of entanglements that makes the application on existing hardware challenging. Prospects for rapid improvements with more capable hardware are, however, very encouraging.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Probabilistic Neural Computing with Stochastic Devices

Abstract The brain has effectively proven a powerful inspiration for the development of computing architectures in which processing is tightly integrated with memory, communication is event‐driven, and analog computation can be performed at scale. These neuromorphic systems increasingly show an ability to improve the efficiency and speed of scientific computing and artificial intelligence applications. Herein, it is proposed that the brain's ubiquitous stochasticity represents an additional source of inspiration for expanding the reach of neuromorphic computing to probabilistic applications. To date, many efforts exploring probabilistic computing have focused primarily on one scale of the microelectronics stack, such as implementing probabilistic algorithms on deterministic hardware or developing probabilistic devices and circuits with the expectation that they will be leveraged by eventual probabilistic architectures. A co‐design vision is described by which large numbers of devices, such as magnetic tunnel junctions and tunnel diodes, can be operated in a stochastic regime and incorporated into a scalable neuromorphic architecture that can impact a number of probabilistic computing applications, such as Monte Carlo simulations and Bayesian neural networks. Finally, a framework is presented to categorize increasingly advanced hardware‐based probabilistic computing technologies.

Misra, Shashank↗

The case for data science in experimental chemistry: examples and recommendations

The physical sciences community is increasingly taking advantage of the possibilities offered by modern data science to solve problems in experimental chemistry and potentially to change the way we design, conduct and understand results from experiments. Successfully exploiting these opportunities involves considerable challenges. In this Expert Recommendation, we focus on experimental co-design and its importance to experimental chemistry. We provide examples of how data science is changing the way we conduct experiments, and we outline opportunities for further integration of data science and experimental chemistry to advance these fields. Our recommendations include establishing stronger links between chemists and data scientists; developing chemistry-specific data science methods; integrating algorithms, software and hardware to ‘co-design’ chemistry experiments from inception; and combining diverse and disparate data sources into a data network for chemistry research.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Performance Evaluation of a Novel Sequence-Based Directional Detection Strategy for Protection of Active Distribution Networks

Directional elements are relied on to achieve selectivity in fault detection in power systems. Although such elements have been deployed successfully for many years, there is an increased need for novel methods to deal with the unique challenges of directional protection in modern distribution networks. This article analyzes the impact of inverter-based resources (IBRs) on existing directional protection methods in distribution systems. It identifies parts of such elements that pose a risk of misoperation when IBRs are used in distribution networks. The authors have developed a new directional detection method for unbalanced faults in such networks using superimposed symmetrical sequence quantities. The phase angle of the superimposed negative sequence admittance is used to determine fault direction. The paper also presents a real-time co-simulation platform between a simulated distribution system and physical protection relay, using OPAL-RT. An SEL-411L relay is used to program the detection algorithm. This hardware-in-the-loop (HIL) setup is used to verify the performance of the method and the results are compared with existing directional methods

24 POWER TRANSMISSION AND DISTRIBUTION↗

Automated Generation of Integrated Digital and Spiking Neuromorphic Machine Learning Accelerators

The growing numbers of application areas for artificial intelligence (AI) methods have led to an explosion of domain-specific accelerators that could support every new machine learning (ML) algorithm advancement, clearly highlighting the need for a capability to quickly and automatically transition from algorithm definition to hardware implementation and explore design space along a variety of SWaP (size, weight and Power). The software defined architectures (SODA) synthesizer implements a compiler-based modular infrastructure for the end-to-end generation of machine learning accelerators from high-level frameworks to hardware description language. At the same time, neuromorphic computing, by mimicking how the brain operates, promises to perform artificial intelligence tasks at efficiencies orders of magnitude higher than the current conventional tensor-processing based accelerators, as demonstrated by a variety of specialized designs leveraging Spiking Neural Networks (SNNs). Nevertheless, the mapping of an artificial neural network (ANN) to solutions supporting SNNs is still a non-trivial and very device-specific task, and completely lack the possibility to design hybrid systems that integrate conventional and spiking neural models. In this paper we discuss the support for such an integrated generation leveraging the SODA Synthesizer framework and its modular structure. In particular, we present a new MLIR dialect (part of the SODA frontend) that allows expressing spiking neural network features (e.g., available resources, spiking sequences, analog signal reading, etc.) and illustrate how it enables mapping to Spiking Neurons and deployment to the related specialized hardware (which, in the digital domain, could be generated through the other existing layers of the SODA Synthesizer). We then discuss the opportunities for even deeper integration afforded by the hardware compilation infrastructure, providing a path towards the generation of complex heterogeneous artificial intelligence systems.

Curzel, Serena↗

BoBa

BoBa is a C++ software library for working with large matrices, tensors, and tensor decompositions. The library provides tools for dense matrix and tensor operations, tensor decompositions, and tensor decomposition methods that support modern CPU and GPU architectures. It includes portable abstractions for linear algebra, tensor algebra, and multidimensional computation. BoBa is intended for scientific computing applications that involve large multidimensional data sets or high dimensional mathematical models. Its capabilities support tasks such as data compression, linear algebra, efficient numerical computation, and the development of scalable algorithms for heterogeneous hardware. Tutorials, tests, and example applications are included to help users learn and apply the library.

Yao, Jin [Lawrence Livermore National Laboratory (↗

An HPC benchmark survey and taxonomy for characterization

The field of High-Performance Computing (HPC) is defined by providing computing devices with highest performance for a variety of demanding scientific users. The tight co-design relationship between HPC providers and users propels the field forward, paired with technological improvements, achieving continuously higher performance and resource utilization. A key device for system architects, architecture researchers, and scientific users are benchmarks, allowing for well-defined assessment of hardware, software, and algorithms. Many benchmarks exist in the community, from individual niche benchmarks testing specific features, to large-scale benchmark suites for whole procurements. We survey the available HPC benchmarks, summarizing them in table form with key details and concise categorization, also through an interactive website. For categorization, we present a benchmark taxonomy for well-defined characterization of benchmarks.

Benchmarking↗

VQE method: a short survey and recent developments

Abstract The variational quantum eigensolver (VQE) is a method that uses a hybrid quantum-classical computational approach to find eigenvalues of a Hamiltonian. VQE has been proposed as an alternative to fully quantum algorithms such as quantum phase estimation (QPE) because fully quantum algorithms require quantum hardware that will not be accessible in the near future. VQE has been successfully applied to solve the electronic Schrödinger equation for a variety of small molecules. However, the scalability of this method is limited by two factors: the complexity of the quantum circuits and the complexity of the classical optimization problem. Both of these factors are affected by the choice of the variational ansatz used to represent the trial wave function. Hence, the construction of an efficient ansatz is an active area of research. Put another way, modern quantum computers are not capable of executing deep quantum circuits produced by using currently available ansatzes for problems that map onto more than several qubits. In this review, we present recent developments in the field of designing efficient ansatzes that fall into two categories—chemistry–inspired and hardware–efficient—that produce quantum circuits that are easier to run on modern hardware. We discuss the shortfalls of ansatzes originally formulated for VQE simulations, how they are addressed in more sophisticated methods, and the potential ways for further improvements.

Fedorov, Dmitry A. (ORCID:0000000316598580)↗

Advances in Mixed Precision Algorithms: 2021 Edition

Over the last year, the ECP xSDK-multiprecision effort has made tremendous progress in developing and deploying new mixed precision technology and customizing the algorithms for the hardware deployed in the ECP flagship supercomputers. The effort also has succeeded in creating a cross-laboratory community of scientists interested in mixed precision technology and now working together in deploying this technology for ECP applications. In this report, we highlight some of the most promising and impactful achievements of the last year. Among the highlights we present are: Mixed precision IR using a dense LU factorization and achieving a 1.8× speedup on Spock; results and strategies for mixed precision IR using a sparse LU factorization; a mixed precision eigenvalue solver; Mixed Precision GMRES-IR being deployed in Trilinos, and achieving a speedup of 1.4× over standard GMRES; compressed Basis (CB) GMRES being deployed in Ginkgo and achieving an average 1.4× speedup over standard GMRES; preparing hypre for mixed precision execution; mixed precision sparse approximate inverse preconditioners achieving an average speedup of 1.2×; and detailed description of the memory accessor separating the arithmetic precision from the memory precision, and enabling memory-bound low precision BLAS 1/2 operations to increase the accuracy by using high precision in the computations without degrading the performance. We emphasize that many of the highlights presented here have also been submitted to peer-reviewed journals or established conferences, and are under peer-review or have already been published.

97 MATHEMATICS AND COMPUTING↗

Advances in Mixed Precision Algorithms: 2021 Edition

Over the last year, the ECP xSDK-multiprecision effort has made tremendous progress in developing and deploying new mixed precision technology and customizing the algorithms for the hardware deployed in the ECP flagship supercomputers. The effort also has succeeded in creating a cross-laboratory community of scientists interested in mixed precision technology and now working together in deploying this technology for ECP applications. In this report, we highlight some of the most promising and impactful achievements of the last year. Among the highlights we present are • Mixed precision IR using a dense LU factorization and achieving a 1.8× speedup on Spock; • Results and strategies for mixed precision IR using a sparse LU factorization; • A mixed precision eigenvalue solver; • Mixed Precision GMRES-IR being deployed in Trilinos, and achieving a speedup of 1.4× over standard GMRES; • Compressed Basis (CB) GMRES being deployed in Ginkgo and achieving an average 1.4× speedup over standard GMRES; • Preparing hypre for mixed precision execution; • Mixed precision sparse approximate inverse preconditioners achieving an average speedup of 1.2×; • Detailed description of the memory accessor separating the arithmetic precision from the memory precision, and enabling memory-bound low precision BLAS 1/2 operations to increase the accuracy by using high precision in the computations without degrading the performance. We emphasize that many of the highlights presented here have also been submitted to peer-reviewed journals or established conferences, and are under peer-review or have already been published.

97 MATHEMATICS AND COMPUTING↗

Extended Low Load Boiler Operation to Improve Performance and Economics of an Existing Coal Fired Power Plant (Final Report)

The overall goal is to improve the performance and economics of existing coal fired power plants by extending low load boiler operation to lower loads than is currently achievable. The objective of this program is to develop and validate sensor hardware and analytical algorithms to lower plant operating expenses (OPEX) for the currently operating pulverized coal utility boiler fleet. Coal fired utility boilers are increasingly under grid dispatch pressure. In some cases, the coal fired cost of generation is noncompetitive with respect to natural gas generation and subsidized renewable sources. To remain profitable and remain fully compliant with existing environmental regulations, the installed coal fired fleet must find technologies which allow it to move into a more flexible cyclic load dispatch model. Today the installed coal fired utility fleet must be cost of generation competitive, fully emissions compliant, and responsive to the variability inherent in renewable energy generation sources. In the Phase I of the project, GE Steam Power, Inc. (GE) performed modeling of different operating scenarios for low load operation using an existing full plant dynamic model developed for a 660MW steam power plant. Sensors and analytic algorithms to enable a stable and steady coal supply for low load pulverizer operation were identified and tested at the Pulverizer Development Facility (PDF) at GE’s Clean Energy Center in Bloomfield, Connecticut. Sensors and analytic algorithms to enable stable combustion for low load operation were identified and tested at the 15 MWth Industrial Scale Burner facility (ISBF) at GE’s Clean Energy Center. A concept was developed to test the sensors and control algorithms, down selected after testing, at a full-scale coal fired power plant. A budget estimate was then developed, and the concept was implemented at an existing utility power plant. The specific objectives of the experimental work were to: • Identify and select sensors and analytic algorithms for monitoring coal pulverizer operation at lower loads to provide stable operation and appropriate coal fineness at lower coal throughput; Identify and select sensors and analytic algorithms for a Boiler Flame Stability Monitor to better balance air and fuel at each burner. This enables a reduction in a coal boiler’s safe low load power level while maintaining stable flame characteristics; Develop a concept in Phase I for low load operation of a full-scale power plant and develop a budget estimate for testing and execute the test plan at an existing plant in Phase II; Validate the capability of the extended low load boiler system to extend the minimum load operating point in a safe and reliable manner on an existing full-scale utility boiler. At the completion of this experimental study, GE has developed a set of sensors and analytic algorithms, down selected after testing, that have the potential to enable safe low load operation of a utility boiler. GE has also identified a host site for testing these identified sensors and analytic algorithms. GE has generated a full set of deliverables that provide sufficient information to proceed with the next step of testing at a host site. This includes a potential host site and budget estimate for concept testing at host site. In the Phase II of the project, a series of field tests were completed to validate the extended low load boiler operation, which consisted of detailed engineering, installation, commissioning, and testing the additional sensors and analytics for the coal-fired combustion system on an existing full-scale utility boiler. The optimization work has been supported by the host plant and endorsed by their engineering and operation staff.

01 COAL, LIGNITE, AND PEAT↗

Integration of a real-time orientation measurement system for a real-time evaluator (RTE) to measure the position and orientation of crane-lifted components

Prefabrication of building components holds the potential to revolutionize the construction industry. Prefabrication consists of manufacturing building components, modules, and other elements in a factory to be shipped and installed on a construction site. Prefabricated components have been produced for various applications including precast concrete panels for new construction and exterior wall retrofits. The manufacturing process has seen much innovation in recent years; however, the installation process has seen minimal advancements. A real-time evaluator (RTE) was developed to reduce the installation cost of prefabricated components by reducing installation time, decreasing rework, and improving accuracy. The RTE uses off-the-shelf hardware and novel algorithms to assist erectors with component installation by measuring the real-time positions of connections and prefabricated components, providing installation guidance through a graphical user interface, and monitoring the accumulated installation errors. An overview of the RTE and the proposed workflow is presented. Previous on-site demonstrations provided valuable feedback from users on the potential areas for improvement of the system. One common request was real-time measurement of component orientation during lifting, a process that previously required that the component remain stationary while the laser tracker cycled through target prisms. This paper will present the incorporation and testing of a real-time orientation measurement system as it was implemented into the RTE, allowing for measurement of component orientation during movement.

Selvakumar, Balaji [ORNL]↗

Rasterization with Data-Parallel Primitives

Parallel rasterization can suffer from race conditions during fragment generation, which is traditionally addressed by using specialized hardware accessible via vendor graphics APIs. Unfortunately, graphics APIs are increasingly problematic on high-performance computers, either because they are not provided or because of concerns about dependencies with in situ visualization. In response, we present a hardware-agnostic rasterization algorithm that handles race conditions using only data-parallel primitives (DPPs), enabling efficient rendering on HPC systems without graphics API dependencies and aligning with recent efforts to deliver visualization software with DPPs. Our evaluation consists of three phases: (1) evaluating portability across different CPU and GPU architectures, (2) evaluating competitiveness with a community standard, and (3) evaluating performance across varying workloads and available parallelism. The supporting experiments run on both AMD and NVIDIA GPUs, considering data sets as large as 460 million triangles and 160 million pixels. While performance generally falls short of graphics API baselines, it achieves interactive frame rates on most workloads. As a result, we conclude our approach is a viable solution for rasterization on high-performance computers since our approach is portably performant across different architectures without the need for specialized vendor support.

Buckley, Makani [University of Oregon] (ORCID:0009↗

Automated tracking of prefabricated components for areal-time evaluator to optimize and automateinstallation

Trade associations for prefabricated construction estimate that about 50% of prefabricated wall projects have alignment problems that lead to defects and rework. Additionally, component installation times average between 30 and 60 minutes per component. To address these issues, a real-time evaluator (RTE) system was introduced to decrease cost and automate prefabricated component installation by reducing the installation time, decreasing rework, and enhancing energy performance through higher installation quality. The RTE uses commonly available hardware and software to perform autonomous tracking to measure the real-time location and orientation of components as they are crane-lifted and installed. The hardware, software, and algorithms that allow the autonomous tracking of components are detailed. An algorithm to automate the initial search for a component with three attached retroreflectors is proposed. Algorithms to automate the measurement of component position and orientation are also proposed. Simple lab-scale proof-of-concept experiments were conducted to assess the algorithms for automation of component searching, measurement of real-time movement, and measurement of component orientation. With additional development, the system can be used as a tool to generate the commands for autonomous crane operation or single-task construction robots.

Hayes, Nolan↗