Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “processors”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Early Performance Results on 4th Gen Intel(R) Xeon (R) Scalable Processors with DDR and Intel(R) Xeon(R) processors, codenamed Sapphire Rapids with HBM

The Crossroads supercomputer was designed to simulate some of the most complex physical devices in the world. These simulations routinely require 1/2 petabyte or more of system memory running on thousands of compute nodes for months at a time on the most powerful supercomputers. Improvements in time to solutions for these workloads can have major impact on our mission capabilities. In this paper we present early results of representative application workloads on 4th Gen Intel Xeon and Intel Xeon Processors codenamed Sapphire Rapids with HBM. These results demonstrate an extremely promising 8.57x improvement (node to node) over our prior generation Intel Broadwell (BDW) based HPC systems. No code modifications were required to achieve this speedup, providing a compelling path forward toward major reductions in time to solution and the complexity of physical systems that can be simulated in the future.

97 MATHEMATICS AND COMPUTING↗

Fast truncated SVD of sparse and dense matrices on graphics processors

We investigate the solution of low-rank matrix approximation problems using the truncated singular value decomposition (SVD). For this purpose, we develop and optimize graphics processing unit (GPU) implementations for the randomized SVD and a blocked variant of the Lanczos approach. Our work takes advantage of the fact that the two methods are composed of very similar linear algebra building blocks, which can be assembled using numerical kernels from existing high-performance linear algebra libraries. Furthermore, the experiments with several sparse matrices arising in representative real-world applications and synthetic dense test matrices reveal a performance advantage of the block Lanczos algorithm when targeting the same approximation accuracy.

Computer Science↗

Statistical analysis on random quantum circuit sampling by Sycamore and Zuchongzhi quantum processors

Random quantum circuit sampling, a task to sample bit strings from a random quantum circuit, is considered a suitable benchmark task to demonstrate the outperformance of quantum computers even with noisy qubits. Recently, random quantum circuit sampling was performed on the Sycamore quantum processor with 53 qubits [Nature (London) 574, 505 (2019)] and on the Zuchongzhi quantum processor with 56 qubits [Phys. Rev. Lett. 127, 180501 (2021)]. Here, we analyze and compare the statistical properties of the outputs of the random quantum circuit sampling by the Sycamore and Zuchongzhi processors. Using the Marchenko-Pastur law of random matrices of bit strings and the Wasssertein distances between bit strings, we find that the statistical properties of Sycamore bit strings are quite different from those of Zuchongzhi bit strings, while both processors score similar values of linear cross-entropy fidelity for random circuit sampling. Some bit strings sampled by the Zuchongzhi processor pass the NIST random number tests while both Sycamore and Zuchongzhi processors show similar patterns in the heat maps of bit strings. Zuchongzhi bit strings are much closer to classical uniform random bits than those of Sycamore. It is shown that the statistical properties of bit strings of both random quantum circuits change little as the depth of the random quantum circuits increases. Our findings raise a question about the computational reliability of noisy quantum processors because two quantum processors with similar noise levels and similar qubit structures produced statistically different outputs for the same random quantum circuit sampling.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Implementing and Benchmarking the Locally Competitive Algorithm on the Loihi 2 Neuromorphic Processor

Neuromorphic processors have garnered considerable interest in recent years for their potential in enabling energy-efficient and high-speed computing. The Locally Competative Algorithm (LCA) has been utilized for power efficient sparse coding on neuromophic processors, including the first Loihi processor \cite{appletospikes, loihi1}. With the Loihi 2 processor enabling custom neuron models and graded spike communication, more complex implementations of LCA are possible \cite{loihi2}. We present a new implementation of LCA designed for the Loihi 2 processor and perform an initial set of benchmarks comparing it to LCA on CPU and GPU devices. In these experiments LCA on Loihi 2 is faster and orders of magnitude more efficient, while maintaining similar reconstruction quality. We find this performance improvement increases as the LCA parameters are tuned towards greater representation sparsity. Our study highlights the potential of neuromorphic processors, particularly Loihi 2, in enabling intelligent,autonomous, real-time processing on small robots, satellite where there are strict SWaP (small, lightweighr, and low-power) requirement. By demonstrating the superior performance of LCA on Loihi 2 compared to conventional computing device, our study suggests that Loihi 2 could be a valuable tool in advancing these types of applications. Overall, our study highlights the potential of neuromorphic processors for efficient and accurate data processing on resource-constrained devices.

Parpart, Gavin G.↗

Locality-aware and sharing-aware cache coherence for collections of processors

A cache coherence technique for operating a multi-processor system including shared memory includes allocating a cache line of a cache memory of a processor to a memory address in the shared memory in response to execution of an instruction of a program executing on the processor. The technique includes encoding a shared information state of the cache line to indicate whether the memory address is a shared memory address shared by the processor and a second processor, or a private memory address private to the processor, in response to whether the instruction is included in a critical section of the program, the critical section being a portion of the program that confines access to shared, writeable data.

Farmahini Farahani, Amin↗

Techniques for recovering from errors when executing software applications on parallel processors

In various embodiments, a software program uses hardware features of a parallel processor to checkpoint a context associated with an execution of a software application on the parallel processor. The software program uses a preemption feature of the parallel processor to cause the parallel processor to stop executing instructions in accordance with the context. The software program then causes the parallel processor to collect state data associated with the context. After generating a checkpoint based on the state data, the software program causes the parallel processor to resume executing instructions in accordance with the context.

Hukerikar, Saurabh↗

Efficient flexible characterization of quantum processors with nested error models

We present a simple and powerful technique for finding a good error model for a quantum processor. The technique iteratively tests a nested sequence of models against data obtained from the processor, and keeps track of the best-fit model and its wildcard error (a metric of the amount of unmodeled error) at each step. Each best-fit model, along with a quantification of its unmodeled error, constitutes a characterization of the processor. We explain how quantum processor models can be compared with experimental data and to each other. We demonstrate the technique by using it to characterize a simulated noisy two-qubit processor.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Mitigating cosmic-ray-like correlated events with a modular quantum processor

Quantum processors based on superconducting qubits are being scaled to larger qubit numbers, enabling the implementation of small-scale quantum error-correction codes. However, catastrophic chip-scale correlated errors have been observed in these processors, attributed to, e.g., cosmic ray impacts, which challenge conventional error-correction codes such as the surface code. These events are characterized by a temporary but pronounced suppression of the qubit-energy relaxation times. Here, in this study, we explore the potential for modular quantum computing architectures to mitigate such correlated energy decay events. We measure cosmic-ray-like events in a quantum processor comprising a motherboard and two flip-chip bonded daughterboard modules, each module containing two superconducting qubits. We monitor the appearance of correlated qubit decay events within a single module and across the physically separated modules. We find that while decay events within one module are strongly correlated (over 85%), events in separate modules only display approximately 2% correlations. We also report coincident decay events in the motherboard and in either of the two daughterboard modules, providing further insight into the nature of these decay events. These results suggest that modular architectures, combined with bespoke errorcorrection codes, offer a promising approach for protecting future quantum processors from chip-scale correlated errors.

Wu, Xuntao [Univ. of Chicago, IL (United States)] ↗

Efficient frequency allocation for superconducting quantum processors using improved optimization techniques

Building on previous research on frequency allocation optimization for superconducting circuit quantum processors, this work incorporates several techniques to improve overall solution quality. Here, we introduce constraints and imposed edgewise differences help to improve the optimization results. We also introduce optimization variables for the orientation of each edge, defined as the direction from the control qubit to the target qubit, to be chosen during optimization. To scale up to larger processors, multimodule designs are employed with various boundary conditions, thereby enhancing the collective yield. These enhancements allow for greater flexibility in processor design by eliminating the need for handpicked orientations. We support the efficient assembly of large processors with dense connectivity by choosing the best boundary conditions. Examples demonstrate that, at low computational cost, this optimization approach finds a frequency configuration for a square chip with over 1000 qubits and over 10% yield at much larger dispersion levels than required by previous approaches.

Zhang, Zewen [Argonne National Laboratory (ANL), A↗

Quantum Computer-Aided Design: Digital Quantum Simulation of Quantum Processors

With the increasing size of quantum processors, submodules that constitute the processor hardware will become too large to accurately simulate on a classical computer. Therefore, one would soon have to fabricate and test each new design primitive and parameter choice in time-consuming coordination between design, fabrication, and experimental validation. Here we show how one can design and test the performance of next-generation quantum hardware—by using existing quantum computers. Focusing on superconducting transmon processors as a prominent hardware platform, we compute the static and dynamic properties of individual and coupled transmons. We show how the energy spectra of transmons can be obtained by variational hybrid quantum-classical algorithms that are well suited for near-term noisy quantum computers. In addition, single- and two-qubit gate simulations are demonstrated via Suzuki-Trotter decomposition. Our methods pave a promising way towards designing candidate quantum processors when the demands of calculating submodule properties exceed the capabilities of classical computing resources.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Mitigation of Cosmic Rays-Induced Errors in Superconducting Quantum Processors

Environmental radioactivity and cosmic-rays have recently been identified as a source of decoherence in super-conducting quantum bits (qubits). In particular, the absorption of cosmic-ray muons and gamma rays emitted by naturally occurring radioactive isotopes in the qubit substrate leads to correlated errors in superconducting quantum processors, posing significant challenges to quantum error correction. To enable quantum computing to scale, it is therefore necessary the devel-opment of mitigation strategies to prevent, or keep under control, error bursts due to particle impacts in the chip. While most environmental radioactive sources can be effectively suppressed using dedicated shielding, cosmic-ray muons, with their high penetration capability, can only be mitigated by moving the entire facility in a deep underground laboratory. This work explores the potential for developing a novel class of quantum processors equipped with an active veto system to protect superconducting-based quantum computers from the detrimental effects of atmospheric muons. Such a device would enable the identification of an atmospheric muon interaction within the processor and veto all operations performed during the occurrence of such an interaction. By demonstrating high detection efficiency and negligible dead time, we aim to establish that the future of quantum processors can be envisioned in above-around facilities.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Detecting and tracking drift in quantum information processors

Abstract If quantum information processors are to fulfill their potential, the diverse errors that affect them must be understood and suppressed. But errors typically fluctuate over time, and the most widely used tools for characterizing them assume static error modes and rates. This mismatch can cause unheralded failures, misidentified error modes, and wasted experimental effort. Here, we demonstrate a spectral analysis technique for resolving time dependence in quantum processors. Our method is fast, simple, and statistically sound. It can be applied to time-series data from any quantum processor experiment. We use data from simulations and trapped-ion qubit experiments to show how our method can resolve time dependence when applied to popular characterization protocols, including randomized benchmarking, gate set tomography, and Ramsey spectroscopy. In the experiments, we detect instability and localize its source, implement drift control techniques to compensate for this instability, and then demonstrate that the instability has been suppressed.

97 MATHEMATICS AND COMPUTING↗

Logical quantum processor based on reconfigurable atom arrays

Suppressing errors is the central challenge for useful quantum computing, requiring quantum error correction (QEC) for large-scale processing. However, the overhead in the realization of error-corrected ‘logical’ qubits, in which information is encoded across many physical qubits for redundancy, poses substantial challenges to large-scale logical quantum computing. Here we report the realization of a programmable quantum processor based on encoded logical qubits operating with up to 280 physical qubits. Using logical-level control and a zoned architecture in reconfigurable neutral-atom arrays, our system combines high two-qubit gate fidelities, arbitrary connectivity, as well as fully programmable single-qubit rotations and mid-circuit readout. Operating this logical processor with various types of encoding, we demonstrate improvement of a two-qubit logic gate by scaling surface-code distance from d = 3 to d = 7, preparation of colour-code qubits with break-even fidelities, fault-tolerant creation of logical Greenberger–Horne–Zeilinger (GHZ) states and feedforward entanglement teleportation, as well as operation of 40 colour-code qubits. Finally, using 3D [[8,3,2]] code blocks, we realize computationally complex sampling circuits with up to 48 logical qubits entangled with hypercube connectivity with 228 logical two-qubit gates and 48 logical CCZ gates. We find that this logical encoding substantially improves algorithmic performance with error detection, outperforming physical-qubit fidelities at both cross-entropy benchmarking and quantum simulations of fast scrambling. These results herald the advent of early error-corrected quantum computation and chart a path towards large-scale logical processors.

74 ATOMIC AND MOLECULAR PHYSICS↗

The Octopus processor for the CMS L1 muon trigger for High Luminosity LHC

The upgraded L1 muon trigger system of the CMS experiment in the High Luminosity Large Hadron Collider is based on custom processors featuring large Field Programmable Gate Arrays (FPGAs) connected by large numbers of optical links. These provide the I/O bandwidth and power necessary to process the complex algorithms used during the collection of physics data. The design and performance requirements of these processors creates significant challenges in signal integrity, power delivery, and thermal management. In this paper we describe the Octopus processor, featuring a large Xilinx Virtex Ultrascale+ FPGA and up to 128 links interfaced to optics through high quality twin-ax copper cables. Results on signal integrity at 25 Gb/s and the first demonstration of 50+ Gb/s links with pluggable optics in CMS are also shown, demonstrating bit error rates below 10 –15 at a 95% confidence level. The thermal performance is measured inside an Advanced-TCA crate with acceptable thermal margins up to 200 W of chip power. Future improvements are mentioned, potentially allowing operation at up to 300 W.

Instruments & Instrumentation↗

Hybrid Oscillator-Qubit Quantum Processors: Instruction Set Architectures, Abstract Machine Models, and Applications

This tutorial offers a pedagogical guide to hybrid quantum processors that integrate discrete-variable (DV) qubits and continuous-variable (CV) oscillators. Aimed at computer scientists, engineers, and physicists, it provides an overview of the experimental, algorithmic, and architectural aspects of this novel and rapidly developing hardware model. Experimental realizations of this model include superconducting, trapped-ion, and neutral-atom platforms. By combining DV and CV components, hybrid oscillator-qubit processors enable a powerful new paradigm that offers complementary strengths for quantum control, error correction, computation, and simulation. Working toward the goal of a full-stack system connecting applications to CV-DV hardware, we define and formulate abstract machine models and instruction set architectures. These essential abstractions enable codesign of hardware and software, and resource estimation for exploring the potential of current and future hardware for computational and simulation tasks. Using these abstractions, we present both new and existing examples that illustrate the benefits of hybrid CV-DV processors relative to traditional DV-only hardware in computation as well as quantum simulation of physical models. Examples include algorithms for transferring states between DV and CV systems, performing the quantum Fourier transform, and simulation of lattice gauge theories. Relative to qubit-only hardware, the bosonic degrees of freedom natively available in hybrid architectures can substantially reduce the circuit complexity of simulations for physical models containing bosons. A key technique is the extension of quantum signal processing ideas to CV-DV systems. This work is intended to serve as a timely and comprehensive guide to this relatively unexplored yet promising approach to quantum computation and to provide a road map to guide future development.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Design Methodologies for Integrated Quantum Frequency Processors

We report frequency-encoded quantum information offers intriguing opportunities for quantum communications and networking, with the quantum frequency processor paradigm—based on electro-optic phase modulators and Fourier-transform pulse shapers—providing a path for scalable construction of quantum gates. Yet all experimental demonstrations to date have relied on discrete fiber-optic components that occupy significant physical space and impart appreciable loss. In this article, we introduce a model for the design of quantum frequency processors comprising microring resonator-based pulse shapers and integrated phase modulators. We estimate the performance of single and parallel frequency-bin Hadamard gates, finding high fidelity values that extend to frequency bins with relatively wide bandwidths. By incorporating multi-order filter designs as well, we explore the limits of tight frequency spacings, a regime extremely difficult to obtain in bulk optics. Overall, our model is general, simple to use, and extendable to other material platforms, providing a much-needed design tool for future frequency processors in integrated photonics.

97 MATHEMATICS AND COMPUTING↗

Two-qubit silicon quantum processor with operation fidelity exceeding 99%

Silicon spin qubits satisfy the necessary criteria for quantum information processing. However, a demonstration of high-fidelity state preparation and readout combined with high-fidelity single- and two-qubit gates, all of which must be present for quantum error correction, has been lacking. We use a two-qubit Si/SiGe quantum processor to demonstrate state preparation and readout with fidelity greater than 97%, combined with both single- and two-qubit control fidelities exceeding 99%. The operation of the quantum processor is quantitatively characterized using gate set tomography and randomized benchmarking. Finally, our results highlight the potential of silicon spin qubits to become a dominant technology in the development of intermediate-scale quantum processors.

97 MATHEMATICS AND COMPUTING↗

Massively parallel and universal approximation of nonlinear functions using diffractive processors

Nonlinear computation is essential for a wide range of information processing tasks, yet implementing nonlinear functions using optical systems remains a challenge due to the weak and power-intensive nature of optical nonlinearities. Overcoming this limitation without relying on nonlinear optical materials could unlock unprecedented opportunities for ultrafast and parallel optical computing systems. Here, we demonstrate that large-scale nonlinear computation can be performed using linear optics through optimized diffractive processors composed of passive phase-only surfaces. In this framework, the input variables of nonlinear functions are encoded into the phase of an optical wavefront—e.g., via a spatial light modulator (SLM)—and transformed by an optimized diffractive structure with spatially varying point-spread functions to yield output intensities that approximate a large set of unique nonlinear functions–all in parallel. We provide proof establishing that this architecture serves as a universal function approximator for an arbitrary set of bandlimited nonlinear functions, also covering wavelength-multiplexed nonlinear functions as well as multi-variate and complex-valued functions that are all-optically cascadable. Our analysis also indicates the successful approximation of typical nonlinear activation functions commonly used in neural networks, including the sigmoid, tanh, ReLU (rectified linear unit), and softplus. We numerically demonstrate the parallel computation of one million distinct nonlinear functions, accurately executed at wavelength-scale spatial density at the output of a diffractive optical processor. Furthermore, we experimentally validated this framework using in situ optical learning and approximated 35 unique nonlinear functions in a single shot using a compact setup consisting of an SLM and an image sensor. These results establish diffractive optical processors as a scalable platform for massively parallel universal nonlinear function approximation, paving the way for new capabilities in analog optical computing based on linear materials.

Rahman, Md Sadman Sakib [University of California,↗