Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “compilation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Resilience–runtime tradeoff relations for quantum algorithms

Abstract A leading approach to algorithm design aims to minimize the number of operations in an algorithm’s compilation. One intuitively expects that reducing the number of operations may decrease the chance of errors. This paradigm is particularly prevalent in quantum computing, where gates are hard to implement and noise rapidly decreases a quantum computer’s potential to outperform classical computers. Here, we find that minimizing the number of operations in a quantum algorithm can be counterproductive, leading to a noise sensitivity that induces errors when running the algorithm in non-ideal conditions. To show this, we develop a framework to characterize the resilience of an algorithm to perturbative noises (including coherent errors, dephasing, and depolarizing noise). Some compilations of an algorithm can be resilient against certain noise sources while being unstable against other noises. We condense these results into a tradeoff relation between an algorithm’s number of operations and its noise resilience. We also show how this framework can be leveraged to identify compilations of an algorithm that are better suited to withstand certain noises.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Evaluating the Benefits of Bayesian Hierarchical Methods for Analyzing Heterogeneous Environmental Datasets: A Case Study of Marine Organic Carbon Fluxes

Large compilations of heterogeneous environmental observations are increasingly available as public databases, allowing researchers to test hypotheses across datasets. Statistical complexities arise when analyzing compiled data due to unbalanced spatial sampling, variable environmental context, mixed measurement techniques, and other reasons. Hierarchical Bayesian modeling is increasingly used in environmental science to describe these complexities, however few studies explicitly compare the utility of hierarchical Bayesian models to simpler and more commonly applied methods. Here we demonstrate the utility of the hierarchical Bayesian approach with application to a large compiled environmental dataset consisting of 5,741 marine vertical organic carbon flux observations from 407 sampling locations spanning eight biomes across the global ocean. We fit a global scale Bayesian hierarchical model that describes the vertical profile of organic carbon flux with depth. Profile parameters within a particular biome are assumed to share a common deviation from the global mean profile. Individual station-level parameters are then modeled as deviations from the common biome-level profile. The hierarchical approach is shown to have several benefits over simpler and more common data aggregation methods. First, the hierarchical approach avoids statistical complexities introduced due to unbalanced sampling and allows for flexible incorporation of spatial heterogeneitites in model parameters. Second, the hierarchical approach uses the whole dataset simultaneously to fit the model parameters which shares information across datasets and reduces the uncertainty up to 95% in individual profiles. Third, the Bayesian approach incorporates prior scientific information about model parameters; for example, the non-negativity of chemical concentrations or mass-balance, which we apply here. We explicitly quantify each of these properties in turn. We emphasize the generality of the hierarchical Bayesian approach for diverse environmental applications and its increasing feasibility for large datasets due to recent developments in Markov Chain Monte Carlo algorithms and easy-to-use high-level software implementations.

54 ENVIRONMENTAL SCIENCES↗

Experimental Characterization of OpenMP Offloading Memory Operations and Unified Shared Memory Support

The OpenMP specification recently introduced support for unified shared memory, allowing implementation to leverage underlying system software to provide a simpler GPU offloading model where explicit mapping of variables is optional. Support for this feature is becoming more available in different OpenMP implementations on several hardware platforms. A deeper understanding of the different implementation’s execution profile and performance is crucial for applications as they consider the performance portability implications of adopting a unified memory offloading programming style. This work introduces a benchmark tool to characterize unified memory support in several OepnMP compilers and runtimes, with emphasis on identifying discrepancies between different OpenMP implementations as to how they various memory allocation strategies interact with unified shared memory. The benchmark tool is used to characterize OpenMP compilers on three leading High Performance Computing platforms supporting different CPU and device architectures. The benchmark tool is used to assess the impact of enabling unified shared memory on the performance of memory-bound code, highlighting implementation differences that should be accounted for when applications consider performance portability across platforms and compilers.

Elwasif, Wael↗

OpenACC Unified Programming Environment for Multi-hybrid Acceleration with GPU and FPGA

Accelerated computing in HPC such as with GPU, plays a central role in HPC nowadays. However, in some complicated applications with partially different performance behavior is hard to solve with a single type of accelerator where GPU is not the perfect solution in these cases. We are developing a framework and transpiler allowing the users to program the codes with a single notation of OpenACC to be compiled for multi-hybrid accelerators, named MHOAT (Multi-Hybrid OpenACC Translator) for HPC applications. MHOAT parses the original code with directives to identify the target accelerating devices, currently supporting NVIDIA GPU and Intel FPGA, dispatching these specific partial codes to background compilers such as NVIDIA HPC SDK for GPU and OpenARC research compiler for FPGA, then assembles binaries for the final object with FPGA bitstream file. In this paper, we present the concept, design, implementation, and performance evaluation of a practical astrophysics simulation code where we successfully enhanced the performance up to 10 times faster than the GPU-only solution.

Boku, Taisuke↗

Profile Generation for GPU Targets

GPU accelerators are ubiquitous, but their ecosystem is far less evolved than the host one. Compiler heuristics are often tuned for CPUs and reused for GPU. Similarly, tooling and more evolved optimization techniques are historically not available on GPU targets. In this work, we address one of these shortcomings and enable profile generation and profile-guided optimizations (PGO) for GPU targets. While this is only a single step towards a CPU equivalent ecosystem for offload devices, it shows how old misconceptions on the limitations of GPUs are often not warranted anymore. Through our implementation in LLVM/Offload, we enable device-side PGO for full scientific applications and open up tooling opportunities, including code coverage analysis and compiler-built-in roofline analysis. Our evaluation highlights the performance implications of profile generation, the insights gained from these profiles, and the (missed) opportunities in utilizing the information for GPU compilation.

McDonough, Ethan Luis [Lawrence Livermore National↗

Enabling Fortran Standard Parallelism in GAMESS for Accelerated Quantum Chemistry Calculations

The performance of Fortran 2008 DO CONCURRENT (DC) relative to OpenACC and OpenMP target offloading (OTO) with different compilers is studied for the GAMESS quantum chemistry application. Specifically, DC and OTO are used to offload the Fock build, which is a computational bottleneck in most quantum chemistry codes, to GPUs. The DC Fock build performance is studied on NVIDIA A100 and V100 accelerators and compared with the OTO versions compiled by the NVIDIA HPC, IBM XL, and Cray Fortran compilers. The results show that DC can speed up the Fock build by 3.0× compared with that of the OTO model. Finally, with similar offloading efforts, DC is a compelling programming model for offloading Fortran applications to GPUs.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Quantum logic gate synthesis as a Markov decision process

Reinforcement learning has witnessed recent applications to a variety of tasks in quantum programming. The underlying assumption is that those tasks could be modeled as Markov decision processes (MDPs). Here, we investigate the feasibility of this assumption by exploring its consequences for single-qubit quantum state preparation and gate compilation. By forming discrete MDPs, we solve for the optimal policy exactly through policy iteration. We find optimal paths that correspond to the shortest possible sequence of gates to prepare a state or compile a gate, up to some target accuracy. Our method works in both the absence and presence of noise and compares favorably to other quantum compilation methods, such as the Ross–Selinger algorithm. This work provides theoretical insight into why reinforcement learning may be successfully used to find optimally short gate sequences in quantum programming.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Evaluating Performance Portability with the CMS Heterogeneous Pixel Reconstruction code

In the past years the landscape of tools for expressing parallel algorithms in a portable way across various compute accelerators has continued to evolve significantly. There are many technologies on the market that provide portability between CPU, GPUs from several vendors, and in some cases even FPGAs. These technologies include C++ libraries such as Alpaka and Kokkos, compiler directives such as OpenMP, the SYCL open specification that can be implemented as a library or in a compiler, and standard C++ where the compiler is solely responsible for the offloading. Given this developing landscape, users have to choose the technology that best fits their applications and constraints. For example, in the CMS experiment the experience so far in heterogeneous reconstruction algorithms suggests that the full application contains a large number of relatively short computational kernels and memory transfer operations. In this work we use a stand-alone version of the CMS heterogeneous pixel reconstruction code as a realistic use case of HEP reconstruction software that is capable of leveraging GPUs effectively. We summarize the experience of porting this code base from CUDA to Alpaka, Kokkos, SYCL, std::par, and OpenMP offloading. We compare the event processing throughput achieved by each version on NVIDIA and AMD GPUs as well as on a CPU, and compare those to what a native version of the code achieves on each platform.

Andriotis, Nikolaos↗

Cosmological constraints from H ii starburst galaxy, quasar angular size, and other measurements

ABSTRACT We compare the constraints from two (2019 and 2021) compilations of H ii starburst galaxy (H iiG) data and test the model independence of quasar (QSO) angular size data using six spatially flat and non-flat cosmological models. We find that the new 2021 compilation of H iiG data generally provides tighter constraints and prefers lower values of cosmological parameters than those from the 2019 H iiG data. QSO data by themselves give relatively model-independent constraints on the characteristic linear size, lm, of the QSOs within the sample. We also use Hubble parameter [H(z)], baryon acoustic oscillation (BAO), Pantheon Type Ia supernova (SN Ia) apparent magnitude (SN-Pantheon), and DES-3 yr binned SN Ia apparent magnitude (SN-DES) measurements to perform joint analyses with H iiG and QSO angular size data, since their constraints are not mutually inconsistent within the six cosmological models we study. A joint analysis of H(z), BAO, SN-Pantheon, SN-DES, QSO, and the newest compilation of H iiG data provides almost model-independent summary estimates of the Hubble constant, $H_0=69.7\pm 1.2\ \rm {km\,s^{-1}\,Mpc^{-1}}$, the non-relativistic matter density parameter, $\Omega _{\rm m_0}=0.293\pm 0.021$, and lm = 10.93 ± 0.25 pc.

79 ASTRONOMY AND ASTROPHYSICS↗

Do quasar X-ray and UV flux measurements provide a useful test of cosmological models?

ABSTRACT The recent compilation of quasar (QSO) X-ray and ultraviolet (UV) flux measurements include QSOs that appear to not be standardizable via the X-ray luminosity and UV luminosity (LX–LUV) relation and so should not be used to constrain cosmological model parameters. Here, we show that the largest of seven sub-samples in this compilation, the SDSS-4XMM QSOs that contribute about 2/3 of the total QSOs, have LX–LUV relations that depend on the cosmological model assumed and also on redshift, and is the main cause of the similar problem discovered earlier for the full QSO compilation. The second and third biggest sub-samples, the SDSS-Chandra and XXL QSOs that together contribute about 30 per cent of the total QSOs, appear standardizable, but provide only weak constraints on cosmological parameters that are not inconsistent with the standard spatially flat ΛCDM model or with constraints from better-established cosmological probes.

79 ASTRONOMY AND ASTROPHYSICS↗

Gamma-ray burst data strongly favour the three-parameter fundamental plane (Dainotti) correlation over the two-parameter one

ABSTRACT Gamma-ray bursts (GRBs), observed to redshift z = 9.4, are potential probes of the largely unexplored z ∼ 2.7–9.4 part of the early Universe. Thus, finding relevant relations among GRB physical properties is crucial. We find that the Platinum GRB data compilation, with 50 long GRBs (with relatively flat plateaus and no flares) in the redshift range 0.553 ≤ z ≤ 5.0, and the LGRB95 data compilation, with 95 long GRBs in 0.297 ≤ z ≤ 9.4, as well as the 145 GRB combination of the two, strongly favour the 3D Fundamental Plane (Dainotti) correlation (between the peak prompt luminosity, the luminosity at the end of the plateau emission, and its rest-frame duration) over the 2D one (between the luminosity at the end of the plateau emission and its duration). The 3D Dainotti correlations in the three data sets are standardizable. We find that while LGRB95 data have ∼50 per cent larger intrinsic scatter parameter values than the better-quality Platinum data, they provide somewhat tighter constraints on cosmological-model and GRB-correlation parameters, perhaps solely due to the larger number of data points, 95 versus 50. This suggests that when compiling GRB data for the purpose of constraining cosmological parameters, given the quality of current GRB data, intrinsic scatter parameter reduction must be balanced against reduced sample size.

79 ASTRONOMY AND ASTROPHYSICS↗

Simulating non-native cubic interactions on noisy quantum machines

As a milestone for general-purpose computing machines, we demonstrate that quantum processors can be programed to efficiently simulate dynamics that are not native to the hardware. Moreover, on noisy devices without error correction, we show that simulation results are significantly improved when the quantum program is compiled using modular gates instead of a restricted set of standard gates. We demonstrate the general methodology by solving a cubic interaction problem, which appears in nonlinear optics, gauge theories, as well as plasma and fluid dynamics. To encode the non-native Hamiltonian evolution, we decompose the Hilbert space into a direct sum of invariant subspaces in which the nonlinear problem is mapped to a finite-dimensional Hamiltonian simulation problem. Furthermore, in a three-states example, the resultant unitary evolution is realized by a product of approximately 20 standard gates, using which approximately ten simulation steps can be carried out on state-of-the-art quantum hardware before results are corrupted by decoherence. In comparison, the simulation depth is improved by more than an order of magnitude when the unitary evolution is realized as a single cubic gate, which is compiled directly using optimal control. Alternatively, parametric gates may also be compiled by interpolating control pulses. Modular gates thus obtained provide high-fidelity building blocks for quantum Hamiltonian simulations.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

On Using Linux Kernel Huge Pages with FLASH, an Astrophysical Simulation Code

We present efforts at improving the performance of FLASH, a multi-scale, multi-physics simulation code principally for astrophysical applications, by using huge pages on Ookami, an HPE Apollo 80 A64FX platform. FLASH is written principally in modern Fortran and makes use of the PARAMESH library to manage a block-structured adaptive mesh. We explored options for enabling the use of huge pages with several compilers, but we were only able to successfully use huge pages when compiling with the Fujitsu compiler. As a result, the use of huge pages substantially reduced the number of translation lookaside buffer misses, but overall performance gains were marginal.

79 ASTRONOMY AND ASTROPHYSICS↗

A Pulse Generation Framework with Augmented Program-aware Basis Gates and Criticality Analysis

Near-term intermediate scale quantum (NISQ) de- vices are subject to considerable noise and short coherence time. Consequently, it is critical to minimize circuit execution latency. Traditionally, each basis gate of a transpiled circuit is decoded into a fixed episode of the device control pulses. Recently, people started to investigate merged pulse generation for customized gates through quantum optimal control (QOC). However, existing QOC approaches face the challenges of (i) restricted search space due to prohibitive compilation overhead; (ii) suboptimal end-to-end performance due to aggressive local optimization and falsely introduced dependency among the customized gates; (iii) inadequate adaptivity towards system calibration, which is critical for NISQ devices. In this work, we propose PAQOC, a novel QOC framework that can (i) automatically detect frequently encountered gate patterns in the logical circuit by modeling the problem as a subgraph mining process and reuse these patterns to enable much larger search space exploration (i.e., program aware); (ii) systemically construct customized gate-set based on the impact to the overall program latency (i.e., criticality-aware); and (iii) quickly adapt to system re-calibration thanks to the small-scale pattern-based gate generation (i.e., adaptivity-aware). PAQOC achieves a good tradeoff between circuit performance and compilation time, allowing fully automatic, single stop, ad- hoc customized pulse generation for more efficient execution of user programs on NISQ devices. Evaluations using fifteen applications show that PAQOC can achieve on average 1.95× speedup of the circuit latency and achieve on average 36.7% reduction in compilation overhead. With PAQOC, circuits can run faster with reduced noise, allowing deeper circuits to be tested within the coherence time of present NISQ platforms.

Chen, Yanhao↗

Reoptimization of Quantum Circuits via Hierarchical Synthesis

The current phase of quantum computing is in the Noisy Intermediate-Scale Quantum (NISQ) era. On NISQ devices, two-qubit gates such as CNOTs are much noisier than single-qubit gates, so it is essential to minimize their count. Quantum circuit synthesis is a process of decomposing an arbitrary unitary into a sequence of quantum gates, and can be used as an optimization tool to produce shorter circuits to improve overall circuit fidelity. However, the time-to-solution of synthesis grows exponentially with the number of qubits. As a result, synthesis is intractable for circuits on a large qubit scale. In this paper, we propose a hierarchical, block-by-block opti-mization framework, QGo, for quantum circuit optimization. Our approach allows an exponential cost optimization to scale to large circuits. QGo uses a combination of partitioning and synthesis: 1) partition the circuit into a sequence of independent circuit blocks; 2) re-generate and optimize each block using quantum synthesis; and 3) re-compose the final circuit by stitching all the blocks together. We perform our analysis and show the fidelity improvements in three different regimes: small-size circuits on real devices, medium-size circuits on noisy simulations, and large-size circuits on analytical models. Our technique can be applied after existing optimizations to achieve higher circuit fidelity. Further, using a set of NISQ benchmarks, we show that QGo can reduce the number of CNOT gates by 29.9% on average and up to 50% when compared with industrial compiler optimizations such as t|ket). When executed on the IBM Athens system, shorter depth leads to higher circuit fidelity. We also demonstrate the scalability of our QGo technique to optimize circuits of 60+ qubits, Our technique is the first demonstration of successfully employing and scaling synthesis in the compilation tool chain for large circuits. Overall, our approach is robust for direct incorporation in production compiler toolchains to further improve the circuit fidelity.

97 MATHEMATICS AND COMPUTING↗

PowerMappeR: Power-Optimized Mapping of SNNs onto ReRAM Crossbars coupled via Packet-Switched NoCs

Many recent efforts in developing hardware-accelerated spiking neural networks (SNNs) are characterized by deep co-design between algorithms, architectures, and devices. Architectural advances overcome device constraints by coupling together many small resistive-RAM (ReRAM) crossbars via a network-on-chip (NoC) for neuromorphic component operation. Concurrently, improved SNN training methods increase accuracy and structural sparsity in networks despite growing problem sizes. Finally, compilers leverage these attributes to minimize area and inter-crossbar communication while mapping large SNNs to sophisticated architectures. However, for compiler-driven co-design to realize increasingly complex and profitable optimizations, a compile-time view of power consumption is critical. We present PowerMappeR to express and optimize over mapping-, architecture-, and device-specific power consumption information. By modeling the dynamic power of well-established components, we develop an integer linear programming (ILP)-based, encoding-agnostic, parametric power estimation model. Using this model, we demonstrate practical improvements in area and inter-crossbar communication by 0%–9.5% and 1.4%–5.1%, respectively. We also limit hotspot formation during optimization, achieving comparable or better results in targeted metrics with up to 96.4%–97.1% restriction of hotspot magnitude. Finally, we introduce profile-guided formulations to reduce worst-case and expected-case hotspot magnitude by 40.7%–69.5% and 40.6%–56.3%, respectively. Optimizing worst-case hotspot magnitude incidentally improves expected-case magnitude by 10.85%–33.45%. Reciprocally, optimizing expected-case magnitude incidentally improves worst-case magnitude by 4.33%–39.87%. Validation against hardware simulators confirms that PowerMappeR can decrease dynamic power consumption by 12.6%–27.3%.

Pohl, Devin [ORNL] (ORCID:0009000040149027)↗

Let Each Quantum Bit Choose Its Basis Gates

Near-term quantum computers are primarily limited by errors in quantum operations (or gates) between two quantum bits (or qubits). A physical machine typically provides a set of basis gates that include primitive 2-qubit (2Q) and 1-qubit (1Q) gates that can be implemented in a given technology. 2Q entangling gates, coupled with some 1Q gates, allow for universal quantum computation. In superconducting technologies, the current state of the art is to implement the same 2Q gate between every pair of qubits (typically an XX-or XY-type gate). This strict hardware uniformity requirement for 2Q gates in a large quantum computer has made scaling up a time and resource-intensive endeavor in the lab. We propose a radical idea – allow the 2Q basis gate(s) to differ between every pair of qubits, selecting the best entangling gates that can be calibrated between given pairs of qubits. This work aims to give quantum scientists the ability to run meaningful algorithms with qubit systems that are not perfectly uniform. Scientists will also be able to use a much broader variety of novel 2Q gates for quantum computing. We develop a theoretical framework for identifying good 2Q basis gates on “nonstandard” Cartan trajectories that deviate from “standard” trajectories like XX. We then introduce practical methods for calibration and compilation with nonstandard 2Q gates, and discuss possible ways to improve the compilation. To demonstrate our methods in a case study, we simulated both standard XY-type trajectories and faster, nonstandard trajectories using an entangling gate architecture with far-detuned transmon qubits. We identify efficient 2Q basis gates on these nonstandard trajectories and use them to compile a number of standard benchmark circuits such as QFT and QAOA. Furthermore, our results demonstrate an 8x improvement over the baseline 2Q gates with respect to speed and coherence-limited gate fidelity.

quantum computing↗

Extending XACC for Quantum Optimal Control

Quantum computing vendors are beginning to open up application programming interfaces for direct pulse-level quantum control. With this, programmers can begin to describe quantum kernels of execution via sequences of arbitrary pulse shapes. This opens new avenues of research and development with regards to smart quantum compilation routines that enable direct translation of higher-level digital assembly representations to these native pulse instructions. In this work, we present an extension to the XACC system-level quantum-classical software framework that directly enables this compilation lowering phase via user-specified quantum optimal control techniques. This extension enables the translation of digital quantum circuit representations to equivalent pulse sequences that are optimal with respect to the backend system dynamics. Our work is modular and extensible, enabling third party optimal control techniques and strategies in both C++ and Python. We demonstrate this extension with familiar gradient-based methods like gradient ascent pulse engineering (GRAPE), gradient optimization of analytic controls (GOAT), and Krotov's method. Our work serves as a foundational component of future quantum-classical compiler designs that lower high-level programmatic representations to low-level machine instructions.

Nguyen, Thien↗