Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel computer architecture”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20

Editorial: Neuroscience, computing, performance, and benchmarks: Why it matters to neuroscience how fast we can compute

At the turn of the millennium the computational neuroscience community realized that neuroscience was in a software crisis: software development was no longer progressing as expected and reproducibility declined. The International Neuroinformatics Coordinating Facility (INCF) was inaugurated in 2007 as an initiative to improve this situation. The INCF has since pursued its mission to help the development of standards and best practices. In a community paper published this very same year, Brette et al. tried to assess the state of the field and to establish a scientific approach to simulation technology, addressing foundational topics, such as which simulation schemes are best suited for the types of models we see in neuroscience. In 2015, a Frontiers Research Topic “Python in neuroscience” by Muller et al. triggered and documented a revolution in the neuroscience community, namely in the usage of the scripting language Python as a common language for interfacing with simulation codes and connecting between applications. The review by Einevoll et al. documented that simulation tools have since further matured and become reliable research instruments used by many scientific groups for their respective questions. Open source and community standard simulators today allow research groups to focus on their scientific questions and leave the details of the computational work to the community of simulator developers. A parallel development has occurred, which has been barely visible in neuroscientific circles beyond the community of simulator developers: Supercomputers used for large and complex scientific calculations have increased their performance from ~10 TeraFLOPS (10 13 floating point operations per second) in the early 2000s to above 1 ExaFLOPS (10 18 floating point operations per second) in the year 2022. This represents a 100,000-fold increase in our computational capabilities, or almost 17 doublings of computational capability in 22 years. Moore's law (the observation that it is economically viable to double the number of transistors in an integrated circuit every other 18–24 months) explains a part of this; our ability and willingness to build and operate physically larger computers, explains another part. It should be clear, however, that such a technological advancement requires software adaptations and under the hood, simulators had to reinvent themselves and change substantially to embrace this technological opportunity. It actually is quite remarkable that—apart from the change in semantics for the parallelization—this has mostly happened without the users knowing. The current Research Topic was motivated by the wish to assemble an update on the state of neuroscientific software (mostly simulators) in 2022, to assess whether we can see more clearly which scientific questions can (or cannot) be asked due to our increased capability of simulation, and also to anticipate whether and for how long we can expect this increase of computational capabilities to continue.

biophysically detailed models↗

Accuracy of the explicit energy-conserving particle-in-cell method for under-resolved simulations of capacitively coupled plasma discharges

The traditional explicit electrostatic momentum-conserving particle-in-cell algorithm requires strict resolution of the electron Debye length to deliver numerical stability and accuracy. The explicit electrostatic energy-conserving particle-in-cell algorithm alleviates this constraint with minimal modification to the traditional algorithm, retaining its simplicity, ease of parallelization, and acceleration on modern supercomputing architectures. In this article, we apply the algorithm to model a one-dimensional radio frequency capacitively coupled plasma discharge relevant to industrial applications. The energy-conserving approach closely matches the results from the momentum-conserving algorithm and retains accuracy even for cell sizes up to 8 times the electron Debye length. For even larger cells, the algorithm loses accuracy due to poor resolution of steep gradients within the radio frequency sheath. Accuracy can be recovered by adopting a non-uniform grid, which resolves the sheath and allows for cell sizes up to 32 times the electron Debye length in the quasi-neutral bulk of the discharge. The effect is an up to 8 times reduction in the number of required simulation cells, an improvement that can compound in higher-dimensional simulations. We therefore consider the explicit energy-conserving algorithm as a promising approach to significantly reduce the computational cost of full-scale device simulations and a pathway to delivering kinetic simulation capabilities of use to industry.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Parallel Runtime Interface for Fortran (PRIF) Specification (Rev. 0.5)

This document specifies an interface to support the parallel features of Fortran, named the Parallel Runtime Interface for Fortran (PRIF). PRIF is a proposed solution in which the runtime library is primarily responsible for implementing coarray allocation, deallocation and accesses, image synchronization, atomic operations, events, teams and collective subroutines. In this interface, the compiler is responsible for transforming the invocation of Fortran-level parallel features into procedure calls to the necessary PRIF subroutines. The interface is designed for portability across shared- and distributed-memory machines, different operating systems, and multiple architectures. Implementations of this interface are intended as an augmentation for the compiler's own runtime library. With an implementation-agnostic interface, alternative parallel runtime libraries may be developed that support the same interface. One benefit of this approach is the ability to vary the communication substrate. A central aim of this document is to define a parallel runtime interface in standard Fortran syntax, which enables us to leverage Fortran to succinctly express various properties of the procedure interfaces, including argument attributes.

97 MATHEMATICS AND COMPUTING↗

Learning local equivariant representations for large-scale atomistic dynamics

Abstract A simultaneously accurate and computationally efficient parametrization of the potential energy surface of molecules and materials is a long-standing goal in the natural sciences. While atom-centered message passing neural networks (MPNNs) have shown remarkable accuracy, their information propagation has limited the accessible length-scales. Local methods, conversely, scale to large simulations but have suffered from inferior accuracy. This work introduces Allegro, a strictly local equivariant deep neural network interatomic potential architecture that simultaneously exhibits excellent accuracy and scalability. Allegro represents a many-body potential using iterated tensor products of learned equivariant representations without atom-centered message passing. Allegro obtains improvements over state-of-the-art methods on QM9 and revMD17. A single tensor product layer outperforms existing deep MPNNs and transformers on QM9. Furthermore, Allegro displays remarkable generalization to out-of-distribution data. Molecular simulations using Allegro recover structural and kinetic properties of an amorphous electrolyte in excellent agreement with ab-initio simulations. Finally, we demonstrate parallelization with a simulation of 100 million atoms.

74 ATOMIC AND MOLECULAR PHYSICS↗

Oxide-nitride heteroepitaxy for low-loss dielectrics in superconducting quantum circuits

Superconducting qubits show great promise for the realization of fault-tolerant quantum computing, but lossy, amorphous dielectrics limit current technology. Identifying highly crystalline and stoichiometric dielectrics with intrinsically low microwave loss is therefore a central materials challenge, yet experimentally validated platforms remain scarce. In this work, we integrate a crystalline dielectric into a heteroepitaxial TiN/$γ$-Al$_2$O$_3$/TiN trilayer grown via pulsed laser deposition. Correlative high-resolution imaging, diffraction, and spectroscopy measurements confirm the single-crystal quality and chemical integrity of all layers, with minimal defects and limited anion interdiffusion across the oxide-nitride interfaces. Using microwave lumped-element resonators with parallel-plate capacitors, we report the first direct measurement of the dielectric loss of epitaxial $γ$-Al$_2$O$_3$, for which we find a low intrinsic two-level system loss, $δ_{\text{TLS}}^0 = (2.8 \pm 0.1) \times 10^{-5}$. These results establish heteroepitaxial oxides on transition metal nitrides as an attractive materials platform for superconducting quantum circuits, particularly for integration into compact device architectures such as merged-element transmons and microwave kinetic inductance detectors.

Garcia-Wetten, David A. [Northwestern U.]↗

VTK-m User's Guide (V.1.8)

Extreme scale systems require a new programming model and a fundamental change in how we design algorithms. To address these issues we created VTK-m: the visualization toolkit for multi-/many-core architectures. VTK-m supports a number of algorithms and the ability to design further algorithms through a top-down design with an emphasis on extreme parallelism. VTK-m also provides support for finding and building links across topologies, making it possible to perform operations that determine manifold surfaces, interpolate generated values, and find adjacencies. Although VTK-m provides a simplified high-level interface for programming, its template-based code removes the overhead of abstraction. VTK-m simplifies the development of parallel scientific visualization algorithms by providing a framework of supporting functionality that allows developers to focus on visualization operations. Additionally, VTK-m provides data abstractions not provided by other libraries that make code written in VTK-m more versatile.

97 MATHEMATICS AND COMPUTING↗

DYFLOW: A flexible framework for orchestrating scientific workflows on supercomputers

Modern scientific workflows are increasing in complexity with growth in computation power, incorporation of non-traditional computation methods, and advances in technologies enabling data streaming to support on-the-fly computation. These workflows have unpredictable runtime behaviors, and a fixed, predetermined resource assignment on supercomputers can be inefficient for overall performance and throughput. Inability to change resource assignments further limits the scientists to avail of science-driven opportunities or respond to failures.We introduce DYFLOW, a flexible framework that orchestrates scientific workflows on supercomputers based on user-designed policies. DYFLOW compartmentalizes orchestration stages into simplified constructs, and end-users can program and reuse them according to their workflow requirements through an easy-to-use interface. These constructs hide the intricacies involved in runtime management from end-users, for instance, procurement of information to understand the workflow state, assessment, and supervision of the runtime changes. DYFLOW is designed to work alongside existing workflow management systems and reuse the available (static) support for workflow management. We have integrated DYFLOW with an existing workflow management tool as a demonstration. With experiments performed on use cases from three types of scientific workflows and two different parallel architectures, we show that DYFLOW achieves the desired orchestration incurring a small cost to carry out the runtime changes.

Singhal, Swati↗

PipeSight: A High-Performance Computing Platform for Pipeline Integrity Management

The Phase I feasibility study completed as part of this project has led to a number of innovative technologies being developed and has laid the foundation for a successful Phase II effort to commercialize a platform for managing the integrity of pipelines for the damage mechanisms of the new, hybrid-energy based economy. To ground the development efforts and direction of the project, an extensive market research and customer discovery effort was undertaken early in Phase I. Through this effort, a number of pipeline owners and operators were interviewed, and the following key findings were discovered about the pipeline industry: • Small pipeline operators do not have the central engineering groups necessary to perform their own independent analysis of inspection data, but instead rely on summarized tally sheets provided to them by inspection service providers. • The time it takes to go from an inspection to a completed engineering assessment, even for small segments of pipeline, can take anywhere from 30-120 days. During this delay, critical threats can (and have been known to) cause failures. • Uncertainty is often not accounted for in the assessment of pipeline integrity. The tally sheets provided by third-party service providers are almost always deterministic in nature, identifying threats that present a concern only to the current (not the future) integrity of the pipeline. • It is uncommon to apply the latest technologies to perform advanced assessments of damaged pipelines. There is a desire to use more advanced analysis capabilities to assess threats. Many pipeline operators indicated that they would often excavate a pipeline to perform an inspection and find that the damage was not as bad as they anticipated, thus using limited resources unnecessarily. Companies are not consistent in their use of inspection data to determine corrosion rates, and those that do only calculate deterministic corrosion rates. • The industry has prominently relied on time-based inspections but has recently started to transition to risk-based inspections. However, there appears to be no uniform guidance on how to do so while properly accounting for all sources of uncertainty. • Companies are not storing inspection data in a manner that allows for the ready determination of temporal trends. • Predictive maintenance principles and practices are beginning to be used by early adopters • Some pipelines are being re-purposed to transport different process fluids than they were designed for, e.g., H 2 and CO 2 rich process streams to serve the new hybrid-energy based economy, which are presenting new integrity concerns for the existing pipeline network that crisscrosses the United States. As a result of these discoveries, we were able to target the development efforts in Phase I to best serve the needs of the industry. In Phase I, we developed a way to correlate multiple large-scale scans of the pipeline to determine a probabilistic corrosion rate that accounts for all sources of error and uncertainty in the inspection process. This probabilistic corrosion rate can be used to predict the future thickness distribution of the pipe wall. We demonstrate how this analysis may be performed in an analytical fashion and has been implemented in such a manner that it can be readily distributed using GPU computing through integration of the Kokkos programming model. We also make a very novel extension of the analytical corrosion rate model to Bayesian Networks (an explainable AI technique) that can account for non-parametric distributions of corrosion rates. With the predictions made above for the probabilistic corrosion rate and corresponding future distribution of the pipe wall thickness, we can assess the integrity of the pipeline through the use of a probabilistic engineering assessment. We developed a novel screening data analysis approach that can rapidly identify ‘hotspots’ (local thin areas) where the integrity of the pipeline is a concern. Once more, we implemented this screening approach in C++ to leverage GPU computing via the Kokkos programming model. After the critical hotspots are identified, we developed a program that can automatically generate an advanced finite element model of the damaged regions. Since the number of damaged regions that require advanced analysis can number in the thousands, we integrated an open-source container-native workflow engine for orchestrating parallel jobs on the cloud. Initially, these advanced numerical models were only designed to account for loading due to internal pressure. However, in a slight pivot from the initial Phase I proposal, we developed a complete pipe stress analysis program (called Simflex) which can simulate the complete pipeline and its response to thermal expansion, pressure, thermal bowing, weight, wind, earthquake, support displacement, support friction and external forces. This pipe stress analysis program was written generically, to handle any piping system, but contains the features needed to model long pipelines (i.e., it incorporates a model for soil mechanics and can account for the nonlinear boundary conditions necessary to simulate long underground pipelines). This pipe stress analysis program can simulate any segment of the pipeline (simple or complex) under any set of conditions and loads, to determine the supplemental loads (axial forces and bending moments) at the location of damage. This enables the most accurate state of stress to be accounted for in the pipeline, which can prove critical when evaluating the integrity of a damaged region. In the process of developing the technologies to perform the integrity assessment of the pipeline, we also extended one of the industry standard approaches for performing the assessment of local thin areas that extend more in the circumferential direction than the longitudinal direction of the pipeline. This approach was presented to the API 579-1/AS ME FFS-1 steering committee in November 2021 for consideration in the next edition of the industry standard for Fitness-For-Service (expected to be released in 2023). To help pipeline operators make decisions with the results on any integrity assessment, we developed a new approach to the life-cycle management of pipelines which uses a Bayesian Decision Network. The network is designed to help pipeline operators plan and prioritize inspection activities and ultimately make smarter, more cost-effective decisions. The Bayesian approach accounts for all sources of uncertainty and carries them through to the final optimal decisions, providing a probabilistic framework for optimizing inspection intervals. The proof-of-concept networks developed in the feasibility study are complete, verified, and are focused on a subset of the pipeline. To expand this novel approach to the scale necessary for an entire network of pipelines in Phase II, we will leverage the DOE-funded Bengi solver for industrial-scale decision making with Bayesian Networks [22]. Once implemented, we will be able to provide the pipeline industry with a much-needed tool for optimal inspection planning using truly explainable artificial intelligence (XAI). To handle all of these advanced capabilities into a cloud-based platform, the architecture of the Equity Engineering Cloud (EEC) was extended to include Argo Workflows, a framework capable of distributing and managing a massive number of jobs that consume their own resources, such that thousands of serial finite element simulations can be run in parallel. As part of this substantial undertaking, we also integrated Argo Continuous Delivery (CD) into the EEC, to aid with the rapid prototyping and iterations that will be imperative to the success of the PipeSight platform’s Agile development process in Phase II. As part of the pipe stress analysis program, we also developed a custom visualizer that leverages the DOE-funded VTK visualization library. We added custom contouring capabilities and a means for interacting visually with both the inputs and outputs of the pipe stress analysis program. We also developed routines for automating the post-processing of the finite element simulations to determine if any failure criteria are met and to visualize the deformations, stresses and strains in ParaView using the exodus II file format (a subset of netCDF).

24 POWER TRANSMISSION AND DISTRIBUTION↗

Developing And Scaling an OpenFOAM Model to Study Turbulent Flow in a HFIR Coolant Channel

Improving the understanding of how computational fluid dynamics (CFD) direct numerical simulations (DNS) of flows in the High Flux Isotope Reactor (HFIR) perform when run in parallel using the high performance computing (HPC) platform Summit at the Oak Ridge Leadership Computing Facility (OLCF) is of particular importance to boost the computational tools used to support HFIR conversion to low enriched fuel (LEU). Evaluation of scaling performance was driven by the increasing importance of graphics processing unit (GPU) usage in HPC, which is becoming the standard for modern supercomputers such as Summit. The desired results are to obtain a strong positive correlation between the computational resources dedicated to a problem and the relative speed-up of the simulation in comparison to a benchmark. This capability will allow substantially improvement in HFIR flow analytical capabilities, specifically when predicting turbulence properties at high Reynolds numbers. The study leverages previous simulation results performed with code PHASTA (finite element) on HPC platforms Cori (NERSC) and Theta (ALCF) [1] with computing options provided in the computing platform OpenFOAM (finite volume) at OLCF. Transitioning from PHASTA to OpenFOAM will (1) eliminate dependence on third-party software for mesh generation and manipulation, (2) reduce resource needs by employing modern architectures, and (3) build expertise for future modeling of HFIR-specific problems like heat transfer in involute geometry, entrance effects, flow structure in channel corners, and so on—all important issues when defining the available thermal margins in the transition to LEU. CPUs and GPUs differ significantly in their architecture and utilization, as discussed in the literature [2]. The most important differences are in the approach to computations and their memory. A single GPU contains a large quantity of cores, enabling it to perform with a much higher throughput than a CPU, but execution requires a different approach. GPU codes execute instructions using the Single-Instruction Multiple-Thread (SIMT) approach in which a single instruction is used for groups of threads called warps. A warp typically consists of 32 threads which must execute the same set of instructions, although on separate threads. Alternately, a CPU has far fewer cores that are much more flexible in their operation, excelling at quickly performing more complex serial computations. This is why GPUs have greater throughput when properly utilized. The second important difference is seen when comparing their memory spaces. Limited memory allocations and CPU–GPU communications cause a significant bottleneck in GPU-accelerated programs. Further study was required to properly take advantage of GPU resources. A comprehensive analysis of code performance and the model-specific features of turbulence constitutes the core of this work. In this study, a DNS simulation of HFIR channel turbulence was performed with the finite volume CFD code OpenFOAM v2112 and CUDA v11.0 on Red Hat Enterprise Linux v8.2. The OpenFOAM installation had AMGx integrated to enable GPU acceleration and utilizes the PETSc4FOAM library. The computational resources and the problem size were scaled on CPU and CPU + GPU architectures to gain a better understanding of the performance of a DNS problem on modern computing hardware. The study aimed to analyze the scaling of the code exclusively on CPUs and then to examine the scaling of the codes with GPU acceleration enabled. Scaling studies included CPU and GPU acceleration on a mesh of varying resolution to analyze the impact of problem size relative to computational resources. In the course of preparing the GPU configuration on Summit, mainly using the AMGX solvers, difficulties were encountered stemming from constant changes resulting from extensive ongoing development activities and the changing environment. This resulted in the inability to complete the GPU portion of the work. The code was compiled and tested, but production runs to assess acceleration were not performed because the used discretional compute time allocation expired as year-end approached. The Summit HPC platform is scheduled for decommissioning in 2024, making it unattractive for future use with Nvidia-based GPUs. Therefore, the work will be moved onto NERSC machines in FY24. An application was prepared and submitted, and sufficient node-hours were awarded to continue the research in the next calendar year. This report summarizes work performed thus far, which mostly focused on CPU OpenFOAM computing.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Parallel Runtime Interface for Fortran (PRIF) Specification (Rev. 0.6)

This document specifies an interface to support the multi-image parallelism features of Fortran, named the Parallel Runtime Interface for Fortran (PRIF). PRIF is a solution in which a runtime library is primarily responsible for implementing coarray allocation, deallocation and accesses, image synchronization, atomic operations, events, teams and collective subroutines. The Fortran compiler is responsible for transforming the invocation of Fortran-level multi-image parallelism features into procedure calls to the necessary PRIF subroutines. The interface is designed for portability across shared- and distributed-memory machines, different operating systems, and multiple architectures. Implementations of this interface are intended as an augmentation for the compiler's own runtime library. With an implementation-agnostic interface, alternative parallel runtime libraries may be developed that support the same interface. One benefit of this approach is the ability to vary the communication substrate. A central aim of this document is to define a parallel runtime interface in standard Fortran syntax, which enables us to leverage Fortran to succinctly express various properties of the procedure interfaces, including argument attributes.

97 MATHEMATICS AND COMPUTING↗

Experiences with SYCL on AMD GPUs with Kokkos

With the recent diversification of the hardware landscape in the high-performance computing (HPC) community, performance-portability solutions are becoming more and more important. One of the most popular choices is Kokkos, which recently became a Linux Foundation project. Most of its development is supported by the US Department of Energy and the French Alternative Energies and Atomic Energy Commission. Kokkos is implemented as a C++ library with multiple backends to support CPUs as well as various GPU architectures. These backends include OpenMP, CUDA, HIP, and also SCYL. This approach enables users to leverage the preferred vendor toolchain for the respective platform (e.g. CUDA, ROCm, OneAPI). The SYCL backend is used to target Intel GPUs, in particular to support the Aurora exascale supercomputer. However, SYCL itself also offers a large degree of portability, and in fact Kokkos’ CI for SYCL has been running on NVIDIA hardware due to a lack of access to Intel GPUs. In this report, we describe our experience with using Kokkos SYCL backend on AMD GPUs targeting the Frontier supercomputer at Oak Ridge National Laboratory. The two major SYCL implementations are DPC++ and AdaptiveCpp. While the Kokkos SYCL backend has been implemented using the former, the latter was the first implementation to target AMD GPUs. We will discuss the experience with both of these SYCL implementations in terms of functionality and performance. Using Kokkos to evaluate SYCL toolchains has a number of benefits. Kokkos’ use of SYCL is fairly complex, exercising features such as graphs, relocatable device functions, atomics – including for non-arithmetic types, as well as pinned and page migratable memory allocations. Kokkos also needs to implement capabilities such as Kokkos’ hierarchical parallelism that are not a straight-forward mapping to SYCL capabilities. Furthermore, a large number of libraries and applications that represent diverse use cases are implemented in Kokkos, providing readily available test cases for a toolchain evaluation. Preliminary results show that support for AMD GPUs in DPC++ is much less mature than for NVIDIA GPUs or Intel GPUs. While the situation has improved significantly over the last year, we still encounter many runtime failures, dispatching problems, and code generation issues. With AdaptiveCpp the challenges arise even earlier in the evaluation process. Since Kokkos’ SYCL implementation is largely focused on supporting Intel GPUs, we opted to leverage SYCL extensions which are available in DPC++ but not in AdaptiveCpp. Furthermore, AdaptiveCpp appears to be less conformant with the SYCL2020 standard which Kokkos relies on. In some cases, we are able to work around the lack of feature support, in other cases we have to disable certain Kokkos capabilities to evaluate the toolchain. Our evaluation will leverage Kokkos’ unit tests to establish basic functionality and feature completeness. We then use simple benchmarks for components of a CG implementation as a measure of usability and performance of the SYCL toolchains.

97 MATHEMATICS AND COMPUTING↗

Scalable Bayesian optimization with randomized prior networks

Several fundamental problems in science and engineering consist of global optimization tasks involving unknown high-dimensional (black-box) functions that map a set of controllable variables to the outcomes of an expensive experiment. Bayesian Optimization (BO) techniques are known to be effective in tackling global optimization problems using a relatively small number objective function evaluations, but their performance suffers when dealing with high-dimensional outputs. To overcome the major challenge of dimensionality, here we propose a deep learning framework for BO and sequential decision making based on bootstrapped ensembles of neural architectures with randomized priors. Using appropriate architecture choices, we show that the proposed framework can approximate functional relationships between design variables and quantities of interest, even in cases where the latter take values in high-dimensional vector spaces or even infinite-dimensional function spaces. In the context of BO, we augmented the proposed probabilistic surrogates with re-parameterized Monte Carlo approximations of multiple-point (parallel) acquisition functions, as well as methodological extensions for accommodating black-box constraints and multi-fidelity information sources. We test the proposed framework against state-of-the-art methods for BO and demonstrate superior performance across several challenging tasks with high-dimensional outputs, including a constrained multi-fidelity optimization task involving shape optimization of rotor blades in turbo-machinery.

97 MATHEMATICS AND COMPUTING↗

Exascale models of stellar explosions: Quintessential multi-physics simulation

The ExaStar project aims to deliver an efficient, versatile, and portable software ecosystem for multi-physics astrophysics simulations run on exascale machines. The code suite is a component-based multi-physics toolkit, built on the capabilities of current simulation codes (in particular Flash-X and Castro), and based on the massively parallel adaptive mesh refinement framework AMReX. It includes modules for hydrodynamics, advanced radiation transport, thermonuclear kinetics, and nuclear microphysics. The code will reach exascale efficiency by building upon current multi- and many-core packages integrated into an orchestration system that uses a combination of configuration tools, code translators, and a domain-specific asynchronous runtime to manage performance across a range of platform architectures. The target science includes multi-physics simulations of astrophysical explosions (such as supernovae and neutron star mergers) to understand the cosmic origin of the elements and the fundamental physics of matter and neutrinos under extreme conditions.

79 ASTRONOMY AND ASTROPHYSICS↗

Observability of fidelity decay at the Lyapunov rate in few-qubit quantum simulations

In certain regimes, the fidelity of quantum states will decay at a rate set by the classical Lyapunov exponent. This serves both as one of the most important examples of the quantum-classical correspondence principle and as an accurate test for the presence of chaos. While detecting this phenomenon is one of the first useful calculations that noisy quantum computers without error correction can perform, a thorough study of the quantum sawtooth map reveals that observing the Lyapunov regime is just beyond the reach of present-day devices. We prove that there are three bounds on the ability of any device to observe the Lyapunov regime and give the first quantitatively accurate description of these bounds: (1) the Fermi golden rule decay rate must be larger than the Lyapunov rate, (2) the quantum dynamics must be diffusive rather than localized, and (3) the initial decay rate must be slow enough for Lyapunov decay to be observable. This last bound, which has not been recognized previously, places a limit on the maximum amount of noise that can be tolerated. The theory implies that an absolute minimum of 6 qubits is required. Recent experiments on IBM-Q and IonQ imply that some combination of a noise reduction by up to 100x per gate and large increases in connectivity and gate parallelization are also necessary. Finally, scaling arguments are given that quantify the ability of future devices to observe the Lyapunov regime based on trade-offs between hardware architecture and performance.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Neural Networks for Nuclear Reactions in MAESTROeX

We demonstrate the use of neural networks to accelerate the reaction steps in the MAESTROeX stellar hydrodynamics code. A traditional MAESTROeX simulation uses a stiff ODE integrator for the reactions; here, we employ a ResNet architecture and describe details relating to the architecture, training, and validation of our networks. Our customized approach includes options for the form of the loss functions, a demonstration that the use of parallel neural networks leads to increased accuracy, and a description of a perturbational approach in the training step that robustifies the model. We test our approach on millimeter-scale flames using a single-step, 3-isotope network describing the first stages of carbon fusion occurring in Type Ia supernovae. We train the neural networks using simulation data from a standard MAESTROeX simulation, and show that the resulting model can be effectively applied to different flame configurations. This work lays the groundwork for more complex networks, and iterative time-integration strategies that can leverage the efficiency of the neural networks.

79 ASTRONOMY AND ASTROPHYSICS↗

A High Performance Sparse Tensor Algebra Compiler in MLIR

Sparse tensor algebra is widely used in many applications, including scientific computing, machine learning, and data analytics. The performance of sparse tensor algebra kernels strongly depends on the intrinsic characteristics of the input tensors, hence many storage formats are designed for tensors to achieve optimal performance for particular applications/architectures, which makes it challenging to implement and optimize every tensor operation of interest on a given architecture. We propose a tensor algebra domain-specific language (DSL) and compiler framework to automatically generate kernels for mixed sparse-dense tensor algebra operations. The proposed DSL provides high-level programming abstractions that resemble the familiar Einstein notation to represent tensor algebra operations. The compiler introduces a new Sparse Tensor Algebra dialect built on top of LLVM's extensible MLIR compiler infrastructure for efficient code generation while covering a wide range of tensor storage formats. Our compiler also leverages input-dependent code optimization to enhance data locality for better performance. Our results show that the performance of automatically generated kernels outperforms the state-of-the-art sparse tensor algebra compiler, with up to 20.92x, 6.39x, and 13.9x performance improvement over state-of-the-art tensor algebra compilers, for parallel SpMV, SpMM, and TTM, respectively.

Tian, Ruiqin↗

A two-level GPU-accelerated incomplete LU preconditioner for general sparse linear systems

This paper presents a parallel preconditioning approach based on incomplete LU (ILU) factorizations in the framework of Domain Decomposition (DD) for general sparse linear systems. We focus on distributed memory parallel architectures, specifically, those that are equipped with graphic processing units (GPUs). In addition to block-Jacobi, we present general purpose two-level ILU Schur complement-based approaches, where different strategies are presented to solve the coarse-level reduced system. These strategies are combined with modified ILU methods in the construction of the coarse-level operator, in order to effectively remove smooth errors by targeting an algebraically smooth vector. We leverage available GPU-based sparse matrix kernels to accelerate the setup and the solve phases of the proposed ILU preconditioner. We evaluate the efficiency of the proposed methods as a smoother for algebraic multigrid (AMG) and as a preconditioner for Krylov subspace methods on challenging anisotropic diffusion problems and a collection of general sparse matrices.

97 MATHEMATICS AND COMPUTING↗

ICED: An Integrated CGRA Framework Enabling DFVS-Aware Acceleration

oarse-grained reconfigurable arrays (CGRAs) are a promising solution to enable energy-efficient acceleration of applications from different domains. By leveraging reconfiguration at the functional level, they can adapt to significantly different computational patterns. Existing CGRA mapping approaches extract instruction-level parallelism, exploit loop-pipelining opportunities, guarantee the data dependency, and target high throughput of a given loop. However, the recurrence data-dependency in the DFG and the mismatch between required and available computing/communication resources complicate the mapping, and might lead to significant unbalances in the utilization of the CGRA's tiles. This results in wasted power for tiles with low utilization. Applying dynamic voltage and frequency scaling (DVFS) can potentially solve this challenge and improve energy efficiency by adjusting voltage and frequency of different tiles independently. CGRAs have also been successful in accelerating data-dependent streaming applications. However, in these applications, the execution time of each kernel in the pipeline might dynamically vary depending on the characteristics of the input. This also leads to under-utilization of resources for the dynamically changing kernels that do not limit the application throughput. DVFS can also improve energy efficiency for these applications by dynamically changing the voltage and frequency levels of tiles that host non performance-constraining kernels. This paper proposes ICEDTEA -- an integrated DVFS-aware framework to map applications on CGRAs that support power islands. ICEDTEA proposes a CGRA architecture supporting DVFS islands at varying granularity (from a single tile to a group of tiles) and the related DVFS-aware compilation and mapping toolchain. ICEDTEA is the first work that introduces DVFS support for spatio-temporal CGRAs at power-island levels. The experimental evaluation shows that ICEDTEA improves average utilization by 2.3$\times$ and energy-efficiency by 1.32$\times$ over a conventional CGRA. With streaming applications, ICEDTEA improves energy efficiency by 1.12$\times$ over a state-of-the-art CGRA that introduces partial dynamic reconfiguration to adapt to variations in kernels' throughput.

Tan, Cheng↗