Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “spatial accelerator”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Evaluating Spatial Accelerator Architectures with Tiled Matrix-Matrix Multiplication

There is a growing interest in custom spatial accelerators for machine learning applications. These accelerators employ a spatial array of processing elements (PEs) interacting via custom buffer hierarchies and networks-on-chip. The efficiency of these accelerators comes from employing optimized dataflow (i.e., spatial/temporal partitioning of data across the PEs and fine-grained scheduling) strategies to optimize data reuse. The focus of this work is to evaluate these accelerator architectures using a tiled general matrix-matrix multiplication (GEMM) kernel. To do so, we develop a framework that finds optimized mappings (dataflow and tile sizes) for a tiled GEMM for a given spatial accelerator and workload combination, leveraging an analytical cost model for runtime and energy. Our evaluations over five spatial accelerators demonstrate that the tiled GEMM mappings systematically generated by our framework achieve high performance on various GEMM workloads and accelerators.

43 PARTICLE ACCELERATORS↗

Evaluating Spatial Accelerator Architectures with Tiled Matrix-Matrix Multiplication.

There is a growing interest in custom spatial accelerators for machine learning applications. These accelerators employ a spatial array of processing elements (PEs) interacting via custom buffer hierarchies and networks-on-chip. The efficiency of these accelerators comes from employing optimized dataflow (i.e., spatial/temporal partitioning of data across the PEs and fine-grained scheduling) strategies to optimize data reuse. The focus of this work is to evaluate these accelerator architectures using a tiled general matrix-matrix multiplication (GEMM) kernel. To do so, we develop a framework that finds optimized mappings (dataflow and tile sizes) for a tiled GEMM for a given spatial accelerator and workload combination, leveraging an analytical cost model for runtime and energy. Finally, our evaluations over five spatial accelerators demonstrate that the tiled GEMM mappings systematically generated by our framework achieve high performance on various GEMM workloads and accelerators.

42 ENGINEERING↗

Understanding the Design Space of Sparse/Dense Multiphase Dataflows for Mapping Graph Neural Networks on Spatial Accelerators

Graph Neural Networks (GNNs) have garnered a lot of recent interest because of their success in learning representations from graph-structured data across several critical applications in cloud and HPC. Owing to their unique compute and memory characteristics that come from an interplay between dense and sparse phases of computations, the emergence of reconfigurable dataflow (aka spatial) accelerators offers promise for acceleration by mapping optimized dataflows (i.e., computation order and parallelism) for both phases. The goal of this work is to characterize and understand the design-space of dataflow choices for running GNNs on spatial accelerators in order for the compilers to optimize the dataflow based on the workload. Specifically, we propose a taxonomy to describe all possible choices for mapping the dense and sparse phases of GNNs spatially and temporally over a spatial accelerator, capturing both the intra-phase dataflow and the inter-phase (pipelined) dataflow. Using this taxonomy, we do deep-dives into the cost and benefits of several dataflows and perform case studies on implications of hardware parameters for dataflows and value of flexibility to support pipelined execution.

97 MATHEMATICS AND COMPUTING↗

Union: A Unified HW-SW Co-Design Ecosystem in MLIR for Evaluating Tensor Operationson Spatial Accelerators

To meet the extreme compute demands for deep learning across commercial and scientific applications, dataflow accelerators are becoming increasingly popular. While these“domain-specific” accelerators are not fully programmable like CPUs and GPUs, they retain varying levels of flexibility with respect to data orchestration, i.e., dataflow and tiling optimizations to enhance efficiency. There are several challenges when designing new algorithms and mapping approaches to execute the algorithms for a target problem on new hardware. Previous works have addressed these challenges individually. To address this challenge as a whole, in this work, we present an HW-SW co-design ecosystem for spatial accelerators called Union within the popular MLIR compiler infrastructure. Our framework allows exploring different algorithms and their mappings on several accelerator cost models. Union also includes a plug-and-play library of accelerator cost models and mappers which can easily be extended. The algorithms and accelerator cost models are connected via a novel mapping abstraction that captures the map space of spatial accelerators which can be systematically pruned based on constraints from the hardware, workload, and mapper. We demonstrate the value of Union for the community with several case studies which examine offloading different tensor operations (CONV/GEMM/Tensor Contraction) on diverse accelerator architectures using different mapping schemes.

Jeong, Geonhwa↗

Spatially Accelerated Winding Numbers for Curved Geometry

The generalized winding number (GWN) is a scalar field that supports robust containment queries on curved geometry, including non-watertight, overlapping, and nested boundary representations. While queries can be easily parallelized over samples, direct evaluation on parametric curves and surfaces remains costly for large and complex models. Fast, state-of-the-art GWN approaches leverage a spatial index to approximate the GWN, typically coupled with a Taylor expansion which approximates the GWN contribution for far clusters of geometric primitives. However, such methods operate only on discrete inputs such as triangle meshes and point clouds, and would introduce containment errors near boundaries if applied to curved input. We extend support for fast GWN evaluation over arbitrary collections of NURBS curves in 2D and trimmed NURBS patches in 3D via a Bounding Volume Hierarchy that stores efficiently precomputed moment data in the hierarchy nodes. When querying the hierarchy, approximations for far clusters are used alongside direct evaluation for nearby NURBS primitives, achieving sub-linear complexity while preserving the geometric features in the vicinity of the query point. Central to our performance improvements is an adaptive subdivision strategy for NURBS primitives during a preprocessing phase, creating better spatial partitions while retaining the same accuracy for containment decisions as a direct evaluation. We demonstrate the performance and accuracy of our approach across a large collection of 2D and 3D datasets.

Computer science↗

Exploring the Use of Novel Spatial Accelerators in Scientific Applications

Driven by the need to find alternative accelerators which can viably replace GPUs in next-generation Supercomputing systems, this paper proposes a methodology to enable agile application/hardware co-design. The application-first methodology provides the ability to come up with design of accelerators while working with real-world workloads, available accelerators, and system software. The iterative design process targets a set of kernels in a workload for performance estimates that can prune the design space for later phases of detailed architectural evaluations. To this effect, in this paper, a novel data-parallel device model is introduced that simulates the latency of performance-sensitive operations in an accelerator including data transfers and kernel computation using multi-core CPUs. The use of off-the-shelf simulators, such as pre-RTL simulator Aladdin or multiple tools available for exploring the design of deep neural network accelerators (e.g., Timeloop) is demonstrated for evaluation of various accelerator designs using applications with realistic inputs. Examples of multiple device configurations that are instantiable in a system are explored to evaluate the performance benefit of deploying novel accelerators. The proposed device is integrated with a programming model and system software to potentially explore the impacts of high-level programming languages/compilers and low-level effects such as task scheduling on multiple accelerators. We analyze our methodology for a set of applications that represent high-performance computing (HPC) and graph analytics. The applications include a computational chemistry kernel realized using tensor contractions, triangle counting, GraphSAGE and Breadth-first Search. These applications include kernels such as dense matrix-dense matrix multiplication, sparse matrix-spare matrix multiplication, and sparse matrix-dense vector multiplication. Our results indicate potential performance benefits and insights for system design by including accelerators that realize these kernels along-side general purpose accelerators.

AI, codesign, Accelerated Computing, Modeling and ↗

Adaptive Spatially Aware I/O for Multiresolution Particle Data Layouts

Large-scale simulations on nonuniform particle distributions that evolve over time are widely used in cosmology, molecular dynamics, and engineering. Such data are often saved in an unstructured format that neither preserves spatial locality nor provides metadata for accelerating spatial or attribute subset queries, leading to poor performance of visualization tasks. Furthermore, the parallel I/O strategy used typically writes a file per process or a single shared file, neither of which is portable or scalable across different HPC systems. We present a portable technique for scalable, spatially aware adaptive aggregation that preserves spatial locality in the output. We evaluate our approach on two supercomputers, Stampede2 and Summit, and demonstrate that it outperforms prior approaches at scale, achieving up to 2.5× faster writes and reads for nonuniform distributions. Furthermore, the layout written by our method is directly suitable for visual analytics, supporting low-latency reads and attribute-based filtering with little overhead.

Usher, Will↗

SHarD: A beam dynamics simulation code for dielectric laser accelerators based on spatial harmonic field expansion

In order to demonstrate acceleration of electrons to relativistic scales by an on chip dielectric laser accelerator (DLA), a ponderomotive focusing scheme capable of capturing and transporting electrons through nanometer-scale apertures over extended interaction lengths has been proposed. Here we present a Matlab-based numerical code (SHarD) utilizing a spatial harmonic expansion of the fields within the dielectric structure to simulate the evolution of the beam phase space distribution in this scheme. The code can be used to optimize key-parameters for the accelerator performance such as the final energy, transverse spot size evolution and total number of electrons accelerated through currently fabricated structures. Eventually, the simulation model will be applied to inform the phase mask profile to be added to a pulse front tilt drive laser pulse using a liquid crystal mask in the experimental setup being assembled at UCLA Pegasus Laboratory.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

An efficient second-order adaptive procedure for inserting CAD geometries into hexahedral meshes using volume fractions

Here, this paper is concerned with inserting three-dimensional computer-aided design (CAD) geometries into meshes composed of hexahedral elements using a volume fraction representation. An adaptive procedure for doing so is presented. The procedure consists of two steps. The first step performs spatial acceleration using a k-d tree. The second step involves subdividing individual hexahedra in an adaptive mesh refinement (AMR)-like fashion and approximating the CAD geometry linearly (as a plane) at the finest subdivision. The procedure requires only two geometric queries from a CAD kernel: determining whether or not a queried spatial coordinate is inside or outside the CAD geometry and determining the closest point on the CAD geometry’s surface from a given spatial coordinate. We prove that the procedure is second-order accurate for sufficiently smooth geometries and sufficiently refined background meshes. We demonstrate the expected order of accuracy is achieved with several verification tests and illustrate the procedure’s effectiveness for several exemplar CAD geometries.

Adaptive↗

Fragmentation analysis of a bar with the Lip-field approach

The Lip-field approach was introduced in Moës and Chevaugeon (2021) as a new way to regularize softening material models. It was tested in 1D quasistatic in Moës and Chevaugeon (2021) and 2D quasistatic in Chevaugeon and Moës (2021): this paper extends it to 1D dynamics, on the challenging problem of dynamic fragmentation. The Lip-field approach formulates the mechanical problem to be solved as an optimization problem, where the incremental potential to be minimized is the non-regularized one. Spurious localization is prevented by imposing a Lipschitz constraint on the damage field. Here, the displacement and damage field at each time step are obtained by a staggered algorithm, that is the displacement field is computed for a fixed damage field, then the damage field is computed for a fixed displacement field. Indeed, these two problems are convex, which is not the case of the global problem where the displacement and damage fields are sought at the same time. The incremental potential is obtained by equivalence with a cohesive zone model, which makes material parameters calibration simple. A non-regularized local damage equivalent to a cohesive zone model is also proposed. It is used as a reference for the Lip-field approach, without the need to implement displacement jumps. These approaches are applied to the brittle fragmentation of a 1D bar with randomly perturbed material properties to accelerate spatial convergence. Both explicit and implicit dynamic implementations are compared. Favorable comparison to several analytical, numerical and experimental references serves to validate the modeling approach.

36 MATERIALS SCIENCE↗

A machine learning approach for particle accelerator errant beam prediction using spatial phase deviation

Particle accelerators are extremely complex systems that are expected to operate on high availability. Predicting impending failures only by utilizing data collected from diagnostic equipment already on board can help operators to avoid installing expensive sensors, unscheduled downtime and associated costs. For this purpose we explore the predictive power of Machine Learning algorithms to detect faulty beams prior to the failure. In this study, we propose a Machine Learning approach to model mapping from a pair of sensors located across the accelerator. While the model is trained to represent normal operation, we evaluate the predictive performance on known faulty beam pulses. We also investigate the model performance on unseen data through k-fold cross-validation. Then we recap the analysis with a neural architecture search and hyperparameter optimization study to fine tune our initial model. In conclusion, this paper will also introduce a sustainable framework that can standardize Machine Learning workflow applied to particle accelerators.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

A Spatially Resolved Evaluation of Accelerated Environmental Aging on Emerging Polypropylene-Based Photovoltaic Backsheets Using Raman Spectroscopy

For this work, accelerated aging was used to assess environmental degradation in emerging co-extruded polypropylene (PP)-based backsheets under three different environmental conditions (65°C/20% relative humidity (RH), 75°C/20% RH, and 75°C/50% RH). Although differential scanning calorimetry did not measure crystallinity changes with exposure, spatially resolved Raman spectroscopy identified crystallinity increases in the core layer of aged samples, indicating a heterogeneous postcrystallization process. The Raman results were in agreement with synchrotron-based microfocused wide-angle X-ray scattering measurements. Cross-sectional nanoindentation was used to correlate localized crystallinity shifts with changes in Young's modulus. A similar trend was found where increased modulus was measured in the core layer, supporting the relationship between modulus and crystallinity. Finally, dielectric characterization was used to assess the impact of these material property changes on performance. While changes in the backsheet material properties and dielectric performance were observed with accelerated aging, these shifts generally equilibrated with time, indicating overall stability in response to environmental stressors. Additionally, the identified heterogeneous material property changes indicate that spatially resolved crystallinity measurements may be a valuable early failure indicator to be used in the assessment of PV backsheet long-term durability.

36 MATERIALS SCIENCE↗

Test Results of a High-Gradient 2.856-GHz Negative Harmonic Accelerating Waveguide

We report the high-gradient tests results of a novel traveling wave accelerating structure for β=0.3 based on a novel approach of operating at the first negative spatial harmonic. Accelerating gradients of 50 MV/m and peak electric fields of 160 MV/m were achieved in a single structure consisting of 15 coupled cells during tests at the advanced photon source. This work was performed by RadiaBeam, in collaboration with the Argonne National Laboratory, as a part of a Research and Development Program for the development of an ultrahigh-gradient linear accelerator, the Advanced Compact Carbon Ion Linac, for hadron therapy.

42 ENGINEERING↗