Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “hardware accelerators”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

IRIS: A Performance-Portable Framework for Cross-Platform Heterogeneous Computing

From edge to exascale, computer architectures are becoming more heterogeneous and complex. The systems typically have fat nodes, with multicore CPUs and multiple hardware accelerators such as GPUs, FPGAs, and DSPs. This complexity is causing a crisis in programming systems and performance portability. Several programming systems are working to address these challenges, but the increasing architectural diversity is forcing software stacks and applications to be specialized for each architecture. As we show, all of these approaches critically depend on their software framework for discovery, execution, scheduling, and data orchestration. To address this challenge, we believe that a more agile and proactive software framework is essential to increase performance portability and improve user productivity. To this end, we have designed and implemented IRIS: a performance-portable framework for cross-platform heterogeneous computing. IRIS can discover available resources, manage multiple diverse programming platforms (e.g., CUDA, Hexagon, HIP, Level Zero, OpenCL, OpenMP) simultaneously in the same execution, respect data dependencies, orchestrate data movement proactively, and provide for user-configurable scheduling. To simplify data movement, IRIS introduces a shared virtual device memory with relaxed consistency among different heterogeneous devices. IRIS also adds an automatic kernel workload partitioning technique using the polyhedral model so that it can resize kernels for a wide range of devices. Our evaluation on three architectures, ranging from Qualcomm Snapdragon to a Summit supercomputer node, shows that IRIS improves portability across a wide range of diverse heterogeneous architectures with negligible overhead.

97 MATHEMATICS AND COMPUTING↗

Toward Large-Scale Image Segmentation on Summit

Semantic segmentation of images is an important computer vision task that emerges in a variety of application domains such as medical imaging, robotic vision and autonomous vehicles to name a few. While these domain-specific image analysis tasks involve relatively small image sizes (~ 10 2 × 10 2 ), there are many applications that need to train machine learning models on image data with extents that are orders of magnitude larger (~10 4 × 10 4 ). Training deep neural network (DNN) models on large extent images is extremely memory-intensive and often exceeds the memory limitations of a single graphical processing unit, a hardware accelerator of choice for computer vision workloads. Here, an efficient, sample parallel approach to train U-Net models on large extent image data sets is presented. Its advantages and limitations are analyzed and near-linear strong-scaling speedup demonstrated on 256 nodes (1536 GPUs) of the Summit supercomputer. Using a single node of the Summit supercomputer, an early evaluation of a recently released model parallel framework called GPipe is demonstrated to deliver ~ 2X speedup in executing a U-Net model with an order of magnitude larger number of trainable parameters than reported before. Performance bottlenecks for pipelined training of U-Net models are identified and mitigation strategies to improve the speedups are discussed. Together, these results open up the possibility of combining both approaches into a unified scalable pipelined and data parallel algorithm to efficiently train U-Net models with very large receptive fields on data sets of ultra-large extent images.

Seal, Sudip↗

Unified many-worlds browsing of arbitrary physics-based animations

Manually tuning physics-based animation parameters to explore a simulation outcome space or achieve desired motion outcomes can be notoriously tedious. This problem has motivated many sophisticated and specialized optimization-based methods for fine-grained (keyframe) control, each of which are typically limited to specific animation phenomena, usually complicated, and, unfortunately, not widely used. In this paper, we propose Unified Many-Worlds Browsing (UMWB), a practical method for sample-level control and exploration of physics-based animations. Our approach supports browsing of large simulation ensembles of arbitrary animation phenomena by using a unified volumetric WORLDPACK representation based on spatiotemporally compressed voxel data associated with geometric occupancy and other low-fidelity animation state. Beyond memory reduction, the WORLDPACK representation also enables unified query support for interactive browsing: it provides fast evaluation of approximate spatiotemporal queries, such as occupancy tests that find ensemble samples ("worlds") where material is either IN or NOT IN a user-specified spacetime region. WORLDPACKS also support real-time hardware-accelerated voxel rendering by exploiting the spatially hierarchical and temporal RLE raster data structure. Our UMWB implementation supports interactive browsing (and offline refinement) of ensembles containing thousands of simulation samples, and fast spatiotemporal queries and ranking. We show UMWB results using a wide variety of physics-based animation phenomena---not just JELL-O ® .

Computer Science↗

MLIR loop optimizations for High-Level Synthesis: a case study

High-Level Synthesis (HLS) tools simplify the design of hardware accelerators by automatically generating Verilog/VHDL code starting from a general purpose software programming language. They include a wide range of optimization techniques in the process, most of them performed on a low-level intermediate representation (IR) of the code. Introducing optimizations on a higher level of abstraction could significantly contribute to the automated design process results; for example, polyhedral techniques for the manipulation of loops could have a significant impact on the generated accelerators when applied on a specialized IR. We use loop pipelining as a case study to explore the introduction of compiler-based transformations on top of an existing HLS process. We leverage the Multi-Level Intermediate Representation (MLIR) framework and an external scheduler to implement the required transformations, and couple them with existing HLS tools to evaluate the improvements that loop pipelining brings to the performance of generated accelerators. The proposed approach can be integrated with other high-level transformations on the MLIR representation, combining different techniques to obtain pre-optimized inputs for HLS that do not have to rely on a specific backend tool.

Curzel, Serena↗

BitGNN: Unlocking the Performance Potential of Binary Graph Neural Networks on GPUs

Graph Neural Networks (GNNs) have shown compelling results in many graph-based learning tasks. They are, however, time-consuming. Recent work has shown a promising direction in improving GNN speed and shrinking the size — network binarization, which binarizes network values and operations. Prior work, however, mainly focused on algorithm designs, leaving it open on how to fully materialize the performance potential. This work fills the gap by proposing techniques to best map binary GNNs and their computations to fit the nature of bit manipulations, optimizations and algorithms to maximize BSpMM kernel efficiency, and solutions to other factors influencing the end-to-end time on GPUs. Results on real-world graphs show that the proposed techniques outperform state of-the-art binary GNN implementations by 21-67× with little accuracy loss.

Chen, Jou-An↗

High-Level Synthesis of Irregular Applications: A Case Study on Influence Maximization

The Influence Maximization problem is the problem of identifying a small cohort of actors from a broader population that, when initially activated in a diffusion process, are expected to result in a large number of activations in the population. While the problem is known to be NP-hard, several approximation algorithms have been devised by leveraging its submodular structure. While these algorithms are theoretically efficient, they are computationally very expensive in practice. This work advances the current state-of-the-art parallelization scheme for the IMM algorithm by devising the adoption of custom hardware accelerators implemented on FPGAs by leveraging High Level Synthesis from OpenCL. We study the performance of our proposed approach by exploring optimizations tailored at improving the parallel efficiency of the accelerators and highlight their effects and limitations in accelerating complex graph analytic applications. Our experimental evaluation shows that FPGA acceleration can improve the performance of the LT diffusion model up to 1.72x for the entire application and up to 2.90x for its most important kernel with respect to a CPU only parallel execution. The FPGA acceleration of the LT model shows also a 1.54x reduction in energy consumption when compared to a parallel CPU only run.

Neff, Reece W.↗

Tensorized Interior Radiative Heat Transfer for a Scalable and Calibrated Building Energy Simulator

Building energy simulation is a critical tool for developing and testing advanced control strategies, such as Reinforcement Learning (RL), to provide demand flexibility and affordable energy costs. The recently introduced Smart Buildings Control Suite (sbsim) provides a lightweight, scalable, and data-calibrated simulation environment based on a 2D finite-difference model. However, the initial model primarily focused on conductive and convective heat transfer, neglecting the significant impact of long-wave radiative heat exchange between interior surfaces. This paper presents a significant extension to the sbsim framework by incorporating a physically-grounded model for interior radiative heat transfer. Our primary contribution is the development and integration of a fully tensorized radiative heat transfer module, which preserves the computational efficiency and scalability of the original simulator. This was achieved by developing a pipeline for view factor calculation, including an algorithm to identify directly seeing surfaces within complex floor plans, and formulating the net radiation equations for efficient execution on modern hardware accelerators. We validate the numerical accuracy of our tensorized implementation by comparing its results against a traditional iterative approach, demonstrating identical outcomes. This enhancement increases the physical fidelity of sbsim, enabling more accurate training of RL agents for building energy optimization.

Ham, Sang woo↗

ESnet SmartNIC v1.0

The ESnet SmartNIC is a collection of Verilog based FPGA design software, as well as drivers to interact with the FPGA. It provides an infrastructure framework for FPGA based hardware acceleration of network packet use cases. Different research and production applications can be easily written for the SmartNIC. Its advantage is that it reduces the development time for new applications by providing a pre-existing shell library for common functions.

Mah, Bruce↗

GenSLMs: Genome-scale language models reveal SARS-CoV-2 evolutionary dynamics

We seek to transform how new and emergent variants of pandemic-causing viruses, specifically SARS-CoV-2, are identified and classified. By adapting large language models (LLMs) for genomic data, we build genome-scale language models (GenSLMs) which can learn the evolutionary landscape of SARS-CoV-2 genomes. By pre-training on over 110 million prokaryotic gene sequences and fine-tuning a SARS-CoV-2-specific model on 1.5 million genomes, we show that GenSLMs can accurately and rapidly identify variants of concern. Thus, to our knowledge, GenSLMs represents one of the first whole-genome scale foundation models which can generalize to other prediction tasks. We demonstrate scaling of GenSLMs on GPU-based supercomputers and AI-hardware accelerators utilizing 1.63 Zettaflops in training runs with a sustained performance of 121 PFLOPS in mixed precision and peak of 850 PFLOPS. We present initial scientific insights from examining GenSLMs in tracking evolutionary dynamics of SARS-CoV-2, paving the path to realizing this on large biological data.

Zvyagin, Maxim↗

Development of message passing-based graph convolutional networks for classifying cancer pathology reports

Abstract Background Applying graph convolutional networks (GCN) to the classification of free-form natural language texts leveraged by graph-of-words features (TextGCN) was studied and confirmed to be an effective means of describing complex natural language texts. However, the text classification models based on the TextGCN possess weaknesses in terms of memory consumption and model dissemination and distribution. In this paper, we present a fast message passing network (FastMPN), implementing a GCN with message passing architecture that provides versatility and flexibility by allowing trainable node embedding and edge weights, helping the GCN model find the better solution. We applied the FastMPN model to the task of clinical information extraction from cancer pathology reports, extracting the following six properties: main site, subsite, laterality, histology, behavior, and grade. Results We evaluated the clinical task performance of the FastMPN models in terms of micro- and macro-averaged F1 scores. A comparison was performed with the multi-task convolutional neural network (MT-CNN) model. Results show that the FastMPN model is equivalent to or better than the MT-CNN. Conclusions Our implementation revealed that our FastMPN model, which is based on the PyTorch platform, can train a large corpus (667,290 training samples) with 202,373 unique words in less than 3 minutes per epoch using one NVIDIA V100 hardware accelerator. Our experiments demonstrated that using this implementation, the clinical task performance scores of information extraction related to tumors from cancer pathology reports were highly competitive.

59 BASIC BIOLOGICAL SCIENCES↗

TRANSFER LEARNING FOR FIELD EMISSION MITIGATION IN CEBAF SRF CAVITIES

The Continuous Electron Beam Accelerator Facility (CEBAF) at Jefferson Lab operates hundreds of super-conducting radio frequency (SRF) cavities in its two linear accelerators (linacs). Field emission (FE) is an ongoing operational challenge in higher gradient SRF cavities. FE generates high levels of neutron and gamma radiation leading to damaged accelerator hardware and a radiation hazard environment. During machine development periods, we performed gradient scans to record data capturing the relationship between cavity gradients and radiation levels measured throughout the linacs. However, the field emission environment at CEBAF varies considerably over time as the configuration of the radio frequency (RF) gradients changes and due to the changing behaviour of field emitters. An artificial intelligence/machine learning (AI/ML) approach with transfer learning could be a valuable tool to mitigate FE and lower the radiation levels. In this work, we mainly focus on leveraging the RF trip data gathered during CEBAF operations. We develop a transfer learning-based surrogate model for radiation detector readings given RF cavity gradients to track the CEBAF?s changing configuration and environment. Then, we could use the developed model as an optimization process for redistributing the RF gradients within a linac to minimize radiation levels.

Ahammed, K.↗

Field Emission Mitigation in CEBAF SRF Cavities Using Deep Learning

The Continuous Electron Beam Accelerator Facility (CEBAF) operates hundreds of superconducting radio frequency (SRF) cavities in its two main linear accelerators. Field emission can occur when the cavities are set to high operating RF gradients and is an ongoing operational challenge. This is especially true in newer, higher gradient SRF cavities. Field emission results in damage to accelerator hardware, generates high levels of neutron and gamma radiation, and has deleterious effects on CEBAF operations. So, field emission reduction is imperative for the reliable, high gradient operation of CEBAF that is required by experimenters. Here we explore the use of deep learning architectures via multilayer perceptron to simultaneously model radiation measurements at multiple detectors in response to arbitrary gradient distributions. These models are trained on collected data and could be used to minimize the radiation production through gradient redistribution. This work builds on previous efforts in developing machine learning (ML) models, and is able to produce similar model performance as our previous ML model without requiring knowledge of the field emission onset for each cavity.

Ahammed, K.↗

jaxhps: An elliptic PDE solver built with machine learning in mind

Elliptic partial differential equations (PDEs) can model many physical phenomena, such as electrostatics, acoustics, wave propagation, and diffusion. In scientific machine learning settings, a high-throughput PDE solver may be required to generate a training dataset, run in the inner loop of an iterative algorithm, or interface directly with a deep neural network. To provide value to machine learning users, such a PDE solver must be compatible with standard automatic differentiation frameworks, scale efficiently when run on graphics processing units (GPUs), and maintain high accuracy for a large range of input parameters. We have designed the jaxhps package with these use-cases in mind by implementing a highly efficient and accurate solver for elliptic problems with native hardware acceleration and automatic differentiation support.

97 MATHEMATICS AND COMPUTING↗

PixelStorm: A Remote Display for Remote Sensing Ground Stations

PixelStorm is a software application for displaying native high-performance applications from remote cloud environments. It is tailored for remote sensing missions that require high framerates, high resolutions, and minimal loss of quality. PixelStorm utilizes hardware-accelerated video compression on graphics processing units and a Sandia-developed streaming network protocol. Using our architecture, we can demonstrate interactive native applications running across two 4K monitors at 60 frames per second while maintaining the visual fidelity required by our missions. This technology allows for the migration of mission critical desktop applications to cloud environments.

97 MATHEMATICS AND COMPUTING↗

ExaSGD: 2021 Kernel Thrust Activities

The Kernel Thrust milestone ADSE22-214 covers the development of device-capable optimization algorithms and solvers technologies required by the ExaSGD project’s software stack in order to solve security-constrained alternating current optimal power flow (SC-ACOPF) problems on emerging exascale architectures. To this extent, in FY21 the main objective of the Kernel Thrust was (i) provide robust optimization solver(s) that run efficiently on hardware accelerator devices (i.e., NVIDIA and AMD GPUs) to perform intra-node computations and (ii) provide coarse-grain parallel optimization capabilities that exploit the decomposition opportunities present in the SC-ACOPF challenge problems to provide exascale-capable solvers.

97 MATHEMATICS AND COMPUTING↗

Field Emission Mitigation in CEBAF SRF Cavities Using Deep Learning

The Continuous Electron Beam Accelerator Facility (CEBAF) operates hundreds of superconducting radio frequency (SRF) cavities in its two main linear accelerators. Field emission can occur when the cavities are set to high operating RF gradients and is an ongoing operational challenge. This is especially true in newer, higher gradient SRF cavities. Field emission results in damage to accelerator hardware, generates high levels of neutron and gamma radiation, and has deleterious effects on CEBAF operations. So, field emission reduction is imperative for the reliable, high gradient operation of CEBAF that is required by experimenters. Here we explore the use of deep learning architectures via multilayer perceptron to simultaneously model radiation measurements at multiple detectors in response to arbitrary gradient distributions. These models are trained on collected data and could be used to minimize the radiation production through gradient redistribution. This work builds on previous efforts in developing machine learning (ML) models, and is able to produce similar model performance as our previous ML model without requiring knowledge of the field emission onset for each cavity.

Ahammed, K.↗

TRANSFER LEARNING FOR FIELD EMISSION MITIGATION IN CEBAF SRF CAVITIES

The Continuous Electron Beam Accelerator Facility (CEBAF) at Jefferson Lab operates hundreds of super-conducting radio frequency (SRF) cavities in its two linear accelerators (linacs). Field emission (FE) is an ongoing operational challenge in higher gradient SRF cavities. FE generates high levels of neutron and gamma radiation leading to damaged accelerator hardware and a radiation hazard environment. During machine development periods, we performed gradient scans to record data capturing the relationship between cavity gradients and radiation levels measured throughout the linacs. However, the field emission environment at CEBAF varies considerably over time as the configuration of the radio frequency (RF) gradients changes and due to the changing behaviour of field emitters. An artificial intelligence/machine learning (AI/ML) approach with transfer learning could be a valuable tool to mitigate FE and lower the radiation levels. In this work, we mainly focus on leveraging the RF trip data gathered during CEBAF operations. We develop a transfer learning-based surrogate model for radiation detector readings given RF cavity gradients to track the CEBAF?s changing configuration and environment. Then, we could use the developed model as an optimization process for redistributing the RF gradients within a linac to minimize radiation levels.

Ahammed, K.↗