Engineering Papers⌕ Search

Engineering topics

Cabrera, Anthony

Publications and source records attributed to Cabrera, Anthony.

Oak Ridge National Laboratory's Strategic Research and Development Insights for Digital Twins

Oak Ridge National Laboratory (ORNL) is pleased to provide our response to the NITRD RFI on Digital Twins Research and Development. Digital twins are virtual representations of physical systems, leveraging real-time data to simulate and predict behaviors. ORNL is advancing digital twin technology across various disciplines, including neutron scattering, networking, science ecosystems, supercomputing, secure facilities, mobility technologies, materials design and discovery, power systems, fusion reactors, biological sciences, and earth observation. These efforts aim to enhance scientific research, operational efficiency, and decision-making processes. ORNL facilities, such as the High Flux Isotope Reactor (HFIR), Grid-C, Spallation Neutron Source (SNS), and Oak Ridge Leadership Computing Facility (OLCF), provide the infrastructure to develop and demonstrate these digital twin technologies. In this document, we lay out key challenges, research gaps, and future opportunities based on our experience with digital twins that aim to serve as useful contributions towards a National Digital Twins R&D Strategic Plan. In the remaining document, we address nine of the thirteen topic areas specified in the RFI.

97 MATHEMATICS AND COMPUTING↗

Errant Beam Detection Using the AMD Versal ACAP and Vitis AI

The prevalence of ML and AI-powered solutions along with the slowing of Moore's Law has given rise to novel hardware platforms aimed at accelerating ML and AI. While programming these hardware platforms can be difficult, particularly for non-hardware experts, hardware vendors provide high-level tooling in an effort to address this difficulty. The Versal ACAP is an SoC designed by AMD that combines CPU cores, FPGA fabric, and a tiled, vector architecture called an AI engine all on the same socket. In an effort to more easily program this heterogeneous system, AMD has provided the Vitis AI development stack. In this work, we leverage Vitis AI to program a Versal ACAP to perform errant beam detection in the Spallation Neutron Source at Oak Ridge National Laboratory. Our initial work shows that after quantization and compilation of the model for the Versal ACAP, the classification accuracy, as measured by the AUC metric, is over 95% accurate while achieving this accuracy in 46 microseconds on average.

Cabrera, Anthony↗

Design and analysis of CXL performance models for tightly-coupled heterogeneous computing

Truly heterogeneous systems enable partitioned workloads to be mapped to the hardware that nets the best performance. However, current practice requires that inter-device communication between different vendors' hardware use host memory as an intermediary step. To date, there are no widely adopted solutions that allow accelerators to directly transfer data. A new cache-coherent protocol, CXL, aims to facilitate easier, fine-grained sharing between accelerators. In this work we analyze existing methods for designing heterogeneous applications that target GPUs and FPGAs working collaboratively, followed by an exploration to show the benefits of a CXL-enabled system. Specifically, we develop a test application that utilizes both an NVIDIA P100 GPU and a Xilinx U250 FPGA to show current communication limitations. From this application, we capture overall execution time and throughput measurements on the FPGA and GPU. We use these measurements as inputs to novel CXL performance models to show that using CXL caching instead of host memory results in a 1.31X speedup, while a more tightly-coupled pipelined implementation using CXL-enabled hardware would result in a speedup of 1.45X.

Cabrera, Anthony↗

Platform Agnostic Streaming Data Application Performance Models

The mapping of computational needs onto execution resources is, by and large, a manual task, and users are frequently guided simply by intuition and past experiences. We present a queueing theory based performance model for streaming data applications that takes steps towards a better understanding of resource mapping decisions, thereby assisting application developers to make good mapping choices. The performance model (and associated cost model) are agnostic to the specific properties of the compute resource and application, simply characterizing them by their achievable data throughput. We illustrate the model with a pair of applications, one chosen from the field of computational biology and the second is a classic machine learning problem.

Faber, Clayton↗

Hardware Evaluation Analytical Modeling and Node Simulation: Benefits of Tighter GPU Integration

In this report, we examine several emerging technologies of interest to the Department of Energy and its computational centers. These include: 1) quantifying the benefit of tighter CPU-GPU integration, 2) quantifying the appropriate CPU core:GPU ratio, 3) quantifying the penalty for CPU-GPU disaggregation, 4) quantifying the benefits of tighter GPU-GPU integration, 5) quantifying the benefits of unified memory, and 6) quantifying the benefits of tighter FPGA-GPU integration.

97 MATHEMATICS AND COMPUTING↗