Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data Transfer”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Impact of the Li 6 asymptotic normalization constant onto α -induced reactions of astrophysical interest

Indirect methods have become the predominant approach in experimental nuclear astrophysics for studying several low-energy nuclear reactions occurring in stars, as direct measurements of many of these relevant reactions are rendered infeasible due to their low reaction probability. Such indirect methods, however, require theoretical input that in turn can have significant poorly quantified uncertainties, which can then be propagated to the reaction rates and have a large effect on our quantitative understanding of stellar evolution and nucleosynthesis processes. Here we present two such examples involving α-induced reactions, 13 C (α,n)⁢ 16 O and 12 C (α,γ)⁢ 16 O, for which the low-energy cross sections have been constrained with ( 6 Li,d) transfer data. In this Letter, we discuss how a first-principle calculation of 6 Li leads to a 21% reduction of the 12 C⁡(α,γ) ⁢ 16 O cross sections with respect to a previous estimation. This calculation further resolves the discrepancy between recent measurements of the 13 C (α,n)⁢ 16 O reaction and points to the need for improved theoretical formulations of nuclear reactions.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Accelerating Scientific Workflows on HPC Platforms with In Situ Processing

Scientific workflows drive most modern large-scale science breakthroughs by allowing scientists to define their computations as a set of jobs executed in a given order based on their data dependencies. Workflow management systems (WMSs) have become key to automating scientific workflows-executing computational jobs and orchestrating data transfers between those jobs running on complex high-performance computing (HPC) platforms. Traditionally, WMSs use files to communicate between jobs: a job writes out files that are read by other jobs. However, HPC machines face a growing gap between their storage and compute capabilities. To address that concern, the scientific community has adopted a new approach called in situ, which bypasses costly parallel filesystem I/O operations with faster in-memory or in-network communications. When using in situ approaches, communication and computations can be interleaved. In this work, we leverage the Decaf in situ dataflow framework to accelerate task-based scientific workflows managed by the Pegasus WMS, by replacing file communications with faster MPI messaging. We propose a new execution engine that uses Decaf to manage communications within a sub-workflow (i.e., set of jobs) to optimize inter-job communications. We consider two workflows in this study: (i) a synthetic workflow that benchmarks and compares file- and MPI-based communication; and (ii) a realistic bioinformatics workflow that computes mu-tational overlaps in the human genome. Experiments show that in situ communication can improve the bioinformatics workflow execution time by 22% to 30% compared with file communication. Our results motivate further opportunities and challenges for bridging traditional WMSs with in situ frameworks.

Decaf↗

Toward Energy-Efficient HPC: Insights from Power Profiling a Cloud-Resolving Earth System Model

Power is a fundamental constraint as supercomputing advances to exascale. Efficient operation within strict power budgets requires application-aware power management based on a detailed understanding of application-level power behavior. This work analyzes the Energy Exascale Earth System Model (E3SM) atmosphere component, SCREAM, on Perlmutter (NERSC) and Frontier (OLCF). We characterize power variation across inputs, concurrency levels, and power caps, evaluate the energy impact of code optimizations, and attribute energy within the code using a newly developed GPU energy model. Results show that SCREAM’s peak power remains stable during its core execution phase and decreases gradually as concurrency increases. Power capping experiments reveal a performance–energy "sweet spot". On Perlmutter, limiting GPU power to 50% of thermal design power (TDP) achieves up to 15% energy savings with a 7% performance penalty. On Frontier, a 40% TDP cap yields up to 10% energy savings with less than 10% performance loss. Code optimizations reduce SCREAM energy by shortening run time without increasing power. Modeling reveals a critical insight: data movement accounts for approximately 70% of SCREAM’s GPU energy. This fundamentally shifts the optimization focus from FLOPS to data transfer reduction for this class of applications, offering the most impactful strategy for improving energy efficiency. This work establishes a foundation for practical, application-aware power management at exascale.

Zhao, Zhengji [Lawrence Berkeley National Laborato↗

Coding the Computing Continuum: Fluid Function Execution in Heterogeneous Computing Environments

Advances in network technologies have greatly decreased barriers to accessing physically distributed computers. This newfound accessibility coincides with increasing hardware specialization, creating exciting new opportunities to dispatch workloads to the best resource for a specific purpose, rather than those that are closest or most easily accessible. We present Delta, a service designed to intelligently schedule function-based workloads across a distributed set of heterogeneous computing resources. Delta implements an extensible architecture in which different predictors and scheduling algorithms can be integrated to provide dynamically evolving estimates of function execution times on different resources-estimates that can be used to determine the most appropriate location for execution. We describe predictors for function runtime, data transfer time, and cold-start resource provisioning and configuration delay; dynamic learning methods that update predictor models over time; and scheduling strategies that take into account both function and endpoint information. We show that these methods can halve workload makespan when compared with a strategy that selects the fastest resource, and decrease makespan by a factor of five when compared to a round robin strategy, when deployed on a heterogeneous testbed with resources ranging from a Raspberry Pi to a GPU node in an academic cloud.

Computing continuum↗

Evaluating Unified Memory Performance in HIP

Heterogeneous unified memory management between a CPU and a GPU is a major challenge in GPU computing. Recently, unified memory (UM) has been supported by software and hardware components on AMD computing platforms. The support could simplify the complexities of memory management. In this paper, we attempt to have a better understanding of UM by evaluating the performance of UM programs on an AMD MI100 GPU. More specifically, we evaluate data migration using UM against other data transfer techniques for the overall performance of an application, assess the impacts of three commonly used optimization techniques on the kernel execution time of a vector add sample, and compare the performance and productivity of selected benchmarks with and without UM. The performance overhead associated with UM is not trivial, but it can improve programming productivity by reducing lines of code for scientific applications. We aim to present early results and feedback on the UM performance to the vendor.

Jin, Zheming↗

Remote Hardware-in-the-Loop Approach for Microgrid Controller Evaluation

Utilities have been installing microgrids because of the increased resilience and reliability advantages they may provide to the distribution system. A microgrid controller is a critical component in microgrids. It is of great benefit to derisk the installation of microgrid controllers before field deployment. Hardware-in-the-loop (HIL) testing is used by controller developers and utilities to evaluate the controllers under stressful conditions. In this work, a microgrid control function developed by the Synchrophasor Grid Monitoring and Automation (SyGMA) laboratory at the University of California, San Diego is tested in a remote HIL (RHIL) setup. The digital real-time simulation of the detailed microgrid system was operated at the National Renewable Energy Laboratory's Energy Systems Integration Facility. Under such RHIL setup, successful controller operation is contingent on understanding and characterizing the communications channel and in particular network latencies. The novelty of this paper is the proposed use of a RHIL setup that leverages existing power system communications protocols to evaluate the controller in conjunction with the simulation capabilities of a remote facility. The work presented here will provide the complete setup of the HIL evaluation platform, the details of the communications protocols used by the setup for data transfer between the two organizations, test cases developed to evaluate the controller, and the results from the experiments.

controller hardware-in-the-loop↗

Improving Signal-to-Noise Ratio (SNR) for Readout Signals Using Adaptive Filters on Reconfigurable Controls Hardware

This study investigates the optimization of Signal-to-Noise Ratio (SNR) in superconducting quantum computing readout signals through adaptive filtering. Quantum computing technology has the potential to revolutionize various fields by delivering exponential speedup in solving certain computational problems. However, the technology's practical implementation is hindered by the difficulty of extracting clean, reliable signals during the readout phase, with various sources of noise presenting a significant barrier to clean signals. This noise, often present in readout profiles due to imperfect isolation, degrades the system's overall SNR, thus impeding the ability to extract the quantum state accurately. The research leverages the power of adaptive filtering to improve the SNR of quantum computing readout signals. Specifically, an adaptive filter is implemented in a PYNQ overlay on an FPGA, and eventually will be connected to a quantum computing system. The system models the noise with a Least Mean Squares (LMS) adaptive filter, and then subtracts the estimated noise from the received signal to improve the SNR. A Direct Memory Access (DMA) channel is used to handle the signal processing, delivering efficient, high-speed data transfer between the PYNQ system and the hardware. The study explores the benefits of this adaptive filtering technique, potentially providing a significant contribution to practical and fast quantum computing.

Johnson, Hans↗

funcX: Federated Function as a Service for Science

Here, funcX is a distributed function as a service (FaaS) platform that enables flexible, scalable, and high performance remote function execution. Unlike centralized FaaS systems, funcX decouples the cloud-hosted management functionality from the edge-hosted execution functionality. funcX's endpoint software can be deployed, by users or administrators, on arbitrary laptops, clouds, clusters, and supercomputers, in effect turning them into function serving systems. funcX's cloud-hosted service provides a single location for registering, sharing, and managing both functions and endpoints. It allows for transparent, secure, and reliable function execution across the federated ecosystem of endpoints-enabling users to route functions to endpoints based on specific needs. funcX uses containers (e.g., Docker, Singularity, and Shifter) to provide common execution environments across endpoints. funcX implements various container management strategies to execute functions with high performance and efficiency on diverse funcX endpoints. funcX also integrates with an in-memory data store and Globus for managing data that may span endpoints. We motivate the need for funcX, present our prototype design and implementation, and demonstrate, via experiments on two supercomputers, that funcX can scale to more than 130000 concurrent workers. We show that funcX's container warming-aware routing algorithm can reduce the completion time for 3,000 functions by up to 61% compared to a randomized algorithm and the in-memory data store can speed up data transfers by up to 3x compared to a shared file system.

97 MATHEMATICS AND COMPUTING↗

Machine-learning-enabled on-the-fly analysis of RHEED patterns during thin film deposition by molecular beam epitaxy

Thin film deposition is a fundamental technology for the discovery, optimization, and manufacturing of functional materials. Deposition by molecular beam epitaxy (MBE) typically employs reflection high-energy electron diffraction (RHEED) as a real-time in situ probe of the growing film. However, the state-of-the-art for RHEED analysis during deposition requires human observation. Here, we present an approach using machine learning (ML) methods to monitor, analyze, and interpret RHEED images on-the-fly during thin film deposition. In the analysis workflow, RHEED pattern images are collected at one frame per second and featurized using a pretrained deep convolutional neural network. The feature vectors are then statistically analyzed to identify changepoints; these changepoints can be related to changes in the deposition mode from initial film nucleation to a transition regime, smooth film deposition, and in some cases, an additional transition to a rough, islanded deposition regime. The feature vectors are additionally analyzed via graph analysis and community classification. The graph is quantified as a stabilization plot, and we show that inflection points in the stabilization plot correspond to changes in the growth regime. The full RHEED analysis workflow is termed RHAAPsody and includes data transfer and output to a visual dashboard. We demonstrate the functionality of RHAAPsody by analyzing the precaptured RHEED images from epitaxial depositions of anatase TiO2 on SrTiO3(001) and show that the analysis workflow can be executed in less than 1 s. Our approach shows promise as one component of ML-enabled real-time feedback control of the MBE deposition process.

36 MATERIALS SCIENCE↗

Exploring the Use of Novel Spatial Accelerators in Scientific Applications

Driven by the need to find alternative accelerators which can viably replace GPUs in next-generation Supercomputing systems, this paper proposes a methodology to enable agile application/hardware co-design. The application-first methodology provides the ability to come up with design of accelerators while working with real-world workloads, available accelerators, and system software. The iterative design process targets a set of kernels in a workload for performance estimates that can prune the design space for later phases of detailed architectural evaluations. To this effect, in this paper, a novel data-parallel device model is introduced that simulates the latency of performance-sensitive operations in an accelerator including data transfers and kernel computation using multi-core CPUs. The use of off-the-shelf simulators, such as pre-RTL simulator Aladdin or multiple tools available for exploring the design of deep neural network accelerators (e.g., Timeloop) is demonstrated for evaluation of various accelerator designs using applications with realistic inputs. Examples of multiple device configurations that are instantiable in a system are explored to evaluate the performance benefit of deploying novel accelerators. The proposed device is integrated with a programming model and system software to potentially explore the impacts of high-level programming languages/compilers and low-level effects such as task scheduling on multiple accelerators. We analyze our methodology for a set of applications that represent high-performance computing (HPC) and graph analytics. The applications include a computational chemistry kernel realized using tensor contractions, triangle counting, GraphSAGE and Breadth-first Search. These applications include kernels such as dense matrix-dense matrix multiplication, sparse matrix-spare matrix multiplication, and sparse matrix-dense vector multiplication. Our results indicate potential performance benefits and insights for system design by including accelerators that realize these kernels along-side general purpose accelerators.

AI, codesign, Accelerated Computing, Modeling and ↗

Design and analysis of CXL performance models for tightly-coupled heterogeneous computing

Truly heterogeneous systems enable partitioned workloads to be mapped to the hardware that nets the best performance. However, current practice requires that inter-device communication between different vendors' hardware use host memory as an intermediary step. To date, there are no widely adopted solutions that allow accelerators to directly transfer data. A new cache-coherent protocol, CXL, aims to facilitate easier, fine-grained sharing between accelerators. In this work we analyze existing methods for designing heterogeneous applications that target GPUs and FPGAs working collaboratively, followed by an exploration to show the benefits of a CXL-enabled system. Specifically, we develop a test application that utilizes both an NVIDIA P100 GPU and a Xilinx U250 FPGA to show current communication limitations. From this application, we capture overall execution time and throughput measurements on the FPGA and GPU. We use these measurements as inputs to novel CXL performance models to show that using CXL caching instead of host memory results in a 1.31X speedup, while a more tightly-coupled pipelined implementation using CXL-enabled hardware would result in a speedup of 1.45X.

Cabrera, Anthony↗

IRIS-MASH: Efficient Multi-device Asynchronous Multi-Stream Heterogeneous Computing

In the rapidly evolving field of high-performance computing (HPC), effectively leveraging heterogeneous devices through asynchronous task programming is paramount. This paper presents a robust asynchronous task programming model tailored for a multi-device, multi-stream execution environment that incorporates a diverse array of heterogeneous computing units, including GPUs from various vendors and other accelerators. Current state-of-the-art task programming models provide methodologies to support asynchronous task executions, but they typically handle homogeneous devices using native programming languages, while support for heterogeneous devices is limited to frameworks like OpenCL. This gap presents significant challenges in abstracting heterogeneous devices to harness their true asynchronous capabilities effectively using their native programming languages. By implementing asynchronous task execution, our model significantly boosts the performance of tiled algorithm task graphs through overlapping data transfers with computation and enabling the simultaneous execution of multiple kernels. We integrate this approach into a heterogeneous Intelligent Runtime System (IRIS) and assess its performance using a suite of tiled algorithm benchmarks from the heterogeneous math kernels library (MatRIS) based on IRIS. Experimental results demonstrate a performance improvement ranging from 1.6 × to 2 × over IRIS without asynchronous support, and a notable 22% performance enhancement compared to established runtime systems such as StarPU and PaRSEC. This approach significantly improves computation efficiency of HPC workflows and provides a solid base for future exploration and development in the area of asynchronous task programming in heterogeneous systems.

Miniskar, Narasinga Rao [ORNL] (ORCID:000000018259↗

Concurrent Relaxation through Accelerated Deep Learning

CRADL captures performance metrics of machine learning algorithms operating on mesh data from multiphysics codes This proxy application is a tool to explore scalability of inference on HPC platforms, and also gather performance metrics for inference on new machine learning specific hardware. CRADL is designed to give users as fine a control as possible over an inference simulation. Users may select the number of cycles, amount of data, and batch size to pass to the accelerator of choice. Additionally the user may select a number of performance optimization libraries and flags. CRADL comes packaged with a repository of anonymized multi-physics simulation data, as well as a pretrained model for inference. The code allows a user to load their own pre-trained model and data if they wish. The code can operate in multiple parallelization schemes, with performance enhancing options such as half-precision libraries, PyTorch benchmarking, and pinned memory with non-blocking data transfers.

Zieb, KristoferJ.↗

ESnet Secure Copy (EScp) v0.6

EScp is a high speed transfer tool with a similar command line syntax to scp. Unlike SCP it is designed to transfer files at high speed, thus far we have been able to show 100gbit/s transfers, although I expect that the throughput should scale in proportion to the network interface, i.e. I expect 400gbit/s performance on our 400gbit/s test bed. EScp achieves good performance through an innovative design (multithreaded, zero copy transfers), along with pluggable filters and I/O engines. As an example, you can switch from POSIX i/O to UIO by checking a different engine. It also natively supports encryption, and cheksums for file verification and transport security. AAA is through standard SSH (same as SCP). By taking advantage of filters, EScp supports transferring unstructured data and/or I/O to non-posix data sources. Examples include streaming data (i.e. from equipment), transferring data to the cloud, and/or supporting non-posix file systems (like HPSS).

Shiflett, Charles↗

HPXML Version Translator

HPXML is a consensus data transfer standard for residential buildings. Over several years, multiple versions of the standard have been released. This tool accepts an HPXML file in an older version and translates it to a newer version.

Merket, Noel↗

TEE-ACM2

TEE-ACM2 is a library that computes matrix chain multiplication efficiently on GPUs, using a blocking strategy to load data. The library will provide an energy efficient algorithm for chain Matrix Multiplication, by minimizing both computations and off chip data transfers on the GPUs.

Lim, Hyun↗

artdaq

The artdaq toolkit is a data-acquisition framework designed for high-energy physics experiments. It provides a flexible, reliable backbone for data transfers and has several locations where users can perform custom analysis tasks using the art framework.

Flumerfelt, EricL. [Fermi National Accelerator Lab↗

DDCP framework

DDCP protocol software 1.0 This repository contains the C++ implementation of version 1.x of the Distributed Data Communications Protocol (DDCP). DDCP provides request/reply, feature discovery, data transfer, control, interrupt, and transaction support for communicating with accelerator instrumentation over UDP. The standard server port is 65000. The framework is a source dependency for services that communicate directly with DDCP hardware. It is not a deployable service by itself.

Joshi, Shreya [Fermi National Accelerator Laborato↗