Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data movement”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

A machine-learning-driven data labeling pipeline for scientific analysis in MLExchange

This study introduces a novel labeling pipeline to accelerate the labeling process of scientific data sets by using artificial intelligence (AI)-guided tagging techniques. This pipeline includes a set of interconnected web-based graphical user interfaces (GUIs), where Data Clinic and MLCoach enable the preparation of machine learning (ML) models for data reduction and classification, respectively, while Label Maker is used for label assignment. Throughout this pipeline, data can be accessed through a direct connection to a file system or through Tiled for access through Hypertext Transfer Protocol (HTTP). Our experimental results present three use cases where this labeling pipeline has been instrumental for the study of large X-ray scattering data sets in the area of pattern recognition, the remote analysis of resonant soft X-ray scattering data and the fine-tuning process of foundation models. These use cases highlight the labeling capabilities of this pipeline, including the ability to label large data sets in a short period of time, to perform remote data analysis while minimizing data movement and to enhance the fine-tuning process of complex ML models with human involvement.

Chavez, Tanny (ORCID:0000000193172896)↗

IRIS: A Portable Runtime System Exploiting Multiple Heterogeneous Programming Systems

Across embedded, mobile, enterprise, and HPC systems, computer architectures are becoming more heterogeneous and complex. This complexity is causing a crisis in programming systems and performance portability. Several programming systems are working to address these challenges, but the increasing architectural diversity is forcing software stacks and applications to specialize for each architecture. As we show, all of these approaches critically depend on their runtime system for discovery, execution, scheduling, and data orchestration. To address this challenge, we believe that a more agile and proactive runtime system is essential to increase performance portability and improve user productivity. In this regard, we have designed and implemented IRIS: a portable runtime system exploiting multiple heterogeneous programming systems. IRIS can discover available resources, manage multiple diverse programming systems (e.g., CUDA, Hexagon, HIP, Level Zero, OpenCL, OpenMP) simultaneously in the same execution, respect data dependencies, orchestrate data movement proactively, and provide for user-configurable scheduling. Our evaluation on three architectures, ranging from Qualcomm Snapdragon to a Summit supercomputer node, shows that IRIS improves portability across a wide range of diverse heterogeneous architectures with negligible overhead.

Kim, Jungwon↗

Adapting In Situ Accelerators for Sparsity With Granular Matrix Reordering

Neural network (NN) inference is an essential part of modern systems and is found at the heart of numerous applications ranging from image recognition to natural language processing. In situ NN accelerators can efficiently perform NN inference using resistive crossbars, which makes them a promising solution to the data movement challenges faced by conventional architectures. Although such accelerators demonstrate significant potential for dense NNs, they often do not benefit from sparse NNs, which contain relatively few non-zero weights. Processing sparse NNs on in situ accelerators results in wasted energy to charge the entire crossbar where most elements are zeros. To address this limitation, this paper proposes Granular Matrix Reordering (GMR): a preprocessing technique that enables an energy-efficient computation of sparse NNs on in situ accelerators. GMR reorders the rows and columns of sparse weight matrices to maximize the crossbars' utilization and minimize the total number of crossbars needed to be charged. The reordering process does not rely on sparsity patterns and incurs no accuracy loss. Finally, GMR achieves an average of 28% and up to 34% reduction in energy consumption over seven pruned NNs across four different pruning methods and network architectures.

97 MATHEMATICS AND COMPUTING↗

GPU Direct I/O with HDF5

Exascale HPC systems are being designed with accelerators, such as GPUs, to accelerate parts of applications. In machine learning workloads as well as large-scale simulations that use GPUs as accelerators, the CPU (or host) memory is currently used as a buffer for data transfers between GPU (or device) memory and the file system. If the CPU does not need to operate on the data, then this is sub-optimal because it wastes host memory by reserving space for duplicated data. Furthermore, this “bounce buffer” approach wastes CPU cycles spent on transferring data. A new technique, NVIDIA GPUDirect Storage (GDS), can eliminate the need to use the host memory as a bounce buffer. Thereby, it becomes possible to transfer data directly between the device memory and the file system. This direct data path shortens latency by omitting the extra copy and enables higher-bandwidth. To take full advantage of GDS in existing applications, it is necessary to provide support with existing I/O libraries, such as HDF5 and MPI-IO, which are heavily used in applications. In this paper, we describe our effort of integrating GDS with HDF5, the top I/O library at NERSC and at DOE leadership computing facilities. We design and implement this integration using a HDF5 Virtual File Driver (VFD). The GDS VFD provides a file system abstraction to the application that allows HDF5 applications to perform I/O without the need to move data between CPUs and GPUs explicitly. We compare performance of the HDF5 GDS VFD with explicit data movement approaches and demonstrate superior performance with the GDS method.

Ravi, J↗

Quantum/AI Topology-Aware Latency-Adaptive HPC Workflow Scheduling Optimization

The growing demand for more powerful high-performance computing (HPC) systems has led to a steady rise in energy consumption by supercomputing worldwide. This study is focused on comparing our Application-Topology Mapper (ATMapper) to the popular Simple Linux Utility for Resource Management (SLURM) for the purpose of exploring methods that can further optimize job-scheduling within HPC systems. ATMapper is an Artificial-Intelligence based approach to job-scheduling that is currently being enhanced with quantum annealing (QA) to generate optimal schedules faster. We are applying QA to speedup our ATMapper process to achieve higher computing efficiency, thereby reducing HPC energy consumption. Here, we examine how four job-scheduling approaches perform in processor node assignment when using an example network architecture of 4 interconnected nodes. Using a specialized script, we are assessing the schedule of a computation flow with 11 interdependent tasks. The data movements among nodes were tracked to count for the number of interactions (network hops) between nodes needed to complete the tasks. The total number of hops and the job completion time were then used to quantify the efficiency of the different mapping approaches. In addition to SLURM, we also compare our ATMapper to the QA-enabled LBNL TIGER and the D-Wave Distributed Computing processor assignment approaches. The preliminary results showed that our topology-aware, latency-adaptive ATMapper is significantly more efficient when compared to the other scheduling approaches due to its load-imbalance network allocation. The scheduler displayed a computing efficiency of 53% by performing significantly fewer network hops than its alternatives. By reducing the number of hops, ATMapper was able to perform all 11 tasks by using only 3 nodes out of given 4. This research indicates the potential to use QA/AI for HPC job-scheduling. Later, we will test a SLURM simulator program to draw further comparisons on the effectiveness of ATMapper's scheduling approach. The results of this comparison will serve as a baseline for later improving SLURM's performance using a QA-enhanced ATMapper approach.

Caraveo, Braulio [University of Huston - Clear Lak↗

To Derive or Not to Derive: I/O Libraries Take Charge of Derived Quantities Computation

The ever-increasing volume of data produced by HPC simulations necessitates scalable methods for data exploration and knowledge extraction. Scientific data analysis often involves complex queries across distributed datasets, requiring manipulation of multiple primary variables and generating derived data that needs to be handled efficiently, creating challenges for applications that need to parse many large datasets. Relying on individual applications to handle all intermediate data generally leads to redundant computations across studies and unnecessary data transfers. In this paper, we investigate the performance of different approaches where applications define derived variables as quantities of interest (QoIs) and offload the computation and transfer of these QoIs to the I/O library. This significantly reduces redundancy and optimizes data movement across the distributed storage and processing infrastructure by allowing control over when and where derived variables are computed. We present a detailed analysis of the performance-storage trade-offs associated with different solutions and showcase results for our study on two large-scale datasets created from climate and combustion simulations.

Gainaru, Ana↗

Scalable Deep Learning-Based Microarchitecture Simulation on GPUs

Cycle-accurate microarchitecture simulators are essential tools for designers to architect, estimate, optimize, and manufacture new processors that meet specific design expectations. However, conventional simulators based on discrete-event methods often require an exceedingly long time-to-solution for the simulation of applications and architectures at full complexity and scale. Given the excitement around wielding the machine learning (ML) hammer to tackle various architecture problems, there have been attempts to employ ML to perform architecture simulations, such as Ithemal and SimNet. However, the direct application of existing ML approaches to architecture simulation may be even slower due to overwhelming memory traffic and stringent sequential computation logic. This work proposes the first graphics processing unit (GPU)-based microarchitecture simulator that fully unleashes the potential of GPUs to accelerate state-of-the-art ML-based simulators. First, considering the application traces are loaded from central processing unit (CPU) to GPU for simulation, we introduce various designs to reduce the data movement cost between CPUs and GPUs. Second, we propose a parallel simulation paradigm that partitions the application trace into sub-traces to simulate them in parallel with rigorous error analysis and effective error correction mechanisms. Combined, this scalable GPU-based simulator outperforms by orders of magnitude the traditional CPU-based simulators and the state-of-the-art ML-based simulators, i.e., SimNet and Ithemal.

97 MATHEMATICS AND COMPUTING↗

Functional Adaptation of LPS-affected Dentoalveolar Fibrous Joints in Rats

We report the functional interplay between cementum of the root and alveolar bone of the socket is tuned by a uniquely positioned 70–80 µm wide fibrous and lubricious ligament in a dentoalveolar joint (DAJ). In this study, structural and biomechanical properties of the DAJ, periodontal ligament space (PDL-space also known as the joint space), alveolar bone of the socket, and cementum of the tooth root that govern the biomechanics of a lipopolysaccharide (LPS)-affected DAJ were mapped both in space and time. The hemi-maxillae from 20 rats (4 control at 6 weeks of age, 4 control and 4 LPS-affected at 12 weeks of age, 4 control and 4 LPS-affected at 16 weeks of age) were investigated using a hybrid technique; micro-X-ray computed tomography (5 µm resolution) in combination with biomechanical testing in situ. Temporal variations in bone and cementum volume fractions were evaluated. Trends in mineral apposition rates (MAR) in additional six Sprague Dawley rats (3 controls, 3 LPS-affected) were revealed by transforming spatial fluorochrome signals to functional growth rates (linearity factor - RW) of bone, dentin, and cementum using a fast Fourier transform on fluorochrome signals from 100-µm hemi-maxillae sections. An overall change in LPS-affected DAJ biomechanics (a 2.5–4.5X increase in tooth displacement and 2X tooth rotation at 6 weeks, no increase in displacement and a 7X increase in rotation at 12 weeks; 27% increase in bone effective strain at 6 weeks and 11% at 12 weeks relative to control) was associated with structural changes in the coronal regions of the DAJ (15% increase in PDL-space from 0 to 6 weeks but only 5% from 6 to 12 weeks compared to control). A significant increase (p < 0.05) in PDL-space between ligated and age-matched control was observed. The bone fraction of ligated at 12 weeks was significantly lower than its age-matched control, and no significant differences (p > 0.05) between groups were observed at 6 weeks. Cementum in the apical regions grew faster but nonlinearly (11% and 20% increase in cementum fraction (CF) at 6 and 12 weeks) compared to control. Alveolar bone revealed site-specific nonlinear growth with an overall increase in MAR (108.5 µm/week to 126.7 µm/week after LPS treatment) compared to dentin (28.3 µm/week in control vs. 26.1 µm/week in LPS-affected) and cementum (126.5 µm/week in control vs. 119.9 µm/week in LPS-affected). A significant increase in CF (p < 0.05) in ligated specimens was observed at 6 weeks of age. Anatomy-specific responses of cementum and bone to the mechano-chemo stimuli, and their collective temporal contribution to observed changes in PDL-space were perpetuated by altered tooth movement. Data highlight the “resilience” of DAJ function through the predominance of nonlinear growth response of cementum, changes in PDL-space, and bone architecture. Despite the significant differences in bone and cementum architectures, data provided insights into the reactionary effects of cementum as a built-in compensatory mechanism to reestablish functional competence of the DAJ. The spatial shifts in architectures of alveolar bone and cementum, and consequently ligament space, highlight adaptations farther away from the site of insult, which also is another novel insight from this study. These adaptations when correlated within the context of joint function (biomechanics) illustrate that they are indeed necessary to sustain DAJ function albeit being pathological.

60 APPLIED LIFE SCIENCES↗

Communication Lower Bounds and Optimal Algorithms for Multiple Tensor-Times-Matrix Computation

Multiple tensor-times-matrix (Multi-TTM) is a key computation in algorithms for computing and operating with the Tucker tensor decomposition, which is frequently used in multidimensional data analysis. Here, we establish communication lower bounds that determine how much data movement is required (under mild conditions) to perform the Multi-TTM computation in parallel. The crux of the proof relies on analytically solving a constrained, nonlinear optimization problem. We also present a parallel algorithm to perform this computation that organizes the processors into a logical grid with twice as many modes as the input tensor. We show that, with correct choices of grid dimensions, the communication cost of the algorithm attains the lower bounds and is therefore communication optimal. Finally, we show that our algorithm can significantly reduce communication compared to the straightforward approach of expressing the computation as a sequence of tensor-times-matrix operations when the input and output tensors vary greatly in size.

HBL-inequalities↗

HQ-Sim: High-performance State Vector Simulation of Quantum Circuits on Heterogeneous HPC Systems

Quantum circuit simulations are applied in more and more circum- stances as the quantum computing community becomes broader. It helps researchers to evaluate quantum algorithms and relieve the burden of limited quantum computing resources. However, most of the state-of-the-art quantum simulators utilize either CPU or GPU to store and calculate the state vector, which results in resource starvation. Moreover, the maximum number of qubits supported by the simulator is bounded by the memory, since the memory utilization increases exponentially with the number of qubits. In this study, we leverage Heterogeneous computing to utilize both CPU and GPU to store and update state vectors. We also integrate lossy data compression to reduce memory requirements. Specifically, we develop a heterogeneous framework that has a dynamic scheduler to fully utilize the computing resources. We apply lossy compression to chunked state vector to make the maximum number of qubits higher than the regular simulators, the compression also benefits the data movement between CPU and GPU.

Zhang, Boyuan↗

cuAlign: Scalable Network Alignment on GPU Accelerators

Given two graphs, the objective of network alignment is to find the best one-to-one mapping of vertices in one graph (??) to vertices in the other (??), such that the number of overlaps is maximized. We say that edges(??, ??) ???and(??', ??') ??? are overlapped if ?? is mapped to ??' and ?? is mapped to??'. Network alignment is an important optimization problem with several applications in bioinformatics, computer vision and ontology matching. Since it is an NP-hard problem, efficient heuristics and scalable implementations are necessary. In this work, we introduce a new framework that combines the concepts of intra-network proximity using vertex embedding,Belief Propagation (BP) and approximate weighted matching, and provides qualitative improvements up to22%over state-of-the-art approaches. We also provide scalable implementations on GPU accelerators, demonstrating up to19×speedup for Belief Propagation and 3× speedup for approximate weighted matching relative to previous multithreaded implementation. A combination of combinatorial and algebraic kernels within the network alignment algorithm poses significant hurdles for parallelization. Load imbalance and irregular DRAM traffic limit achievable performance on GPUs. Our novel approach identifies and exploits unique structural proper-ties of the BP-based algorithm and employs code fusion to reduce data movement between different steps of the algorithm. Using a diverse set of inputs, we demonstrate qualitative improvements of our algorithms, and performance gains of our GPU-accelerated implementation. We believe that our work will enable algorithmic improvements and practical applications of network alignment.

Xiang, Lizhi↗

Accelerated Constrained Sparse Tensor Factorization on Massively Parallel Architectures

This study presents the first constrained sparse tensor factorization (cSTF) framework that optimizes and fully offloads computation to massively parallel GPU architectures, and the first performance characterization of cSTF on GPU architectures. In contrast to prior work on tensor factorization, where the matricized tensor times Khatri-Rao product (MTTKRP) is the primary performance bottleneck, our systematic analysis of the cSTF algorithm on GPUs reveals that adding constraints creates an additional bottleneck in the update operation for many real-world sparse tensors. While executing the update operation on the GPU brings significant speedup over its CPU counterpart, it remains a significant bottleneck. To further accelerate the update operation, we propose cuADMM, a new update algorithm that leverages algorithmic and code optimization strategies to minimize both computation and data movement on GPUs. As a result, our framework delivers significantly improved performance compared to prior state-of-the-art. On 10 real-world sparse tensors, our framework achieves geometric mean speedup of 5.1 × (max 41.59 ×) and 7.01 × (max 58.05 ×) on the NIVIDA A100 and H100 GPUs, respectively, over the state-of-the-art SPLATT library running on a 26-core Intel Ice Lake Xeon CPU.

Soh, Yongseok↗

CGSim: A Simulation Framework for Large Scale Distributed Computing Environment

Large-scale distributed computing infrastructures such as the Worldwide LHC Computing Grid (WLCG) require comprehensive simulation tools for evaluating performance, testing new algorithms, and optimizing resource allocation strategies. However, existing simulators suffer from limited scalability, hardwired algorithms, lack of real-time monitoring, and inability to generate datasets suitable for modern machine learning approaches. We present CGSim, a simulation framework for large-scale distributed computing environments that addresses these limitations. Built upon the validated SimGrid simulation framework, CGSim provides high-level abstractions for modeling heterogeneous grid environments while maintaining accuracy and scalability. Key features include a modular plugin mechanism for testing custom workflow scheduling and data movement policies, interactive real-time visualization dashboards, and automatic generation of event-level datasets suitable for AI-assisted performance modeling. We demonstrate CGSim’s capabilities through a comprehensive evaluation using production ATLAS PanDA workloads, showing significant calibration accuracy improvements across WLCG computing sites. Scalability experiments show near-linear scaling for multi-site simulations, with distributed workloads achieving 6 × better performance compared to single-site execution. The framework enables researchers to simulate WLCG-scale infrastructures with hundreds of sites and thousands of concurrent jobs within practical time budget constraints on commodity hardware.

Vatsavai, Sairam Sri [Brookhaven National Laborato↗

FLYING SERVING: On-the-Fly Parallelism Switching for Large Language Model Serving

Production LLM serving must simultaneously deliver high throughput, low latency, and sufficient context capacity under non-stationary traffic and mixed request requirements. Data parallelism (DP) maximizes throughput by running independent replicas, while tensor parallelism (TP) reduces per-request latency and pools memory for long-context inference. However, existing serving stacks typically commit to a static parallelism configuration at deployment; adapting to bursts, priorities, or long-context requests is often disruptive and slow. We present Flying Serving, a vLLM-based system that enables online DP-TP switching without restarting engine workers. Flying Serving makes reconfiguration practical by virtualizing the state that would otherwise force data movement: (i) a zero-copy Model Weights Manager that exposes TP shard views on demand, (ii) a KV Cache Adaptor that preserves request KV state across DP/TP layouts, (iii) an eagerly initialized Communicator Pool to amortize collective setup, and (iv) a deadlock-free scheduler that coordinates safe transitions under execution skew. Across three popular LLMs and realistic serving scenarios, Flying Serving improves performance by up to 4.79 × under high load and 3.47 × under low load while supporting latency- and memory-driven requests.

Gao, Shouwei [ORNL]↗

BASIC HARTREE-FOCK PROXY APPLICTION

This proxy application simulates the compute load and data-movement of the kernel of the Hartree-Fock method in quantum chemistry. The proxy features a simplified algorithm for computing electron repulsion integrals that is easily offloaded to GPUs.

FLETCHER, GRAHAMD↗

pnnl/arena

CFA ARENA is a novel programming model with the support of a runtime targeting asynchronous data-centric execution paradigm in a distributed system. All the machine nodes in ARENA are connected by a ring network to bring the specialized computation to the data rather than the reverse to minimize data movement. The programming interfaces are implemented using C++

Tan, Cheng↗

Janus v1.0

Janus provides a software framework for lightweight container management and orchestration. It's primary use cases are around deploying containerized services for high-performance data movement needs. Thus, Janus differentiates itself from systems like Kubernetes by tailoring the deployment of containers around network, storage, and host tuning optimizations. Janus uses the concept of profiles to capture repeatable deployment patterns and applies them to container execution across one or more endpoints. A Janus Agent component provides remote host tuning and monitoring capabilities.

Essiari, Abdelilah [Lawrence Berkeley National Lab↗

Bluecrab: Comprehensive Reactor Analysis Bundle

BlueCRAB is a combination of existing codes created in close collaboration with the U.S. NRC useful for various reactor safety analysis simulations. For the purpose of classification it should be noted that BlueCRAB (the bundle) can be split up in several ways. At the heart of the software is "wrapping" and "coupling" code for brining in several non-INL projects from the NRC and Argonne National Laboratory (SAM). These codes facilitate the building and linking of these various codes during compilation and assist with data movement during execution. BlueCRAB can optionally link in several other applications including the following: BISON: fuels performance Griffin: Reactor Physics Pronghorn: CFD IAPS95: EOS for water, helium, nitrogen TRACE: NRC Code for 2 phase flow (system analysis) FAST: NRC Code for fuels performance SAM: ANL Code for single phase system analysis

Permann, Cody↗