Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data Transfer”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Multiplexed Readout for an Experiment with a Large Number of Channels Using Single-Electron Sensitivity Skipper-CCDs

This paper presents the implementation of a multiplexed analog readout electronics system that can achieve single-electron counting using Skipper-CCDs with non-destructive readout. The proposed system allows the best performance of the sensors to be maintained, with sub-electron noise-level operation, while maintaining low-bandwidth data transfer, a minimum number of analog-to-digital converters (ADC) and low disk storage requirement with zero added multiplexing time, even for the simultaneous operation of thousands of channels. These features are possible with a combination of analog charge pile-up, sample and hold circuits and analog multiplexing. The implementation also aims to use the minimum number of components in circuits to keep compatibility with high-channel-density experiments using Skipper-CCDs for low-threshold particle detection applications. Performance details and experimental results using a sensor with 16 output stages are presented along with a review of the circuit design considerations.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Porting the WAVEWATCH III (v6.07) wave action source terms to GPU

Abstract. Surface gravity waves play a critical role in several processes, including mixing, coastal inundation, and surface fluxes. Despite the growing literature on the importance of ocean surface waves, wind–wave processes have traditionally been excluded from Earth system models (ESMs) due to the high computational costs of running spectral wave models. The development of the Next Generation Ocean Model for the DOE’s (Department of Energy) E3SM (Energy Exascale Earth System Model) Project partly focuses on the inclusion of a wave model, WAVEWATCH III (WW3), into E3SM. WW3, which was originally developed for operational wave forecasting, needs to be computationally less expensive before it can be integrated into ESMs. To accomplish this, we take advantage of heterogeneous architectures at DOE leadership computing facilities and the increasing computing power of general-purpose graphics processing units (GPUs). This paper identifies the wave action source terms, W3SRCEMD, as the most computationally intensive module in WW3 and then accelerates them via GPU. Our experiments on two computing platforms, Kodiak (P100 GPU and Intel(R) Xeon(R) central processing unit, CPU, E5-2695 v4) and Summit (V100 GPU and IBM POWER9 CPU) show respective average speedups of 2× and 4× when mapping one Message Passing Interface (MPI) per GPU. An average speedup of 1.4× was achieved using all 42 CPU cores and 6 GPUs on a Summit node (with 7 MPI ranks per GPU). However, the GPU speedup over the 42 CPU cores remains relatively unchanged (∼ 1.3×) even when using 4 MPI ranks per GPU (24 ranks in total) and 3 MPI ranks per GPU (18 ranks in total). This corresponds to a 35 %–40 % decrease in both simulation time and usage of resources. Due to too many local scalars and arrays in the W3SRCEMD subroutine and the huge WW3 memory requirement, GPU performance is currently limited by the data transfer bandwidth between the CPU and the GPU. Ideally, OpenACC routine directives could be used to further improve performance. However, W3SRCEMD would require significant code refactoring to make this possible. We also discuss how the trade-off between the occupancy, register, and latency affects the GPU performance of WW3.

58 GEOSCIENCES↗

Alquimia v1.0: a generic interface to biogeochemical codes – a tool for interoperable development, prototyping and benchmarking for multiphysics simulators

Alquimia v1.0 is a generic interface to geochemical solvers that facilitates development of multiphysics simulators by enabling code coupling, prototyping and benchmarking. The interface enforces the function arguments and their types for setting up, solving, serving up output data and carrying out other common auxiliary tasks while providing a set of structures for data transfer between the multiphysics code driving the simulation and the geochemical solver. Alquimia relies on a single-cell approach that permits operator splitting coupling and parallel computation. We describe the implementation in Alquimia of two widely used open-source codes that perform geochemical calculations: PFLOTRAN and CrunchFlow. We then exemplify its use for the implementation and simulation of reactive transport in porous media by two open-source flow and transport simulators: Amanzi and ParFlow. We also demonstrate its use for the simulation of coupled processes in novel multiphysics applications including the effect of multiphase flow on reaction rates at the pore scale with OpenFOAM, the role of complex biogeochemical processes in land surface models such as the E3SM Land Model (ELM) and the impact of surface–subsurface hydrological interactions on hydrogeochemical export from watersheds with the Advanced Terrestrial Simulator (ATS). These applications make it apparent that the availability of a well-defined yet flexible interface has the potential to improve the software development workflow, freeing up resources to focus on advances in process models and mechanistic understanding of coupled problems.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

HAMR - Heterogeneous Accelerator Memory Resource (HAMR) v1.0

HAMR is a library defining an accelerator technology agnostic memory model that bridges between accelerator technologies (CUDA, HIP, ROCm, OpenMP, Sycl, OpenCL, Kokos, etc) and traditional CPUs in heterogeneous computing environments. HAMR is light weight and implemented in modern C++. HAMR can be used to manage memory with in a single code or as a data model for coupling codes in a technologically agnostic way. HAMR provides a Python module for coupling C++ and Python codes which implements zero-copy data transfers to and from Python using the Numpy array interface and Numba CUDA array interface protocols.

Loring, Burlen↗

Remote Hardware-in-the-Loop Approach for Microgrid Controller Evaluation

Utilities have been installing microgrids because of the increased resilience and reliability advantages they may provide to the distribution system. A microgrid controller is a critical component in microgrids. It is of great benefit to derisk the installation of microgrid controllers before field deployment. Hardware-in-the-loop (HIL) testing is used by controller developers and utilities to evaluate the controllers under stressful conditions. In this work, a microgrid control function developed by the Synchrophasor Grid Monitoring and Automation (SyGMA) laboratory at the University of California, San Diego is tested in a remote HIL (RHIL) setup. The digital real-time simulation of the detailed microgrid system was operated at the National Renewable Energy Laboratory's Energy Systems Integration Facility. Under such RHIL setup, successful controller operation is contingent on understanding and characterizing the communications channel and in particular network latencies. The novelty of this paper is the proposed use of a RHIL setup that leverages existing power system communications protocols to evaluate the controller in conjunction with the simulation capabilities of a remote facility. The work presented here will provide the complete setup of the HIL evaluation platform, the details of the communications protocols used by the setup for data transfer between the two organizations, test cases developed to evaluate the controller, and the results from the experiments.

24 POWER TRANSMISSION AND DISTRIBUTION↗

The updated DESGW processing pipeline for the third LIGO/VIRGO observing run

The DESGW group seeks to identify electromagnetic counterpartsof gravitational wave events seen by the LIGO-VIRGO network, such as thoseexpected from binary neutron star mergers or neutron star- black hole mergers.DESGW was active throughout the first two LIGO observing seasons, followingup several binary black hole mergers and the first binary neutron star merger,GW170817. We describe the modifications to the observing strategy generationand image processing pipeline between the second (ending in August 2017)and third (beginning in April 2019) LIGO observing seasons. The modifica-tions include a more robust observing strategy generator, further parallelizationof the image reduction software and dierence imaging processing pipeline,data transfer streamlining, and a web page listing identified counterpart candi-dates that updates in real time. Taken together, the additional parallelizationsteps enable us to identify potential electromagnetic counterparts within fullycalibrated search images in less than one hour, compared to the 3-5 hours itwould typically take during the first two seasons. These performance improve-ments are critical to the entire EM followup community, as rapid identification(or rejection) of candidates enables detailed spectroscopic followup by multipleinstruments as soon as possible, leading to more information about the environ-ment immediately following such gravitational wave events.

Herner, Kenneth R.↗

High-Level Model Articulation with BuildingSync and OpenStudio

The use of the BuildingSync schema for describing the contents of buildings is becoming more common with its recent integration into the Audit Template tool as well as ASHRAE’s Building Energy Quotient (bEQ) web portal. Although BuildingSync was initially created to store and transfer data related to building energy audits (as defined by ASHRAE Standard 211), it has since been expanded to store the data needed to articulate fully defined physics-based building energy models. BuildingSync combined with abstracted high-level input methods defined in OpenStudio’s Standards project and the newly developed BuildingSync gem allows BuildingSync eXtensible Markup Language (XML) (Bray, Paoli, Sperberg-McQueen, Maler, Eve (Sun Microsystems, & Francois, 2008) document to be converted into OpenStudio models. Automatic generation of a building energy model from data collected during an audit 1) eliminates the need to separately generate an energy model, 2) improves consistency between model generation, and 3) simplifies evaluation of various energy efficiency measures. This manuscript will discuss this open-source project, including: development of the validation infrastructure necessary to provide formalized expectations of informational requirements for BuildingSync documents; development of the BuildingSync gem for translation of BuildingSync documents to OpenStudio models as well as example implementations; and finally, the manuscript will elaborate on the advantages and disadvantages of using high-level models generated from BuildingSync as surrogates to detailed models.

30 DIRECT ENERGY CONVERSION↗

C-SAW: a framework for graph sampling and random walk on GPUs

Many applications require to learn, mine, analyze and visualize large-scale graphs. These graphs are often too large to be addressed efficiently using conventional graph processing technologies. Fortunately, recent research efforts find out graph sampling and random walk, which significantly reduce the size of original graphs, can benefit the tasks of learning, mining, analyzing and visualizing large graphs by capturing the desirable graph properties. This paper introduces C-SAW, the first framework that accelerates Sampling and Random Walk framework on GPUs. Particularly, C-SAW makes three contributions: First, our framework provides a generic API which allows users to implement a wide range of sampling and random walk algorithms with ease. Second, offloading this framework on GPU, we introduce warp-centric parallel selection, and two novel optimizations for collision migration. Third, towards supporting graphs that exceed the GPU memory capacity, we introduce efficient data transfer optimizations for out-of-memory and multi-GPU sampling, such as workload-aware scheduling and batched multi-instance sampling. Taken together, our framework constantly outperforms the state of the art projects in addition to the capability of supporting a wide range of sampling and random walk algorithms.

97 MATHEMATICS AND COMPUTING↗

Overcoming obstacles to IPv6 on WLCG

The transition of the Worldwide Large Hadron Collider Computing Grid (WLCG) storage services to dual-stack IPv6/IPv4 is almost complete; all Tier-1 and 94% of Tier-2 storage are IPv6 enabled. While most data transfers now use IPv6, a significant number of IPv4 transfers still occur even when both endpoints support IPv6. This paper presents the ongoing efforts of the HEPiX IPv6 working group to steer WLCG toward IPv6-only services by investigating and fixing the obstacles to the use of IPv6 and identifying cases where IPv4 is used when IPv6 is available. Removing IPv4 use is essential for the long-term agreed goal of IPv6-only access to resources within WLCG, thus eliminating the complexity and security concerns associated with dual-stack services. We present our achievements and ongoing challenges as we navigate the final stages of the transition from IPv4 to IPv6 within WLCG.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Investigation of CAD-based Geometry Workflows for Multiphysics Fusion Problems Using OpenMC and MOOSE

Fusion system designs are complex and require intricate and accurate meshes to be properly modeled. In this study, we investigate the use of CAD-based geometry workflows in fusion systems multiphysics problems. A simplified tokamak was introduced and modeled in CAD using a multiphysics coupling of OpenMC Monte Carlo transport and MOOSE heat conduction. The meshed geometry was prepared using direct accelerated geometry Monte Carlo (DAGMC) for particle transport, and a volumetric mesh was also prepared to be used in MOOSE and to tally OpenMC results. Cardinal was used to run OpenMC Monte Carlo particle transport within MOOSE framework. The heat source distribution and tritium production were calculated in OpenMC. The data transfer system was used to transfer heat source and temperature distribution between OpenMC and MOOSE. Two computational studies related to mesh refinement were performed: (1) refining the DAGMC and volumetric meshes used for tallying results and solving heat conduction and (2) only refining the DAGMC particle transport mesh. The refinement of the tally mesh has a much larger effect on the runtime compared to the refinement of the DAGMC particle transport surface mesh.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Systems Innovation: Modernization & Efficiencies for ESH&Q Reviews

Environmental compliance reviews at INL have traditionally been managed through fragmented systems, relying on multiple spreadsheets and manual processes. This inefficiency led to time-consuming status updates and redundant tasks, such as manually sending reminder emails and transferring data from Excel to the Environmental Review Process (ERP). Initial attempts to streamline these processes using Power Automate and Excel revealed significant limitations, necessitating a more comprehensive solution. To address these immediate inefficiencies, automated workflows were developed using Power Automate. These workflows were designed to send scheduled status update reminders and capture responses through standardized forms, with submitted data flowing directly into centralized Excel trackers. This automation reduced the administrative burden, improved data accuracy, and enabled faster, more consistent reporting. Specifically, email automation achieved a 65% efficiency gain, while data integration saw a 48% improvement, resulting in 91% of project statuses being updated within two months. Despite the improvements brought by Power Automate, the fragmented nature of the review processes persisted. To further enhance efficiency and accuracy, the Integrated Review Tool (IRT) was developed. The IRT aims to centralize review initiation and connect team systems, creating an interconnected data infrastructure that preserves team autonomy while enhancing overall efficiency. This tool automates email reminders, centralizes reviews, and streamlines data integration, significantly improving the accuracy and efficiency of environmental compliance reviews. The design and development of the IRT involved advanced systems methodology, process mapping, project management, and collaboration with subject matter experts. The minimum viable product design is 100% complete, and system development is currently underway, with expected outcomes including a centralized entry point for all ESH&Q reviews, automated routing, real-time tracking and analytics, AI integration, and a user-friendly interface. This project demonstrates the potential of leveraging automation and integrated systems to enhance efficiency, accuracy, and decision-making in environmental reporting and compliance processes at INL.

99 - GENERAL AND MISCELLANEOUS↗

Latency Analysis of the Nexus Digital Twin Framework

Real-time digital catalogs are increasingly relied upon to track metadata and connect disparate data sources for cloud-based data integration efforts. One such tool, Deeplynx Nexus is supporting real-time digital twin efforts through event-driven data integration and time-series queries. Nexus’s usefulness for these applications depends critically on how quickly individual records can be uploaded and downloaded, since delays directly affect the responsiveness of any system built on top of it. However, the actual latency a user should expect from Nexus has not been systematically measured before, particularly for the small, frequent transactions typical of live sensor feeds. Here we show that single-record round-trip latency is 61.1 ms on a local Nexus instance and 391.7 ms on the hosted production infrastructure, a roughly 6.4x difference driven primarily by fixed per-request overhead rather than data volume. This overhead dominates at small scale: comparing single-record and ten-record trials suggests approximately 56 ms of each single-record request is fixed connection and authentication cost rather than data-transfer time, meaning batching even a handful of records is substantially more efficient than transmitting them individually. At large batch sizes, this pattern reverses for uploads, which converge to near parity between local and hosted environments by 25,000-50,000 records, while download latency remains persistently 5.7-6.4x slower on hosted infrastructure even at scale. These results suggest that Nexus deployments intended for real-time digital twin applications should prioritize record batching over single-record transactions, and that download-path optimization on hosted infrastructure offers the largest remaining opportunity to reduce latency at scale. We anticipate these baseline measurements will serve as a reference point for future digital twin projects evaluating whether Nexus’s latency profile meets their real-time requirements, and as a benchmark for tracking the effect of future infrastructure or API changes.

99 - GENERAL AND MISCELLANEOUS↗

Transferability of data-driven, many-body models for CO 2 simulations in the vapor and liquid phases

Here, extending on the previous work by Riera et al. [J. Chem. Theory Comput. 16, 2246–2257 (2020)], we introduce a second generation family of data-driven many-body MB-nrg models for CO 2 and systematically assess how the strength and anisotropy of the CO 2 –CO 2 interactions affect the models’ ability to predict vapor, liquid, and vapor–liquid equilibrium properties. Building upon the many-body expansion formalism, we construct a series of MB-nrg models by fitting one-body and two-body reference energies calculated at the coupled cluster level of theory for large monomer and dimer training sets. Advancing from the first generation models, we employ the charge model 5 scheme to determine the atomic charges and systematically scale the two-body energies to obtain more accurate descriptions of vapor, liquid, and vapor–liquid equilibrium properties. Challenges in model construction arise due to the anisotropic nature and small magnitude of the interaction energies in CO 2 , calling for the necessity of highly accurate descriptions of the multidimensional energy landscape of liquid CO 2 . These findings emphasize the key role played by the training set quality in the development of transferable, data-driven models, which, accurately representing high-dimensional many-body effects, can enable predictive computer simulations of molecular fluids across the entire phase diagram.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Nuclear Physics Network Requirements Review Report

The Energy Sciences Network (ESnet) is the Office of Science’s high-performance network user facility, delivering highly reliable data transport capabilities optimized for the requirements of data-intensive science. In essence, ESnet is the circulatory system that enables the U.S. Department of Energy (DOE) science mission by connecting each and every DOE lab and its user facilities. ESnet is funded and stewarded by the Advanced Scientific Computing Research (ASCR) Program and managed and operated by the Scientific Networking Division at Lawrence Berkeley National Laboratory (LBNL). ESnet is widely regarded as a global leader in the research and education networking community. ESnet connects DOE national laboratories, user facilities, and major experiments so scientists can use remote instruments and computing resources as well as share data with collaborators, transfer large data sets, and access distributed data repositories. While ESnet provides network connectivity, it cannot be characterized as an internet service provider as it is specifically built to provide a range of network services that are tailored to meet the unique requirements of DOE’s data-intensive science.

97 MATHEMATICS AND COMPUTING↗

GPU-based Image Compression for Efficient Compositing in Distributed Rendering Applications

Visualizations of large-scale data sets are often created on graphics clusters that distribute the rendering task amongst many processes. When using real-time GPU-based graphics algorithms, the most time-consuming aspect of distributed rendering is typically the com-positing phase - combining all partial images from each rendering process into the final visualization. Compo siting requires image data to be copied off the GPU and sent over a network to other processes. While compression has been utilized in existing distributed rendering compositors to reduce the data being sent over the network, this compression tends to occur after the raw images are transferred from the GPU to main memory. In this paper, we present work that leverages OpenGL / CUDA interoperability to compress raw images on the GPU prior to transferring the data to main memory. This approach can significantly reduce the device-to-host data transfer time, thus enabling more efficient compositing of images generated by distributed rendering applications.

Lipinksi, Riley↗

Operational Evolution of FTS3: A DevOps Driven Approach to Elastic Operations

The File Transfer Service (FTS3) is a distributed data movement service developed at CERN and widely used to transfer data across the Worldwide LHC Computing Grid (WLCG). At Fermilab, FTS3 supports data transfers for multiple experiments, including Intensity Frontier experiments such as DUNE, enabling reliable data movement between WebDAV endpoints in Europe and the Americas.​ At CHEP 2021, we reported on the initial containerized deployment of FTS3 on OKD, the community Kubernetes distribution of Red Hat OpenShift. In this work, we present the subsequent evolution of this deployment, focusing on new operational capabilities introduced to improve scalability, robustness, and long-term maintainability.​ We describe the adoption of more secure and reproducible container build workflows, the integration of DevOps-driven operational practices, and enhancements in monitoring and automation. A key new result is the introduction of horizontal scaling and elastic resource management, allowing FTS3 components to dynamically adapt to workload variations while maintaining service reliability. We also discuss improvements in fault tolerance and operational procedures derived from production experience.​ Finally, we summarize lessons learned from operating FTS3 as a Kubernetes-native service and outline how these developments have improved the resilience and efficiency of data movement operations at Fermilab.

Munoz Flores, Victor Leopoldo [Fermilab]↗