Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data Transfer”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Unprecedented cloud resolution in a GPU-enabled full-physics atmospheric climate simulation on OLCF’s summit supercomputer

Clouds represent a key uncertainty in future climate projection. While explicit cloud resolution remains beyond our computational grasp for global climate, we can incorporate important cloud effects through a computational middle ground called the Multi-scale Modeling Framework (MMF), also known as Super Parameterization. This algorithmic approach embeds high-resolution Cloud Resolving Models (CRMs) to represent moist convective processes within each grid column in a Global Climate Model (GCM). The MMF code requires no parallel data transfers and provides a self-contained target for acceleration. This study investigates the performance of the Energy Exascale Earth System Model-MMF (E3SM-MMF) code on the OLCF Summit supercomputer at an unprecedented scale of simulation. Hundreds of kernels in the roughly 10K lines of code in the E3SM-MMF CRM were ported to GPUs with OpenACC directives. A high-resolution benchmark using 4600 nodes on Summit demonstrates the computational capability of the GPU-enabled E3SM-MMF code in a full physics climate simulation.

58 GEOSCIENCES↗

WATTS: Workflow and template toolkit for simulation

Modeling and simulation in many science and engineering domains often involves the execution and/or iteration of a sequence of applications, with data transfer between applications typically required. These applications often do not have a formal application programming interface (API). Instead, executing an application requires first writing a text-based input file, the format of which is typically defined in a user’s manual. While text-based input files are suitable for simple one-off calculations, they can become cumbersome if a user wants to execute the applications multiple times and systematically vary input parameters, especially when a complex workflow is involved. In this case, they must resort to either manually making changes in the input file or developing their own script that modifies the input file and executes the application. Depending on the format of the input file, writing such a script can be a non-trivial and error-prone task.

97 MATHEMATICS AND COMPUTING↗

Analysis of the performance of a hybrid CPU/GPU 1D2D coupled model for real flood cases

Coupled 1D2D models emerged as an efficient solution for a two-dimensional (2D) representation of the floodplain combined with a fast one-dimensional (1D) schematization of the main channel. At the same time, high-performance computing (HPC) has appeared as an efficient tool for model acceleration. In this work, a previously validated 1D2D Central Processing Unit (CPU) model is combined with an HPC technique for fast and accurate flood simulation. Due to the speed of 1D schemes, a hybrid CPU/GPU model that runs the 1D main channel on CPU and accelerates the 2D floodplain with a Graphics Processing Unit (GPU) is presented. Since the data transfer between sub-domains and devices (CPU/GPU) may be the main potential drawback of this architecture, the test cases are selected to carry out a careful time analysis. Here, the results reveal the speed-up dependency on the 2D mesh, the event to be solved and the 1D discretization of the main channel. Additionally, special attention must be paid to the time step size computation shared between sub-models. In spite of the use of a hybrid CPU/GPU implementation, high speed-ups are accomplished in some cases.

54 ENVIRONMENTAL SCIENCES↗

Understanding HPC Benchmark Performance on Intel Broadwell and Cascade Lake Processors

Hardware platforms in high performance computing are constantly getting more complex to handle even when considering multicore CPUs alone. Numerous features and configuration options in the hardware and the software environment that are relevant for performance are not even known to most application users or developers. Microbenchmarks, i.e., simple codes that fathom a particular aspect of the hardware, can help to shed light on such issues, but only if they are well understood and if the results can be reconciled with known facts or performance models. The insight gained from microbenchmarks may then be applied to real applications for performance analysis or optimization. In this paper we investigate two modern Intel x86 server CPU architectures in depth: Broadwell EP and Cascade Lake SP. We highlight relevant hardware configuration settings that can have a decisive impact on code performance and show how to properly measure on-chip and off-chip data transfer bandwidths. The new victim L3 cache of Cascade Lake and its advanced replacement policy receive due attention. Finally we use DGEMM, sparse matrix-vector multiplication, and the HPCG benchmark to make a connection to relevant application scenarios.

97 MATHEMATICS AND COMPUTING↗

A Systems Engineering Analysis of National Ignition Facility Industrial Controls Systems and Safety Interlock Systems Remote Input/Output Networking Migration from ControlNet to EtherNet/IP

The ControlNet industrial communications protocol and modules used in the Industrial Control System (ICS) and Safety Interlock System (SIS) at the National Ignition Facility (NIF) are no longer necessary and the ICS and SIS would be better served by migrating the communications structure to use EtherNet/Industrial Protocol (IP) and EtherNet bridge modules instead. By the admission of the vendor of ControlNet hardware, Rockwell Automation, in literature by Bill Petro [1], “Moving forward, customers will be able to optimize their asset utilization better using EtherNet/IP protocol than with ControlNet.” The NIF is one of the key elements of the Inertial Confinement Fusion (ICF) program at Lawrence Livermore National Laboratory (LLNL), a federally funded research and development center (FFRDC). The NIF contains the systems and provides the operational capacity to perform ICF, high energy density (HED), and discovery science experiments utilizing 192 individual beamlines, a host of diagnostics, and all the industrial systems required to facilitate these beamlines and diagnostics. The industrial systems are governed by the ICS and SIS, with the ICS providing control and the SIS providing monitoring and permissives. Construction on the NIF began in 1997 and was certified complete in 2009 and, as a result, the ICS and SIS were developed during this time using the tools that were available then. This includes the communications structure and protocols for these systems, much of which was, and still is, ControlNet. ControlNet, particularly during the time that the ICS and SIS were being built, has several attractive features. ControlNet hardware is exclusive to Rockwell Automation, which was the automation hardware chosen for the ICS and SIS. One feature that could be considered an advantage or a disadvantage depending on the communication needs of the system is that ControlNet also utilizes no active network components, excluding repeaters which are not always necessary. According to the architect of the ICS system at the NIF, Gordon Lau, one of the most attractive features of the ControlNet protocol during development of the ICS and SIS was that it is deterministic, providing timing of data transfer that is executed exactly as it is defined by the developer.

42 ENGINEERING↗

Remapping of Data Between One-Dimensional Meshes

In this report we present two approaches to data remapping between one-dimensional meshes implemented with the c++ programming language. Our goal was to test the performance of two search algorithms, linear and binary, and verify the accuracy of our implementations of the two methods. We first introduce the concept of data remap and meshing components, as well as their various uses. We then delve into the differences between point-wise and conservative remap, the algorithms used in the implementations, and lastly confirm the implementations work as intended when given various inputs. We expect that, after profiling, the binary search algorithm will be more efficient than the linear algorithm for sorted sets of data, the point-wise remap implementation to accurately approximate the data transfer between two meshes, and the conservative remap implementation to conserve the area underneath the curve of two distinct meshes.

97 MATHEMATICS AND COMPUTING↗

Scalable Skipper-CCD Analog Readout Electronics (SSCARE) for the Massive Multichannel OSCURA Experiment

We present a multiplexed analog readout electronics system for Skipper-CCDs based on the MIDNA ASIC. It allows for subelectron noise-level operation while maintaining a minimal number of acquisition channels. In addition, it requires low-disk storage and low-bandwidth data transfer with zero added multiplexing time during the simultaneous operation of thousands of channels. We describe the implementation of such a system in a new instrument composed by 16 sensors operated with a two-stage analog multiplexed readout scheme. The instrument is a part of the R&D effort of the OSCURA experiment.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Improving Signal-to-Noise Ratio (SNR) for Readout Signals Using Adaptive Filters on Reconfigurable Controls Hardware

This study investigates the optimization of Signal-to-Noise Ratio (SNR) in superconducting quantum computing readout signals through adaptive filtering. Quantum computing technology has the potential to revolutionize various fields by delivering exponential speedup in solving certain computational problems. However, the technology's practical implementation is hindered by the difficulty of extracting clean, reliable signals during the readout phase, with various sources of noise presenting a significant barrier to clean signals. This noise, often present in readout profiles due to imperfect isolation, degrades the system's overall SNR, thus impeding the ability to extract the quantum state accurately. The research leverages the power of adaptive filtering to improve the SNR of quantum computing readout signals. Specifically, an adaptive filter is implemented in a PYNQ overlay on an FPGA, and eventually will be connected to a quantum computing system. The system models the n oise with a Least Mean Squares (LMS) adaptive filter, and then subtracts the estimated noise from the received signal to improve the SNR. A Direct Memory Access (DMA) channel is used to handle the signal processing, delivering efficient, high-speed data transfer between the PYNQ system and the hardware. The study explores the benefits of this adaptive filtering technique, potentially providing a significant contribution to practical and fast quantum computing.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Low-Flow Marine Hydrokinetic Turbine for Small Autonomous Unmanned Mobile Recharge Stations

A prototype low-flow marine current turbine for deployment from a small unmanned mobile floating platform has been developed for autonomously seeking and harnessing tidal/coastal currents. The support platform is an unmanned surface vehicle (USV), in the form of a catamaran with two electric outboard motors and with capabilities for autonomous navigation. The USV utilized is a WAM-V 16 vehicle that has been developed separately with support from the Office of Naval Research (ONR) [1]. The marine current turbine is based on a freestream waterwheel (FSWW), also known as an undershot waterwheel (FSWW), mounted on the stern of the USV. The concept of operation involves the USV autonomously navigating to a designated marine current resource. Upon arrival, the USV anchors itself, aligns with the current, and deploys the FSWW turbine using a custom cable-lift mechanism. The turbine harnesses the local current, and an onboard power-take-off (PTO) device converts the mechanical energy into electricity, which is stored in an onboard battery bank. When energy harvesting is completed, the turbine and the anchor are retrieved and the USV navigates to a selected location. These unmanned at-sea platforms can provide power to other unmanned maritime systems. Specifically, in this project, the power generated onboard can be used to charge aerial drones via a custom flight deck that has been developed for the USV. The recharging capabilities offered by a fleet of such strategically placed recharging stations can significantly benefit aerial drones operating in the maritime domain by eliminating the need to travel back and forth to land or ship based charging stations. The project has resulted in the development of subcomponents, including the FSWW turbine, a novel PTO, an automated anchoring system for the USV, an automated turbine deployment system, and a flight deck with capabilities onboard the USV for landing, direct-contact charging and takeoff of aerial drones. The design and development of these subsystems have culminated in the overall prototype marine hydrokinetic platform (MHK Platform, Fig. 1). Comprehensive lab and field testing have been conducted to validate the functionality and performance of the platform and its components. The project demonstrates the potential for autonomous, unmanned systems to harness renewable energy from marine currents, and provide sustainable power solutions for maritime applications such as coastal surveillance and environmental monitoring; shoreline mapping; search and rescue; oceanographic research; inspection and maintenance of offshore energy installations like wind turbines and oil rigs; oil spill response; maritime disaster response; and aerial surveys, as well as facilitation of data transfer drones and shore stations.

16 TIDAL AND WAVE POWER↗

Implementing One Sided Partitioned Communication in Open MPI

This report introduces partitioned communication, a new MPI 4.0 interface that enables early bird communication by overlapping communication and computation. By partitioning messages into smaller sub-messages, MPI can start partial data transfers early. Performance studies show that the RMA implementation outperforms the Persistent implementation, despite some constraints. This report details a new opt-in RMA implementation, offering a high-performance option for partitioned communication that imposes some additional limitations.

97 MATHEMATICS AND COMPUTING↗

CalTestBed - Delphire - Testing and Evaluation of Delphire Sentinel System (CRADA Final Report)

The Delphire Sentinel is a modular fire detection and communications system operating as a mobile field unit, with low voltage DC power supplied by onboard photovoltaics (PV) and batteries. The Sentinel addresses several aspects of fire detection, communications and data analysis. The Sentinel's mobility enables it to be rapidly deployed and operate independently of existing power and communications networks. The duration of independent operation depends critically on the energy consumption of the systems and performance of the onboard PV and battery. The purpose of this testing is to ascertain the power draw and energy consumption of the Delphire Sentinel prototype system under several operational states, including various data transfer packet sizes, transmission time and frequencies, and communication pathways (Wi-Fi, cellular, satellite) expected to be encountered in field deployments. It will also include procedures to test the ability of the Sentinel to operate for extended periods without loss of functionality. Based on results from energy and power measurements, and anticipated duty cycles in field deployments, we will model annual system autonomy (e.g. loss of load probability) for off-grid operation in representative locations.

47 OTHER INSTRUMENTATION↗

Software Quality Assurance for the MOOSE-Based Open-Source Multiphysics Code Cardinal - An Expanded CI Testing Suite

Cardinal is a wrapping of the GPU-oriented spectral element Computational Fluid Dynamics (CFD) code NekRS and the Monte Carlo particle transport code OpenMC within the Multiphysics Object-Oriented Simulation Environment (MOOSE). Cardinal provides high-resolution thermal-hydraulics and/or radiation transport feedback to MOOSE multiphysics simulations. Multiphysics feedback is implemented in a geometry-agnostic manner which eliminates the need for rigid one-to-one mappings. A generic data transfer implementation also allows NekRS and OpenMC to couple to any MOOSE application, enabling a broad set of multiphysics capabilities. Cardinal simulations can also leverage combinations of MPI, OpenMP, and GPU resources. Cardinal continuous development and improvement efforts have led to the software being considered as a high-fidelity design and licensing tool for key areas of nuclear reactor relevant physics, including neutron transport, fluid flow, heat transfer, and mechanical processes. The fast development and expansion of the software from a pure R&D framework towards its application in the nuclear industry and regulation require a focus on developing, enhancing and, maintaining Cardinal’s software quality through strict adherence to a Software Quality Assurance (SQA) framework and SQA program. To facilitate compliance with SQA standards, the Cardinal SQA Program has been initiated during Fiscal Year 2023 (FY23). During the development of the Cardinal SQA Program, multiple gaps have been identified. These gaps are primarily related to model verification and code pedigree as they relate to the use of Cardinal as a safety analysis tool. These gaps have been captured in a report published in 2023. A second report highlighted the progress made during Fiscal Year 2024 (FY24) and described Argonne’s effort to document and integrate software verification within Cardinal’s software development process. This report documents a snapshot of the verification test cases currently available for Cardinal and NekRS in their assimilation into a Continuous Integration (CI) platform. Following the CI practice permits the integrating of source code changes frequently and ensuring that the integrated codebase clears the verification testing for the software. It should be noted that the SQA program itself, including the program plans, procedures, configuration management, and testing strategies, need to be developed in a future step of this task.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Progress Towards NQA-1 for Cardinal in FY25

Cardinal is a wrapping of the GPU-oriented spectral element Computational Fluid Dynamics (CFD) code NekRS and the Monte Carlo particle transport code OpenMC within the Multiphysics Object-Oriented Simulation Environment (MOOSE). Cardinal provides high-resolution thermal-hydraulics and/or radiation transport feedback to MOOSE multiphysics simulations. Multiphysics feedback is implemented in a geometry-agnostic manner which eliminates the need for rigid one-to-one mappings. A generic data transfer implementation also allows NekRS and OpenMC to couple to any MOOSE application, enabling a broad set of multiphysics capabilities. Cardinal simulations can also leverage combinations of MPI, OpenMP, and GPU resources. Cardinal continuous development and improvement efforts have led to the software being considered as a high-fidelity design and licensing tool for key areas of nuclear reactor relevant physics, including neutron transport, fluid flow, heat transfer, and mechanical processes. The fast development and expansion of the software from a pure R&D framework towards its application in the nuclear industry and regulation require a focus on developing, enhancing,and maintaining Cardinal’s software quality through strict adherence to a Software Quality Assurance (SQA) framework and SQA program. To facilitate compliance with SQA standards, the Cardinal SQA Program was initiated during Fiscal Year 2023 (FY23). During the development of the Cardinal SQA Program, multiple gaps have been identified. These gaps are primarily related to model verification and code pedigree as they relate to the use of Cardinal as an analysis tool. These gaps were captured in a report published in 2023. A second report highlighted the progress made during Fiscal Year 2024 (FY24) and described Argonne’s effort to document and integrate software verification within Cardinal’s software development process. This report documents the progress made towards NQA-1 for Cardinal in the Fiscal Year 2025 (FY25). All cases in the expanded Continuous Integration (CI) suite of NekRS are included in this report which test the solvers and modules available in NekRS exhaustively. The NekRS tests are integrated with the Cardinal CI suite and made available in publicly accessible Github documentation. Following the CI practice permits integrating of source code changes frequently and ensuring that the integrated codebase clears the verification testing for the software. Also in this report is a brief overview of the development of the Cardinal Software Quality Assurance Plan (SQAP) that was done in FY25, though it should be noted that the rest of the documentation for the SQA program needs to be developed in a future step of this task.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Casing Annulus Monitoring of CO 2 Injection Using Wireless Autonomous Distributed Sensor Networks

Effective and secure carbon subsurface storage, involving the deep underground injection of CO 2 into geological formations where it is permanently trapped, is paramount to mitigating CO 2 emissions (Figure I). Ensuring the integrity of these storage sites and detecting potential leakage through the casing annulus necessitates robust monitoring. This work provides the first integrated demonstration of a wireless casing-annulus monitoring architecture that can operate in highly attenuating cement-brine environments relevant to CO 2 storage. This project focused on developing and validating a novel sensor system for integration with autonomous monitoring near the cement reservoir interface. The goal was a fully integrated Technology Readiness Level (TRL) 4/5 field validation of a distributed wireless intelligent sensor system providing real-time, direct subsurface formation measurements to enhance fluid movement monitoring in the cemented casing annulus. Achieving this objective required the development and integration of 1) wireless autonomous microsensor technology by California Institute of Technology (Caltech); 2) sensor packaging and emplacement technology by Research Triangle Institute (RTI); and 3) smart well completions using wireless active casing collars and NOV pipe by the Sandia National Lab (SNL). The collaboration with the Caltech team in this project aimed to develop millimeter-scale radio frequency identification (RFID) sensors capable of detecting CO 2 , pH, and/or methane levels. These sensors are engineered to be impervious to fluids, allowing them to be mixed with cement and installed within the casing annulus. They operate using RFID protocols at frequencies of 902–928 MHz for both power and communication. A Sandia National Laboratories’ team engaged their expertise in the development of a Smart Collar system designed for the wireless data collection from these RFID sensors embedded in the cement annulus and transmission of this information to the ground surface via IntelliPipe/IntelliServ NOV drill pipe. This is accomplished through inductive coupling at the collar, which facilitates data transfer through each segment of the pipe. Because the system cannot transmit a direct current signal to power the Smart Collar, both power and communication were implemented using alternating current and electromagnetic signals at varying frequencies. Furthermore, the developed microsensor technology had to be demonstrated and validated in comparison with reference transducer measurements in a field test site at The University of Texas at Austin (UT-Austin). Although the full sensor suite did not reach field-deployment readiness, the system-level integration achieved in this project establishes a validated pathway for future incorporation of advanced microsensors.

47 OTHER INSTRUMENTATION↗

Cloud-Based Demonstration of the Eastern Interconnection Situational Awareness Monitoring System (ESAMS)

This report describes a cloud-based implementation and field demonstration of the Eastern Interconnection Situational Awareness and Monitoring System (ESAMS). ESAMS was developed to support the detection and source localization of forced oscillations using synchrophasor measurements from tie-lines connecting areas served by different reliability coordinators (RCs), so that RCs could better coordinate their response to wide-area events. A previous effort had identified deployment barriers associated with hosting shared situational awareness tools at a single RC. To address these barriers, ESAMS was migrated to Amazon Web Services and evaluated in a six-month field demonstration. ISO New England (ISO-NE) and PJM streamed data to the platform using AWS Direct Connect and a site-to-site VPN, respectively. The resulting multi-utility measurement footprint enabled regional source localization across major portions of the U.S. Eastern Interconnection and supported routine identification of oscillation events. During the final three months of the trial, 24 events above 2 MW/MVAR were detected. The largest detected oscillation approached a 25 MW peak-to-peak amplitude, and the longest persisted intermittently for more than 11 hours. The demonstration also assessed operational considerations—including data transfer volumes, end-to-end latency, and cloud computing costs—and found that network and compute requirements were modest relative to typical cloud capabilities while providing performance comparable to prior on-premises deployments. Overall, the results indicate that cloud hosting can provide a practical path to shared interconnection-wide oscillation monitoring. The cloud ESAMS demonstration establishes a foundation for broader utility participation and for building future wide-area analytics that leverage measurements across organizational boundaries.

Follum, James D.↗

Optimizing Metadata Exchange: Leveraging DAOS for ADIOS Metadata I/O

In HPC I/O middleware like the Adaptable I/O System (ADIOS) often mediates data transfers between applications. The metadata I/O generated by such systems often presents significant scaling and performance limitations. This work seeks improvement opportunities for metadata I/O by leveraging the DAOS storage systems, a recent storage system solution deployed on high-end systems such as the Aurora supercomputer. We investigate the tradeoffs and the design space for integrating I/O engines for the ADIOS middleware based on the different storage mechanisms supported by DAOS. We present a new DAOS-Array-ChunkSize-aligned engine which provides up to 2.3× improved performance than when using the existing DAOS-POSIX interface, without requiring any application modifications.

Venkatesh, Ranjan Sarpangala↗

Optimization of Asynchronous Communication Operations through Eager Notifications

UPC++ is a C++ library implementing the Asynchronous Partitioned Global Address Space (APGAS) model. We propose an enhancement to the completion mechanisms of UPC++ used to synchronize communication operations that is designed to reduce overhead for on-node operations. Our enhancement permits eager delivery of completion notification in cases where the data transfer semantics of an operation happen to complete synchronously, for example due to the use of shared-memory bypass. This semantic relaxation allows removing significant overhead from the critical path of the implementation in such cases. We evaluate our results on three different representative systems using a combination of microbenchmarks and five variations of the the HPCChallenge RandomAccess benchmark implemented in UPC++ and run on a single node to accentuate the impact of locality. We find that in RMA versions of the benchmark written in a straightforward manner (without manually optimizing for locality), the new eager notification mode can provide up to a 25% speedup when synchronizing with promises and up to a 13.5x speedup when synchronizing with conjoined futures. We also evaluate our results using a graph matching application written with UPC++ RMA communication, where we measure overall speedups of as much as 11% in single-node runs of the unmodified application code, due to our transparent enhancements.

Kamil, Amir↗

Defining quantum-ready primitives for hybrid HPC-QC supercomputing: a case study in Hamiltonian simulation

As computational demands in scientific applications continue to rise, hybrid high-performance computing (HPC) systems integrating classical and quantum computers (HPC-QC) are emerging as a promising approach to tackling complex computational challenges. One critical area of application is Hamiltonian simulation, a fundamental task in quantum physics and other large-scale scientific domains. This paper investigates strategies for quantum-classical integration to enhance Hamiltonian simulation within hybrid supercomputing environments. By analyzing computational primitives in HPC allocations dedicated to these tasks, we identify key components in Hamiltonian simulation workflows that stand to benefit from quantum acceleration. To this end, we systematically break down the Hamiltonian simulation process into discrete computational phases, highlighting specific primitives that could be effectively offloaded to quantum processors for improved efficiency. Our empirical findings provide insights into system integration, potential offloading techniques, and the challenges of achieving seamless quantum-classical interoperability. We assess the feasibility of quantum-ready primitives within HPC workflows and discuss key barriers such as synchronization, data transfer latency, and algorithmic adaptability. These results contribute to the ongoing development of optimized hybrid solutions, advancing the role of quantum-enhanced computing in scientific research.

97 MATHEMATICS AND COMPUTING↗