Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data Transfer”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Enabling Low-Overhead HT-HPC Workflows at Extreme Scale using GNU Parallel

GNU Parallel is a versatile and powerful tool for process parallelization widely used in scientific computing. This paper demonstrates its effective application in high-performance computing (HPC) environments, particularly focusing on its scalability and efficiency in executing large-scale high-throughput high-performance computing (HT-HPC) workflows. Through real-world examples, we highlight GNU Parallel’s performance across various HPC workloads, including GPU computing, container-based workloads, and node-local NVMe storage. Our results on two leading supercomputers, OLCF’s Frontier and NERSC’s Perlmutter, showcase GNU Parallel’s rapid process dispatching ability and its capacity to maintain low overhead even at extreme scales. We explore GNU Parallel’s application in massive parallel file transfers using a scheduled Data Transfer Node (DTN) cluster, emphasizing its broad utility in diverse scientific workflows. Beyond its direct application as a viable workflow manager, GNU Parallel can be employed in conjunction with other workflow systems as a "last-mile" parallelizing driver and as a quick prototyping tool to design and extract parallel profiles from application executions. We then argue that the potential for GNU Parallel to transform workflow management at extreme scales is substantial, paving the way for more efficient and effective scientific discoveries.

Maheshwari, Ketan↗

Enhancing Cluster Identification in Atom Probe Tomography Data Using Transfer Learning

Atom probe tomography (APT) has enabled the direct visualization of solute clusters, providing valuable insights into material structures. This clustering is crucial for understanding the nanoscale composition and behavior of materials, which can significantly influence their mechanical and physical properties. However, the widely used clustering methods in the APT community face challenges such as subjective parametric selection and limited applicability, particularly in dealing with overlapping clusters, nested clusters, and artifacts across different scales, such as precipitates and dislocations. To address these challenges, we present a framework based on density-based cluster analysis that aims to be less dependent on user input, reproducible, and robust.

Density-based clustering↗

Portable and Adaptable Neutron Diagnostics for Advancing Fusion Energy Science Addendum

Activation detectors developed at LLNL for measuring real-time neutron fluence from fusion sources are used in the broader fusion community. The recommended fluence operating range of this diagnostic is 5x10 2 – 1x10 6 n/cm2. The upper limit on this fluence range is set by the dead time caused by data transfer between the detector and data acquisition computer. Delaying the start of counting is a possible strategy to operate these detectors in higher fluences.

42 ENGINEERING↗

Distributed Computing for the Project 8 Experiment

The Project 8 collaboration aims to measure the absolute neutrino mass or improve on the current limit by measuring the tritium beta decay electron spectrum. We present the current distributed computing model for the Project 8 experiment. Project 8 is in its second phase of data taking with a near continuous data rate of 1Gbps. The current computing model uses DIRAC (Distributed Infrastructure with Remote Agent Control) for its workflow and data management. A detailed meta-data assignment using the DIRAC File Catalog is used to automate raw data transfers and subsequent stages of data processing. The DIRAC system is deployed on containers managed using a Kubernetes cluster to provide a scalable infrastructure. A modified DIRAC Site Director provides the ability to submit jobs using Singularity on opportunistic High-Performance Computing (HPC) sites.

Distributed Computing, Kubernetes, DIRAC, Project ↗

FPGA-based computing system for processing data in size, weight, and power constrained environments

Technologies that are well-suited for use in size, weight, and power (SWAP)-constrained environments are described herein. A host controller dispatches data processing instructions to hardware acceleration engines (HAEs) of one or more field programmable gate arrays (FPGAs) and further dispatches data transfer instructions to a memory controller, such that the HAEs perform processing operations on data stored in local memory devices of the HAEs in parallel with other data being transferred from external memory devices coupled to the FPGA(s) to the local memory devices.

Napier, Matthew↗

A novel transfer learning framework for sorghum biomass prediction using UAV-based remote sensing data and genetic markers

Yield for biofuel crops is measured in terms of biomass, so measurements throughout the growing season are crucial in breeding programs, yet traditionally time- and labor-consuming since they involve destructive sampling. Modern remote sensing platforms, such as unmanned aerial vehicles (UAVs), can carry multiple sensors and collect numerous phenotypic traits with efficient, non-invasive field surveys. However, modeling the complex relationships between the observed phenotypic traits and biomass remains a challenging task, as the ground reference data are very limited for each genotype in the breeding experiment. In this study, a Long Short-Term Memory (LSTM) based Recurrent Neural Network (RNN) model is proposed for sorghum biomass prediction. The architecture is designed to exploit the time series remote sensing and weather data, as well as static genotypic information. As a large number of features have been derived from the remote sensing data, feature importance analysis is conducted to identify and remove redundant features. A strategy to extract representative information from high-dimensional genetic markers is proposed. To enhance generalization and minimize the need for ground reference data, transfer learning strategies are proposed for selecting the most informative training samples from the target domain. Consequently, a pre-trained model can be refined with limited training samples. Field experiments were conducted over a sorghum breeding trial planted in multiple years with more than 600 testcross hybrids. The results show that the proposed LSTM-based RNN model can achieve high accuracies for single year prediction. Further, with the proposed transfer learning strategies, a pre-trained model can be refined with limited training samples from the target domain and predict biomass with an accuracy comparable to that from a trained-from-scratch model for both multiple experiments within a given year and across multiple years.

36 MATERIALS SCIENCE↗

Complete and Correct Transfer of Information (CACTI)

Many distributed systems, file transfer mechanisms, and message passing systems offer reliability mechanisms such as acknowledgements, retries, and durability. While these tools may be “good enough” for their typical use cases, they may not offer sufficient coverage for the wide range of faults that impact data transfers and communication. A gap in the reliability measures may lead to some small amount of data loss. Some high-consequence systems cannot tolerate the loss or corruption of even a single record. We present seven principles that will counter a wide range of faults and protect against data loss and corruption. These principles bring together lessons learned from a wide range of technologies and can inform appropriate system design and application usage. These principles will help readers reason on how prevent data loss in a multi-hop pipeline and how to properly use tools that may have a deficiency in reliability.

97 MATHEMATICS AND COMPUTING↗

Engaging the Regulatory Community to Aid Environmental Consenting/Permitting Processes for Marine Renewable Energy

Regulators involved in consenting/permitting marine renewable energy (MRE) have faced multiple challenges due to relatively new, unfamiliar technologies and uncertainty surrounding potential environmental impacts. This has resulted in slow progress for the MRE industry, including long consenting timeframes and extensive and expensive monitoring requirements, which increase financial risk for investors. OES-Environmental has surveyed regulators internationally to understand their key knowledge gaps and perspectives to support the development of the MRE industry. From the results of these surveys a data transferability process and a risk retirement pathway have been developed to assess consenting and monitoring requirements in proportion to risk. A tool for discovering existing data sets by using an online matrix has been developed, along with training materials, regulatory guidance documents, and a strategic outreach plan to engage regulators and advisers. his engagement and the application of these products should lead to a better understanding of the environmental effects of marine energy, and more efficient consenting processes.

16 TIDAL AND WAVE POWER↗

Leveraging History to Predict Infrequent Abnormal Transfers in Distributed Workflows

Scientific computing heavily relies on data shared by the community, especially in distributed data-intensive applications. This research focuses on predicting slow connections that create bottlenecks in distributed workflows. In this study, we analyze network traffic logs collected between January 2021 and August 2022 at the National Energy Research Scientific Computing Center (NERSC). Based on the observed patterns, we define a set of features primarily based on history for identifying low-performing data transfers. Typically, there are far fewer slow connections on well-maintained networks, which creates difficulty in learning to identify these abnormally slow connections from the normal ones. We devise several stratified sampling techniques to address the class-imbalance challenge and study how they affect the machine learning approaches. Our tests show that a relatively simple technique that undersamples the normal cases to balance the number of samples in two classes (normal and slow) is very effective for model training. This model predicts slow connections with an F1 score of 0.926.

97 MATHEMATICS AND COMPUTING↗

Controlling the helicity of light by electrical magnetization switching

Controlling the intensity of emitted light and charge current is the basis of transferring and processing information. By contrast, robust information storage and magnetic random-access memories are implemented using the spin of the carrier and the associated magnetization in ferromagnets. In this study, the missing link between the respective disciplines of photonics, electronics and spintronics is to modulate the circular polarization of the emitted light, rather than its intensity, by electrically controlled magnetization. Here we demonstrate that this missing link is established at room temperature and zero applied magnetic field in light-emitting diodes through the transfer of angular momentum between photons, electrons and ferromagnets. With spin-orbit torque a charge current generates also a spin current to electrically switch the magnetization. This switching determines the spin orientation of injected carriers into semiconductors, in which the transfer of angular momentum from the electron spin to photon controls the circular polarization of the emitted light. The spin-photon conversion with the nonvolatile control of magnetization opens paths to seamlessly integrate information transfer, processing and storage. Our results provide substantial advances towards electrically controlled ultrafast modulation of circular polarization and spin injection with magnetization dynamics for the next-generation information and communication technology, including space-light data transfer. The same operating principle in scaled-down structures or using two-dimensional materials will enable transformative opportunities for quantum information processing with spin-controlled single-photon sources, as well as for implementing spin-dependent time-resolved spectroscopies.

74 ATOMIC AND MOLECULAR PHYSICS↗

2024 OES-Environmental 2024 State of the Science Report, Chapter 8: Marine Renewable Energy Data and Information Systems

As the marine renewable energy (MRE) sector grows, large amounts of environmental and technical data and information are being collected. When these data and information are openly available, they can be used to guide research and development, inform responsible siting and consenting of projects, and increase stakeholder understanding through transparency. For example, quality environmental data collected during the siting, consenting, construction, operation, and decommissioning of MRE projects can all play key roles in better characterizing baseline conditions, developing effective monitoring and mitigation strategies, and retiring environmental risks through data transferability (see Chapter 6). Ensuring that these data and information are easily discoverable and accessible will help the MRE sector make informed decisions and coexist in an increasingly busy ocean environment.

16 TIDAL AND WAVE POWER↗

DoCeph: DPU-Offloaded Messaging in Ceph for Reduced Host CPU Utilization

Ceph is a widely used distributed object store, but its messenger layer imposes substantial CPU overhead on the host. To address this limitation, we propose DoCeph, a DPU-offloaded storage architecture for Ceph that disaggregates the system by offloading the communication-intensive messaging component to the DPU while retaining the storage backend on the host. The DPU efficiently manages communication, using lightweight RPC for metadata operations and DMA for data transfer. Moreover, DoCeph introduces a pipelining technique that overlaps data transmission with buffer preparation, mitigating hardware-imposed transfer size limitations. We implemented DoCeph on a Ceph cluster with NVIDIA BlueField-3 DPUs. Evaluation results indicate that DoCeph cuts host CPU usage by up to 92% while sustaining stable throughput and providing larger performance benefits for object writes over 1 MB.

Park, Kuri [Sogang University]↗

Pressurizer dynamic model and emulated programmable logic controllers for nuclear power plants cybersecurity investigations

This work demonstrates the functionality of the pressurizer using a fast-running three regions, non-equilibrium model and control programs of emulated PLCs. The state variables from an integrated model of primary loop in a representative PWR plant are communicated to the pressurizer’s emulated PLCs using a synchronized data transfer function. In turn, the PLCs communicate back instructions to the pressurizer model to adjust the pressure and water level to remain within preprogramed setpoints. Pressurizer model simultaneously solves the coupled mass and energy conservation equations in the saturated vapor and saturated and subcooled liquid regions using fixed step solver. Results demonstrate the response of the pressurizer model linked to an emulated pressure and water level PLCs in a simulated transient involving surge-in and surge-out events. The 50 ms response delay time for the emulated PLC insignificantly affects operation and predictions of the pressurizer model.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Applying an Oriented Divergence Theorem to Swept Face Remap

Here we present a novel oriented divergence theorem and apply the results to a swept face remap method (conservative data transfer between two meshes) in arbitrary Langrangian–Eulerian hydrodynamics. In our setting, we compute the material flux along swept regions between corresponding faces in the source and target meshes. Since the swept region may add material, subtract material, or do both when it intersects itself, we cannot apply the conventional divergence theorem without accounting for orientation and self-overlaps. In this work, we encode the swept region orientation and geometry with a map from the unit n -dimensional cube, and then apply an oriented analog of divergence theorem to compute the material flux. We present efficient implementation strategies for the presented method. We also provide numerical evidence supporting our results and discuss extensions to more general mesh topologies.

97 MATHEMATICS AND COMPUTING↗

A Scalable PDC Placement Technique for Fast and Resilient Monitoring of Large Power Grids

The wide-area measurement system (WAMS) is a key enabler of real-time monitoring of power grids. The essential goals of WAMS design are fast and resilient data transfer from phasor measurement units (PMU) to phasor data concentrators (PDC). We propose a scalable two-stage PDC placement technique for minimizing the end-to-end delay while maintaining resiliency. In the prescreening stage, the plausible candidates of PDC configurations are identified based on a graph theory-based multi-median function (MMF). Here, in this article, a computationally efficient meta-heuristic algorithm is used to address scalability. In the candidate selection stage, two different algorithms, namely, Suurballe's and Dijkstra's, are employed to identify the best of those plausible PDC configurations as the final design. This technique not only minimizes the hop paths between PMUs and PDCs, but also ensures network resiliency against single PMU, PDC, or communication link failure by incorporating the roles of PMUs in power grid observability into routing policy. Simulation results on the IEEE 57-bus test power system and the 2000-bus test power system demonstrate the effectiveness and scalability of the proposed technique.

24 POWER TRANSMISSION AND DISTRIBUTION↗

In situ compression artifact removal in scientific data using deep transfer learning and experience replay

The massive amount of data produced during simulation on high-performance computers has grown exponentially over the past decade, exacerbating the need for streaming compression and decompression methods for efficient storage and transfer of this data---key to realizing the full potential of large-scale computational science. Lossy compression approaches such as JPEG when applied to scientific simulation data realized as a stream of images can achieve good compression rates but at the cost of introducing compression artifacts and loss of information. This paper develops a unified framework for in situ compression artifact removal in which the fully convolutional neural network architectures are combined with scalable training, transfer learning, and experience replay to achieve superior accuracy and efficiency while significantly decreasing the storage footprint as compared with the traditional optimization-based approaches. We demonstrate the proposed approach and compare it with compressed sensing postprocessing and other baseline deep learning models using climate simulations and nuclear reactor simulations, both of which are driven by hyperbolic partial differential equations. Our approach when applied to remove the compression artifacts on the JPEG-compressed nuclear reactor simulation data (using a transfer-trained model that was pretrained on the climate simulation data and updated incrementally as the nuclear reactor simulation progressed), achieved a significant improvement---mean peak signal-to-noise ratio of 42.438 as compared with 27.725 obtained with the compressed sensing approach.

97 MATHEMATICS AND COMPUTING↗

FIRM: federated image reconstruction using multimodal tomographic data

Here, we propose a federated algorithm for reconstructing images using multimodal tomographic data sourced from dispersed locations, addressing the challenges of traditional unimodal approaches that are prone to noise and reduced image quality, as well as the limitations of centralized multimodal approaches that require extensive data transfer, leading to significant communication overhead, storage demands, and potential data privacy concerns. Our approach formulates a joint inverse optimization problem incorporating multimodality constraints and solves it in a federated framework through local gradient computations complemented by lightweight central operations, thereby ensuring data decentralization. Leveraging the connection between our federated algorithm and the quadratic penalty method, we introduce an adaptive step-size rule with guaranteed sublinear convergence. Numerical results demonstrate superior computational efficiency and improved image reconstruction quality compared to existing approaches.

federated algorithm↗