SEARCH · Engineering Papers
Results for “Data transfer”
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Seamless Communication Between High-Performance Computing System and Electron Microscopes for On-Demand Automated Data Transfer and Remote Control
Explore the source record for details and available documents.
Interpolation data transfer between the models before and after a partial drawdown leach in BH
It has been recognized that as cavern operations become more frequent due to oil sales, field conditions may arise which require a faster turnaround time of analysis to address potential cavern impacts. This letter describes attempts to implement a strategy of transferring an intermediate solution of a Big Hill (BH) geomechanical model from a previous finite element mesh with a specified cavern geometry, to a new mesh with a new cavern geometry created by leaching from an oil sale operation.
Big data analysis of data transfers in multi petabyte distributed storage system [Poster]
Explore the source record for details and available documents.
Experimental Validation of Crosstalk Minimization in Metallic Barriers with Simultaneous Ultrasonic Power and Data Transfer.
Abstract not provided.
Logistics of High Throughput Data Transfer in HEP
Explore the source record for details and available documents.
he The amazing powers of Generalized Moving Least Squares: Applications to PDEs, data transfer and device models.
Abstract not provided.
Data Transfers and Host/Device Communication using OneAPI for FPGA.
Abstract not provided.
Data Transfer between the Models before and after a Partial Drawdown in a Salt Cavern.
Abstract not provided.
Experimental Validation of Crosstalk Minimization in Metallic Barriers with Simultaneous Ultrasonic Power and Data Transfer.
Abstract not provided.
Field Validation of MVA Technology for Offshore CCS: Novel Ultra-High-Resolution 3D Marine Seismic Technology (P-Cable) (Final Report)
The objectives of the proposed study were to deploy and validate a specific monitoring technology, high-resolution 3D marine seismic (HR3D), appropriate for large-demonstration and commercial-scale offshore CCS sites. The project accomplished successful acquisition two HR3D seismic surveys. The first HR3D dataset was over the offshore injection site of the Tomakomai, Japan integrated pilot CCS project, which at the time of survey acquisition was actively injecting CO 2 . The first survey also represented a successful international collaboration between the DOE NETL program and Japan’s national CCS program and was the first successful acquisition and use of HR3D over an active CO 2 injection site (Meckel, Feng et al. 2019). The Tomakomai HR3D survey successfully tested a novel 4-streamer HR3D system array in which, for the first time, no cross-cable (aka “P-Cable”) was utilized and only four GeoEel streamers were used instead of the standard 12-streamer configuration. Consequently, this was not, strictly speaking, a deployment of the “P-Cable” system of (Planke and Berndt 2004) but rather a modified version, thereof, and it is the first known demonstration of the modified system configuration. One very positive outcome from the Japanese collaboration earlier in the project was the ability to learn from the Japanese how they used tail buoys with GPS to determine the position of the seismic source and receivers in time and space. Based on that experience, GCCC designed and built six GPS receivers that could be used to position the streamer receivers and the seismic source via tail buoys. A fundamental advance that was made on the original design, was the ability to directly power the tail buoy GPS units and transfer data through the streamers (i.e., vs. the batteries used at Tomakomai). The bulkiness of the GPS batteries caused drag and episodic surging of the buoys, which affected data quality by lifting up the tail end of the streamers so the receivers were not at the same depth. The units were tested onshore for accuracy and functionality, and the design was subsequently and successfully tested in marine acquisition mode during the SLP survey acquisition. The marine acquisition test and survey satisfied Subtasks 2.2.2, Novel Positioning Technology Selection and Subtask 2.2.3, Novel Positioning Technology Deployment. Results of the novel positioning technology selection (Subtask 2.2.2) were considered successful and will be incorporated in future HR3D seismic acquisition projects to reduce costs, improve deployment safety at sea, and integrate both seismic and data recording via a single data transfer through the streamers to the recording system. The project also established a permitting process through NETL NEPA compliance, which included an Environmental Assessment in a marine setting and is required for conducting these types of surveys using Federal funding. The permitting process charted a “boilerplate,” which can allow future surveys related to other funded projects to move forward more expeditiously. Future improvements that could be considered are more robust seals on the GPS module and stronger materials (especially joints) on tail buoy fabrication. These would increase fixed costs, but would be advisable and probably more economic long-term if multiple HR3D surveys are planned. Project Accomplishments include: • Pre-survey Sensitivity Study • Marine geochemistry methods and data analysis • Successful HR3D seismic dataset acquired @ Tomakomai active CO 2 injection marine site • Developed advanced seismic processing techniques • No NRMS anomalies detected in overburden; Demonstration of containment • Repeatability study • Second survey collected @ San Luis Pass, TX • 4D application using positioning techniques developed in the project for monitoring were successful
An FPGA-based hardware accelerator supporting sensitive sequence homology filtering with profile hidden Markov models
Abstract Background Sequence alignment lies at the heart of genome sequence annotation. While the BLAST suite of alignment tools has long held an important role in alignment-based sequence database search, greater sensitivity is achieved through the use of profile hidden Markov models (pHMMs). Here, we describe an FPGA hardware accelerator, called HAVAC, that targets a key bottleneck step (SSV) in the analysis pipeline of the popular pHMM alignment tool, HMMER. Results The HAVAC kernel calculates the SSV matrix at 1739 GCUPS on a $$\sim$$ ∼ $3000 Xilinx Alveo U50 FPGA accelerator card, $$\sim$$ ∼ 227× faster than the optimized SSV implementation in nhmmer . Accounting for PCI-e data transfer data processing, HAVAC is 65× faster than nhmmer’s SSV with one thread and 35× faster than nhmmer with four threads, and uses $$\sim$$ ∼ 31% the energy of a traditional high end Intel CPU. Conclusions HAVAC demonstrates the potential offered by FPGA hardware accelerators to produce dramatic speed gains in sequence annotation and related bioinformatics applications. Because these computations are performed on a co-processor, the host CPU remains free to simultaneously compute other aspects of the analysis pipeline.
Contrastive Machine Learning with Gamma Spectroscopy Data Augmentations for Detecting Shielded Radiological Material Transfers
Data analysis techniques can be powerful tools for rapidly analyzing data and extracting information that can be used in a latent space for categorizing observations between classes of data. Machine learning models that exploit learned data relationships can address a variety of nuclear nonproliferation challenges like the detection and tracking of shielded radiological material transfers. The high resource cost of manually labeling radiation spectra is a hindrance to the rapid analysis of data collected from persistent monitoring and to the adoption of supervised machine learning methods that require large volumes of curated training data. Instead, contrastive self-supervised learning on unlabeled spectra can enhance models that are built on limited labeled radiation datasets. This work demonstrates that contrastive machine learning is an effective technique for leveraging unlabeled data in detecting and characterizing nuclear material transfers demonstrated on radiation measurements collected at an Oak Ridge National Laboratory testbed, where sodium iodide detectors measure gamma radiation emitted by material transfers between the High Flux Isotope Reactor and the Radiochemical Engineering Development Center. Label-invariant data augmentations tailored for gamma radiation detection physics are used on unlabeled spectra to contrastively train an encoder, learning a complex, embedded state space with self-supervision. A linear classifier is then trained on a limited set of labeled data to distinguish transfer spectra between byproducts and tracked nuclear material using representations from the contrastively trained encoder. The optimized hyperparameter model achieves a balanced accuracy score of 80.30%. Any given model—that is, a trained encoder and classifier—shows preferential treatment for specific subclasses of transfer types. Regardless of the classifier complexity, a supervised classifier using contrastively trained representations achieves higher accuracy than using spectra when trained and tested on limited labeled data.
IRIS-DMEM: Efficient Memory Management for Heterogeneous Computing
This paper proposes an efficient data memory management approach for the Intelligent RuntIme System (IRIS) heterogeneous computing framework along with new data transfer policies. IRIS provides a task-based programming model for extreme heterogeneous computing (e.g., CPU, GPU, DSP, FPGA) with support for today's most important programming languages (e.g., OpenMP, OpenCL, CUDA, HIP, OpenACC). However, the IRIS framework either forces the programmer to introduce data transfer commands for each task or relies on suboptimal memory management for automatic and transparent data transfers. The work described here extends IRIS with novel heterogeneous memory handling and introduces novel data transfer policies by employing the Distributed data MEMory handler (DMEM) for efficient and optimal movement of data among the various computing resources. The proposed approach achieves performance gains of up to 7× for tiled LU factorization and tiled DGEMM (i.e., matrix multiplication) benchmarks. Moreover, this approach also reduces data transfers by up to 71% when compared to previous IRIS heterogeneous memory management handlers. This work compares the performance results of the IRIS framework's novel DMEM with the StarPU runtime and MAGMA math library for GPUs. Experiments show a performance gain of up to 1.95× over StarPU and 2.1× over MAGMA.
Suppressing simulation bias in multi-modal data using transfer learning
Abstract Many problems in science and engineering require making predictions based on few observations. To build a robust predictive model, these sparse data may need to be augmented with simulated data, especially when the design space is multi-dimensional. Simulations, however, often suffer from an inherent bias. Estimation of this bias may be poorly constrained not only because of data sparsity, but also because traditional predictive models fit only one type of observed outputs, such as scalars or images, instead of all available output data modalities, which might have been acquired and simulated at great cost. To break this limitation and open up the path for multi-modal calibration, we propose to combine a novel, transfer learning technique for suppressing the bias with recent developments in deep learning, which allow building predictive models with multi-modal outputs. First, we train an initial neural network model on simulated data to learn important correlations between different output modalities and between simulation inputs and outputs. Then, the model is partially retrained, or transfer learned, to fit the experiments; a method that has never been implemented in this type of architecture. Using fewer than 10 inertial confinement fusion experiments for training, transfer learning systematically improves the simulation predictions while a simple output calibration, which we design as a baseline, makes the predictions worse. We also offer extensive cross-validation with real and carefully designed synthetic data. The method described in this paper can be applied to a wide range of problems that require transferring knowledge from simulations to the domain of experiments.
Data transmission by quantum matter wave modulation
Abstract Classical communication schemes exploiting wave modulation are the basis of our information era. Quantum information techniques with photons enable future secure data transfer in the dawn of decoding quantum computers. Here we demonstrate that also matter waves can be applied for secure data transfer. Our technique allows the transmission of a message by a quantum modulation of coherent electrons in a biprism interferometer. The data is encoded in the superposition state by a Wien filter introducing a longitudinal shift between separated matter wave packets. The transmission receiver is a delay line detector performing a dynamic contrast analysis of the fringe pattern. Our method relies on the Aharonov–Bohm effect but does not shift the phase. It is demonstrated that an eavesdropping attack will terminate the data transfer by disturbing the quantum state and introducing decoherence. Furthermore, we discuss the security limitations of the scheme due to the multi-particle aspect and propose the implementation of a key distribution protocol that can prevent active eavesdropping.
GPU Direct I/O with HDF5
Exascale HPC systems are being designed with accelerators, such as GPUs, to accelerate parts of applications. In machine learning workloads as well as large-scale simulations that use GPUs as accelerators, the CPU (or host) memory is currently used as a buffer for data transfers between GPU (or device) memory and the file system. If the CPU does not need to operate on the data, then this is sub-optimal because it wastes host memory by reserving space for duplicated data. Furthermore, this “bounce buffer” approach wastes CPU cycles spent on transferring data. A new technique, NVIDIA GPUDirect Storage (GDS), can eliminate the need to use the host memory as a bounce buffer. Thereby, it becomes possible to transfer data directly between the device memory and the file system. This direct data path shortens latency by omitting the extra copy and enables higher-bandwidth. To take full advantage of GDS in existing applications, it is necessary to provide support with existing I/O libraries, such as HDF5 and MPI-IO, which are heavily used in applications. In this paper, we describe our effort of integrating GDS with HDF5, the top I/O library at NERSC and at DOE leadership computing facilities. We design and implement this integration using a HDF5 Virtual File Driver (VFD). The GDS VFD provides a file system abstraction to the application that allows HDF5 applications to perform I/O without the need to move data between CPUs and GPUs explicitly. We compare performance of the HDF5 GDS VFD with explicit data movement approaches and demonstrate superior performance with the GDS method.
Tiling Framework for Heterogeneous Computing of Matrix based Tiled Algorithms
Tiling matrix operations can improve the load balancing and performance of applications on heterogeneous computing resources. Writing a tile-based algorithm for each operation with a traditional, hand-tuned tiling approach that uses for loops in C/C++ is cumbersome and error prone. Moreover, it must enable and support the heterogeneous memory management of data objects and also explore architecture-supported, native, tiled-data transfer APIs instead of copying the tiled data to continuous memory before the data transfer. The tiling framework provides a tiled data structure for heterogeneous memory mapping and parameterization to a heterogeneous task specification API. We have integrated our tiled framework into MatRIS (Math kernels library using IRIS). IRIS is a heterogeneous run-time framework with a heterogeneous programming model, memory model, and task execution model. Experiments reveal that the tiled framework for BLAS operations has improved the programmability of tiled BLAS and improved performance by ~20% when compared against the traditional method that copies the data to continuous memory locations for heterogeneous computing.