Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data transfer”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Ocelot: An Interactive, Efficient Distributed Compression-As-a-Service Platform With Optimized Data Compression Techniques

Large volumes of data generated by scientific simulations, genome sequencing, and other applications need to be moved among clusters for data collection/analysis. Data compression techniques have effectively reduced data storage and transfer costs. However, users' requirements on interactively controlling both data quality and compression ratios are non-trivial to fulfill. Here, we propose a novel Compression-as-a-Service (CaaS) platform called Ocelot with four important contributions: (1) It offers real-time visualization, interactive compression, and transfer of scientific datasets. (2) It incorporates new strategies for compressing diverse types of datasets more effectively than traditional methods. (3) It provides an effective method for estimating the compression ratio and execution time of compression tasks. (4) Experiments on multiple real-world datasets on geographically distributed computers show that Ocelot can significantly improve data transfer efficiency with a performance gain of more than 10x in computing clusters with relatively slow networks.

compression as a service (CaaS)↗

Training data selection for accuracy and transferability of interatomic potentials

Abstract Advances in machine learning (ML) have enabled the development of interatomic potentials that promise the accuracy of first principles methods and the low-cost, parallel efficiency of empirical potentials. However, ML-based potentials struggle to achieve transferability, i.e., provide consistent accuracy across configurations that differ from those used during training. In order to realize the promise of ML-based potentials, systematic and scalable approaches to generate diverse training sets need to be developed. This work creates a diverse training set for tungsten in an automated manner using an entropy optimization approach. Subsequently, multiple polynomial and neural network potentials are trained on the entropy-optimized dataset. A corresponding set of potentials are trained on an expert-curated dataset for tungsten for comparison. The models trained to the entropy-optimized data exhibited superior transferability compared to the expert-curated models. Furthermore, the models trained to the expert-curated set exhibited a significant decrease in performance when evaluated on out-of-sample configurations.

36 MATERIALS SCIENCE↗

FTS3: Data Movement Service in containers deployed in OKD [Slides]

The File Transfer Service (FTS3) is a data movement service developed at CERN which is used to distribute the majority of the Large Hadron Collider's data across the Worldwide LHC Computing Grid (WLCG) infrastructure. At Fermilab, we have deployed FTS3 instances for Intensity Frontier experiments (e.g. DUNE) to transfer data in America and Europe, using a container-based strategy. In this article we summarize our experience building docker images based on work from the SLATE project (slateci.io) and deployed in OKD, the community distribution of Red Hat OpenShift. Additionally, we discuss our method of certificate management and maintenance utilizing Kubernetes CronJobs. Finally, we also report on the two different configurations currently running at Fermilab, comparing and contrasting a Docker-based OKD deployment against a traditional RPM-based deployment.

97 MATHEMATICS AND COMPUTING↗

FTS3: Data Movement Service in containers deployed in OKD [Slides]

The File Transfer Service (FTS3) is a data movement service developed at CERN which is used to distribute the majority of the Large Hadron Collider's data across the Worldwide LHC Computing Grid (WLCG) infrastructure. At Fermilab, we have deployed a couple of FTS3 instances for Intensity Frontier experiments (e.g. DUNE) to transfer data in America and Europe, using a container-based strategy. During this talk, we are going to present the two different configurations currently running at Fermilab, comparing and contrasting a Docker-based OKD deployment against a traditional RPM-based deployment and giving an overview of the possible issues encountered. In addition, we discuss our method of certificate management and maintenance utilizing Kubernetes cronjobs.

97 MATHEMATICS AND COMPUTING↗

A case study on parallel HDF5 dataset concatenation for high energy physics data analysis

In High Energy Physics (HEP), experimentalists generate large volumes of data that, when analyzed, helps us better understand the fundamental particles and their interactions. This data is often captured in many files of small size, creating a data management challenge for scientists. In order to better facilitate data management, transfer, and analysis on large scale platforms, it is advantageous to aggregate data further into a smaller number of larger files. However, this translation process can consume significant time and resources, and if performed incorrectly the resulting aggregated files can be inefficient for highly parallel access during analysis on large scale platforms. In this paper, we present our case study on parallel I/O strategies and HDF5 features for reducing data aggregation time, making effective use of compression, and ensuring efficient access to the resulting data during analysis at scale. We focus on NOvA detector data in this case study, a large-scale HEP experiment generating many terabytes of data. Here, the lessons learned from our case study inform the handling of similar datasets, thus expanding community knowledge related to this common data management task.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Data for The utility of transfer learning to improve the performance of deep learning in axon segmentation

The utility of transfer learning to improve the performance of deep learning in axon segmentation Data Data: All the input and labeled volumes tf-logs: Tensorflow logs, view with command "tensorboard --logdir [name of folder]" Model Weights: model_weights: the argument list under variable combo indicate 1) no oversampling, 2) no rotation, 3) no learn scheduler, and 4) flipping on all three dimensions, and the additional values indicate 5) elastic deformation percentage, 6) rotate deformation percentage, 7) layer setting , 8) learning rate, and 9) training/validation/test data division suffix (leave '' if not using suffix). Results: Output from inference segment_total_results_validation_final: All validation results and calculations segment_total_results: All test results and calculations Authors The modified code was created for a paper by: Marjolein Oostrom, Michael A. Muniak, Rogene Eichler West, Sarah Akers, Paritosh Pande, Moses Obiri, Wei Wang, Kasey Bowyer, Zhuhao Wu, Lisa Bramer, Tianyi Mao, Bobbie Jo Webb-Robertson The work is adapted from Github TrailMap, which was created by Albert Pun and Drew Friedmann Acknowledgments MO, RMEW, SA, MO, LB, BJWR were supported by the Laboratory Directed Research and Development at Pacific Northwest National Laboratory (PNNL), a Department of Energy facility operated by Battelle under contract DE-AC05-76RLO01830. WW, KB, and ZW were supported in part by a NIH/BRAIN Initiative Grant RF1MH128969. MAM and TM were supported by two NIH/BRAIN Initiative Grants R01NS104944, RF1MH120119 and NIH R01NS081071. This research is affiliated with the Pacific northwest bioMedical Innovation Co-laboratory (PMedIC) collaboration between OHSU and PNNL.

Oostrom, Marjolein T↗

Extended Heat Transfer Model Dataset

Data repository for the paper: Molerus and Wirth's Heat Transfer Model for Bubbling Fluidized Beds: Proposal for an Extended Model Including Immersed Tube Banks and Particle Cross-Flow Powder Technology 2025

contact time↗

System Modeling Frameworks for Wind Turbines and Plants: Review and Requirements Specifications

System modeling frameworks for wind turbines and plants are used by research groups and industry to design wind energy systems that take into account key trade-offs across performance, cost, and reliability at both the turbine and plant level. The frameworks are exercised using a variety of multi-disciplinary design, analysis and optimization (MDAO) methods. To improve inter-operability and foster collaboration, this report proposes a classification system for the frameworks along dimensions of model fidelity and scope. The classification system is first motivated with reviews the state-of-the-art in the development of software frameworks for integrated wind turbine and plant simulation. Within each major wind turbine and power plant subsystem, a matrix is developed for the disciplines used and the fidelity levels with which each discipline can be modeled. The existing frameworks are then classified according to the matrix. Next, an ontology is proposed that will allow for standardizing how data is transferred between the most common discipline-fidelity combinations used in the frameworks. A common representation of data creates the ability to 1) share system descriptions and analysis results, supporting more transparent benchmarks and comparison, and 2) integrate models together into workflows within and across organizations for improving the efficiency and performance of wind turbine and power plant design processes. Ultimately, this integration leads to better overall wind energy system designs with high performance and low costs.

17 WIND ENERGY↗

Accelerating shared file checkpoint with local burst buffers

A data management system and method for accelerating shared file checkpointing. Written application data is aggregated in an application data file created in a local burst buffer memory at a compute node, and an associated data mapping built index to maintain information related to the offsets into a shared file at which segments of the application data is to be stored in a parallel file system, and where in the buffer those segments are located. The node asynchronously transfers a data file containing the application data and the associated data mapping index to a file server for shared file storage. The data management system and method further accelerates shared file checkpointing in which a shared file, together with a map file that specifies how the shared file is to be distributed, is asynchronously transferred to local burst buffer memories at the nodes to accelerate reading of the shared file.

Gooding, Thomas↗

Remote Radiation Sensing Using Aerial and Ground Platforms

Remote sensing of ionizing radiation has a significant role in waste management, nuclear material management and nonproliferation, and radiation safety. Robotic platforms can surpass the number of tasks that are achieved by humans. With this technique, the operator's radiation exposure can be decreased. Remote sensing allows for the evaluation and monitoring of radiological contamination. Gamma-ray and neutron sensors were integrated onto the robotic platforms. This approach allows for the radiation sensor data to be dynamically tracked and mapped thus enabling further analysis of the radiation flux in temporal and spatial domains. The goal is to complete scheduled tasks while the robot is being irradiated. To achieve this, electronic components must be shielded and radiation hardened. CZT Detector: Cadmium Zinc Telluride (CZT) detector technology has been a promising solution for gamma-ray and x-ray measurements. Detector data is transferred to the Odroid minicomputer that controls and powers the module via the USB. Robot Operating System (ROS) was utilized for data acquisition and data fusion. The Mariscotti method was employed for the spectrum analysis. A function was programmed in ROS for the automatic identification of photopeaks. CLYC Detector: A Cs{sub 2}LiYCl{sub 6}:Ce{sup 3+} (CLYC) detector was used for simultaneous medium-resolution gamma-ray measurements and neutron counting. A 2.54 cm diameter photomultiplier tube (PMT) was equipped with a high voltage supply and a miniature digitizer. Gamma-ray excitation: fast core-to-valence luminescence (CVL) with 1 ns decay constant, and prompt Ce{sup 3+} emission with 50 ns decay constant. Neutron excitation: slow cerium self-trapped excitation (Ce{sup 3+} STE), 1000 ns decay constant. Radiation Source Localization: Maximum Likelihood Estimation (MLE) and gradient-based methods were used to locate the position of a radiation source based on measured radiation intensities. Multi-Particle Transport Code FLUKA: Estimation of radiation damage of the electronic components is important in order to optimize the robot's operational time while it is irradiated. Displacement per atom (DPA) represents the radiation damage in materials exposed to the ionizing radiation. Various shielding layers of different thickness t were analyzed (< 5% statistical error). The model of the controller of the UAS was designed in FLUKA. Conclusion: CZT and CLYC detectors were integrated onto the robotic platforms. Radiation source localization and contour mapping using robotic platforms were studied. Functions for data analysis and fusion were developed in ROS. FLUKA code was utilized to analyze DPA values. Layers of low-density and high-density materials were used to shield the UAS electronics.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Recent Experience with the CMS Data Management System

The CMS[1] experiment manages a large-scale data infrastructure, currently handling over 200 PB of disk and 500 PB of tape storage and transferring more than 1 PB of data per day on average between various WLCG[2] sites. Utilizing Rucio[3] for high-level data management, FTS[4] for data transfers, and a variety of storage and network technologies at the sites, CMS confronts inevitable challenges due to the system’s growing scale and evolving nature. Key challenges include managing transfer and storage failures, optimizing data distribution across different storages based on production and analysis needs, implementing necessary technology upgrades and migrations, and efficiently handling user requests. The data management team has established comprehensive monitoring to supervise this system and has successfully addressed many of these challenges. The team’s efforts aim to ensure data availability and protection, minimize failures and manual interventions, maximize transfer throughput and resource utilization, and provide reliable user support. This paper details the operational experience of CMS with its data management system in recent years, focusing on the encountered challenges, the effective strategies employed to overcome them and the ongoing challenges as we prepare for future demands.

Öztürk, Hasan [CERN]↗

Integrated Methane Monitoring Platform Extension, Volume I: Final Technical Report

The IMMPE project, DE-FE0032284, was to enhance methane monitoring technologies and their applications across various natural gas asset classes. The scope included deploying advanced methane detection and monitoring technologies to identify and mitigate fugitive methane emissions, measuring emission rates, and assessing impacts. The findings included the successful mitigation of identified emissions and quantification of emission rates. A key outcome was the development of a comprehensive template and summary of recommendations for methane emissions monitoring, which is replicable for both upstream and downstream applications. Furthermore, the project emphasized the importance of education by providing training opportunities for technicians and regulators, thereby fostering awareness and promoting the adoption of cost-effective methane emissions monitoring and management techniques.

02 PETROLEUM↗

Trust-Enhancing Probabilistic Transfer Learning for Sparse and Noisy Data Environments

There is an increasing aspiration to utilize machine learning (ML) for various tasks of relevance to national security. ML models have thus far been mostly applied to tasks and domains that, while impactful, have sufficient volume of data. For predictive tasks of national security relevance, ML models of great capacity (ability to approximate nonlinear trends in input-output maps) are often needed to capture the complex underlying physics. However, scientific problems of relevance to national security are often accompanied by various sources of sparse and/or incomplete data, including experiments and simulations, across different regimes of operation, of varying degrees of fidelity, and include noise with different characteristics and/or intensity. State-of-the-art ML models, despite exhibiting superior performance on the task and domain they were trained on, may suffer detrimental loss in performance in such sparse data environments. This report summarizes the results of the Laboratory Directed Research and Development project entitled Trust-Enhancing Probabilistic Transfer Learning for Sparse and Noisy Data Environments. The objective of the project was to develop a new transfer learning (TL) framework that aims to adaptively blend the data across different sources in tackling one task of interest, resulting in enhanced trustworthiness of ML models for mission- and safety-critical systems. The proposed framework determines when it is worth applying TL and how much knowledge is to be transferred, despite uncontrollable uncertainties. The framework accomplishes this by leveraging concepts and techniques from the fields of Bayesian inverse modeling and uncertainty quantification, relying on strong mathematical foundations of probability and measure theories to devise new uncertainty-aware TL workflows.

97 MATHEMATICS AND COMPUTING↗

A ModEx Framework for Watershed Subsurface Investigation With Limited Geophysical Data Using Machine Learning and Hydrologic Modeling

Abstract Subsurface heterogeneity influences watershed hydrology strongly but remains difficult to characterize at catchment scales with sparse and costly field data. Geophysical surveys such as electromagnetic induction (EMI) provide local spatial subsurface images yet scaling them to watershed scales and converting EMI‐derived resistivity into hydraulic properties remains a challenge. We present a Model–Experiment (ModEx) framework that integrates limited EMI data with machine learning (ML) and hydrologic modeling to improve process representation and guide field investigations. Sparse EMI surveys were scaled to the catchment scale using a Random Forest model, and the resulting resistivity fields were combined with nearby borehole constraints to parameterize a hydrologic model. The EMI‐informed hydrological simulations improved predictions of streamflow sustained by subsurface flow and shallow saturation patterns. By combining EMI data and ML with hydrologic modeling, the ModEx framework guides future subsurface surveys, providing a transferable and efficient strategy for data–model integration across diverse watersheds. Plain Language Summary Mapping the underground network of soil and rock that controls water is essential for predicting floods and droughts, but seeing underground is difficult and expensive. We cannot drill everywhere, so scientists use geophysical tools to scan broad areas. There are two key challenges: these geophysical scans are often sparse across the whole watershed, and the geophysical data is hard to translate into water‐related properties. We used artificial intelligence to solve these problems. We taught a computer to find patterns linking the limited geophysical data to the land surface properties. This allowed it to fill in the gaps and create a complete, useful subsurface map for the entire watershed. This new map improves hydrologic simulations, leading to more accurate predictions of water movement in the watershed. It also helps scientists build better models with less data and generates a priority map showing where to measure next, making future investigations more efficient. Key Points Limited EMI scaled with ML improves catchment‐scale subsurface parameterization for hydrologic models The framework integrates hydrologic modeling with limited geophysical data to support subsurface investigation design ModEx framework offers a transferable data–model integration strategy that quantifies and reduces uncertainty guiding watershed studies

Chen, Hang↗

A Variable Eddington Factor Model for Thermal Radiative Transfer with Closure Based on Data-Driven Shape Function

Here, a new variable Eddington factor (VEF) model is presented for nonlinear problems of thermal radiative transfer (TRT). The VEF model is data-driven and acts on known (a-priori) radiation-diffusion solutions for material temperatures in the TRT problem. A linear auxiliary problem is constructed for the radiative transfer equation (RTE) whose emission source and opacities are evaluated at these known material temperatures. The solution to this RTE approximates the specific intensity distribution in phase-space and time. It is applied as a shape function to define the Eddington tensor for the presented VEF model. The shape function computed via the auxiliary RTE problem will capture some degree of transport effects within the TRT problem. The VEF moment equations closed with this approximate Eddington tensor will thus carry with them these captured transport effects. In this study, the temperature data comes from multigroup P 1 , P 1/3 , and flux-limited diffusion radiative transfer models. The proposed VEF model can be interpreted as a transport-corrected diffusion reduced-order model. Numerical results are presented on the Fleck-Cummings test problem which models a supersonic wavefront of radiation. The VEF model is shown to improve accuracy by 1–2 orders of magnitude compared to the considered radiation-diffusion model solutions to the TRT problem.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Preliminary Transfer Learning Results on Israel Data

In this preliminary report, we use publicly available data recorded in Israel to test and expand upon existing machine learning models for seismic-phase detection and arrival-time measurement. We downloaded 3-years of waveform data from Geofon, and cross referenced the waveforms to Israel bulletin picks (Schardong et al., 2021). The initial results using existing models directly generated ubiquitous false detections and that obscured detections of signals that are clearly visible in the waveforms. However, after applying transfer learning (tuning parameters in the existing ML models using one year of the Israel-network data), the results are encouraging, i.e. ML picks agree within a few tenths of a second with bulletin picks and the number of false detections is greatly reduced. The bulletin picks are a good starting point, but they cannot be considered ground-truth. To test potential improvement in picking using ML we would like to relocate the events using the ML picks to see if the events cluster more tightly at known mine locations. However, in order to constrain event locations, we need ML picks for the whole Israeli-Jordanian network, which requires waveforms that are not publicly available.

58 GEOSCIENCES↗

A Transfer Learning Strategy for Improving the Data Efficiency of Deep Reinforcement Learning Control in Smart Buildings

Reinforcement learning (RL) is a powerful tool that has shown promising results in many domains such as robotics and game-playing. Because RL algorithms learn optimal control policies by continuously interacting with their environments, these algorithms require a lot of data to learn, which limits their application to a wide range of domains. For this reason, there is an immense need for improving the training and data efficiency of RL. Towards addressing this research gap, this paper proposes a transfer learning (TL) approach to improve the efficiency of the RL algorithms by reducing data need and, thus, reducing training time. To demonstrate the proposed approach, a knowledge transfer from a set of buildings to another building was conducted. The results show that the proposed TL approach is a promising method that can efficiently harness the information from similar RL tasks and reduce the data needs of RL algorithms.

Amasyali, Kadir↗

Deep structural clustering for single-cell RNA-seq data jointly through autoencoder and graph neural network

Abstract Single-cell RNA sequencing (scRNA-seq) permits researchers to study the complex mechanisms of cell heterogeneity and diversity. Unsupervised clustering is of central importance for the analysis of the scRNA-seq data, as it can be used to identify putative cell types. However, due to noise impacts, high dimensionality and pervasive dropout events, clustering analysis of scRNA-seq data remains a computational challenge. Here, we propose a new deep structural clustering method for scRNA-seq data, named scDSC, which integrate the structural information into deep clustering of single cells. The proposed scDSC consists of a Zero-Inflated Negative Binomial (ZINB) model-based autoencoder, a graph neural network (GNN) module and a mutual-supervised module. To learn the data representation from the sparse and zero-inflated scRNA-seq data, we add a ZINB model to the basic autoencoder. The GNN module is introduced to capture the structural information among cells. By joining the ZINB-based autoencoder with the GNN module, the model transfers the data representation learned by autoencoder to the corresponding GNN layer. Furthermore, we adopt a mutual supervised strategy to unify these two different deep neural architectures and to guide the clustering task. Extensive experimental results on six real scRNA-seq datasets demonstrate that scDSC outperforms state-of-the-art methods in terms of clustering accuracy and scalability. Our method scDSC is implemented in Python using the Pytorch machine-learning library, and it is freely available at https://github.com/DHUDBlab/scDSC.

Gan, Yanglan↗