Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data transfer”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Particle track classification using quantum associative memory

Pattern recognition algorithms are commonly employed to simplify the challenging and necessary step of track reconstruction in sub-atomic physics experiments. Aiding in the discrimination of relevant interactions, pattern recognition seeks to accelerate track reconstruction by isolating signals of interest. In high collision rate experiments, such algorithms can be particularly crucial for determining whether to retain or discard information from a given interaction even before the data is transferred to tape. As data rates, detector resolution, noise, and inefficiencies increase, pattern recognition becomes more computationally challenging, motivating the development of higher efficiency algorithms and techniques. Quantum associative memory is an approach that seeks to exploits quantum mechanical phenomena to gain advantage in learning capacity, or the number of patterns that can be stored and accurately recalled. Here, we study quantum associative memory based on quantum annealing and apply it to the particle track classification. We focus on discrimination models based on Ising formulations of quantum associative memory model (QAMM) recall and quantum content-addressable memory (QCAM) recall. We characterize classification performance of these approaches as a function detector resolution, pattern library size, and detector inefficiencies, using the D-Wave 2000Q processor as a testbed. Discrimination criteria is set using both solution-state energy and classification labels embedded in solution states. We find that energy-based QAMM classification performs well in regimes of small pattern density and low detector inefficiency. In contrast, state-based QCAM achieves reasonably high accuracy recall for large pattern density and the greatest recall accuracy robustness to a variety of detector noise sources.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Biological carbon pump estimate based on multidecadal hydrographic data

The transfer of photosynthetically produced organic carbon from surface to mesopelagic waters draws carbon dioxide from the atmosphere. However, current observation-based estimates disagree on the strength of this biological carbon pump (BCP). Earth system models (ESMs) also exhibit a large spread of BCP estimates, indicating limited representations of the known carbon export pathways. Here we use several decades of hydrographic observations to produce a top-down estimate of the strength of the BCP with an inverse biogeochemical model that implicitly accounts for all known export pathways. Our estimate of total organic carbon (TOC) export at 73.4 m (model euphotic zone depth) is 15.00 ± 1.12 Pg C year –1 , with only two-thirds reaching 100 m depth owing to rapid remineralization of organic matter in the upper water column. Partitioned by sequestration time below the euphotic zone, τ, the globally integrated organic carbon production rate with τ > 3 months is 11.09 ± 1.02 Pg C year –1 , dropping to 8.25 ± 0.30 Pg C year –1 for τ > 1 year, with 81% contributed by the non-advective-diffusive vertical flux owing to sinking particles and vertically migrating zooplankton. Nevertheless, export of organic carbon by mixing and other fluid transport of dissolved matter and suspended particles remains regionally important for meeting the respiratory carbon demand. Here, the temperature dependence of the sequestration efficiency inferred from our inversion suggests that future global warming may intensify the recycling of organic matter in the upper ocean, potentially weakening the BCP.

58 GEOSCIENCES↗

Progressive transfer learning for low-frequency data prediction in full waveform inversion

To effectively overcome the cycle-skipping issue in full waveform inversion (FWI), we developed a deep neural network (DNN) approach to predict the absent low-frequency components by exploiting the hidden physical relation connecting the low- and the high-frequency data. To efficiently solve this challenging nonlinear regression problem, two novel strategies were proposed to design the DNN architecture and to optimize the learning process: (1) dual data feed structure; (2) progressive transfer learning. With the dual data feed structure, not only the high-frequency data, but also the corresponding beat tone data are fed into the DNN to relieve the burden of feature extraction. The second strategy, progressive transfer learning, enables us to train the DNN using a single evolving training dataset. Within the framework of the progressive transfer learning, the training dataset continuously evolves in an iterative manner by gradually retrieving the subsurface information through the physics-based inversion module, progressively enhancing the prediction accuracy of the DNN and propelling the inversion process out of the local minima. Here, the synthetic numerical experiments suggest that, without any a priori geological information, the low-frequency data predicted by the progressive transfer learning are sufficiently accurate for an FWI engine to produce reliable subsurface velocity models free of cycle-skipping artifacts.

02 PETROLEUM↗

Pushing the Dark-Matter Limit: The LUX-Zeplin (LZ) Experiment (Final Report)

This report highlights the accomplishments of the dark-matter group of the University of Rochester during the project period. During the grant period, all of our responsibilities for LUX were completed. Several papers describing the results of LUX analyses carried out by the Rochester group were published. The success of the Rochester trigger electronics, developed for LUX, resulted in it being selected to provide the digital platform for LZ. The PI is leading an excellent team of engineers, scientists, and students to provide the signal processing electronics, the control and monitoring system for the LZ detector, and the online computing infrastructure to process the LZ data and transfer them to the US Data Center at NERSC. During the grant period most of the R&D and prototyping of the LZ electronics was completed.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Transfer Factors for the FRMAC Assessment Manual and Turbo FRMAC to Improve Radiological Dose Assessment

The Turbo FRMAC analysis tool is used to perform complex calculations to quickly evaluate radiological consequences and aid in decision making during an emergency response by assessing impacts to the public, workers, and the food supply. Turbo FRMAC calculations are based on methods established by the Federal Radiological Monitoring and Assessment Center (FRMAC). To be able to assess impacts, input data called transfer factors that describe radionuclide uptake by local plants and animals are required. During the code application exercises in late 2016, identifying, finding, and validating nonstandard transfer factors proved to be time consuming and diverted the teams’ activity away from other critical tasks. As a result, a task was undertaken to dramatically expand the list of available transfer factors (food and non-food) and incorporate these factors into Turbo FRMAC. This will ultimately result in improved efficiency of the assessment team to perform calculations during times when the FRMAC is activated and provide more defensible, vetted data from which to calculate results. As a result of these and subsequent exercises, transfer factors were needed for the following items: bell peppers, Christmas tree, deer, flowers, fresh cucumbers, tomatoes, grapefruit, lichens, mushrooms, oranges, snap beans, squash, strawberries, sugarcane for sugar and seed, sweet corn, tea, tobacco, tree bark, and watermelon. To expand the applicability of the information tables, generic transfer data were also provided for common categories and recommendations were made for expanding the list of chemicals based on chemical similarity. The data presented in this report were compiled from recent literature with most of the data encompassing the period from 2000 to 2018. The following radionuclides were targeted during the literature search: elements associated with reactor accidents or nuclear detonations (Sr 89/90, Cs 134/137, Ce 141/144, Ru 103/106, I 129/131/133, Pu 238/239, Am 241, Zr 95, and Nb 95) and elements associated with industrial accidents or dirty bombs (Ir 192 and Co 60). Data from other elements were evaluated if they were identified during the literature search. In addition, reports were evaluated that were recommended by the research consultants. For each plant or animal transfer factor, the goal was to determine the geometric mean, the geometric standard deviation, the minimum, the maximum, and the number of measurements used. If only one measurement was available that was presented as the mean. Concentration ratios with large geometric standard deviations (GSDs) were generally the result of a paucity of measurements or a few disparate measurements. Measurement disparity was observed for data from different soil types. The geometric mean (GM) is a good reference value for planning and responses purposes, but the location-specific concentrations are unlikely to be similar to the model results. Recommendations for further work include: segregating the data to reflect the influence of soil type, developing approaches to incorporate animal data based on aggregated transfer measurements, and including data for foliar deposition on plants.

61 RADIATION PROTECTION AND DOSIMETRY↗

Classical and quantum compression for edge computing: the ubiquitous data dimensionality reduction

Edge computing aims to address the challenges associated with communicating and transferring large amounts of data generated remotely to a data center in a timely and efficient manner. A central pillar of edge computing is local (i.e., at- or near-source) data processing capability so that data transfer to a data center for processing can be minimized. Data compression at the edge is therefore a natural component of edge workflows. Here we present a survey of data compression algorithms with a focus on edge computing. Not all compression algorithms can accommodate the data type heterogeneity, tight processing and communication time constraints, or energy efficiency requirement characteristics of edge computing. We discuss specific examples of compression algorithms that are being explored in the context of edge computing. We end our review with a brief survey of emerging quantum compression techniques that are of importance in quantum information processing, including the proposed concept of quantum edge computing.

97 MATHEMATICS AND COMPUTING↗

CatCost™ Data Tools [SWR-20-59]

CatCost™ enables catalyst researchers to develop comprehensive estimates of the cost to produce pre-commercial catalysts early in the catalyst development process. It includes both web browser and Excel-based tools, each with distinct strengths. The web tool includes powerful visualization capabilities, while the Excel tool allows for high degrees of control and customization by power users. The CatCost™ Python Data Tools are designed to facilitate the transfer of data from the Excel tool to the web tool, so that users can take advantage of both the powerful visualization in the web tool and the higher degree of customization available in the excel tool. Rather than painstakingly copying entries from estimates, the materials, equipment, and spent catalyst libraries from Excel to a JSON format for the web tool, the CatCost™ Data Tools provide an intuitive applet which allows them to simply choose which library or estimate they wish to transfer, and then returns files that integrate smoothly with the web version of the tool.

Baddour, Frederick↗

Dynamic cache bypassing

A processing system fills a memory access request for data from a processor core by bypassing a cache when a write congestion condition is detected, and when transferring the data to the cache would cause eviction of a dirty cache line. The cache is bypassed by transferring the requested data to the processor core or to a different cache. Accordingly, the processing system can temporarily bypass the cache storing the dirty cache line when filling a memory access request, thereby avoiding the eviction and write back to main memory of a dirty cache line when a write congestion condition exists.

97 MATHEMATICS AND COMPUTING↗

Predicting weather impacts on corn production in a data-limited region using a transfer learning approach

The stability of food supply and prices may depend more on annual changes in yields from year-to-year variability in weather than on longer-term average changes from changing climatic conditions. However, the absence of high-quality data on crop yields at fine spatial resolutions in many regions of the world makes it challenging to statistically model their response to interannual variability in weather patterns. Therefore, there is a need for empirical methods that can project annual crop yield changes even in limited data regions. Here, we propose a transfer learning algorithm that uses high spatial resolution data from one region to project yields in another region with more limited data. The goal of our work is to understand what data types can be beneficial for transferring learning from a source region to a very different target region with more limited data. We utilize Long Short-Term Memory to develop a transfer learning model that is trained on historical county-level corn yield in the United States and predicts district-level corn yield variations in India. Even using smaller amounts of data in India, simulating a data-scarce region, we achieve an average root mean square error of 0.48 bu acre−1 in predicting interannual yield variations. Using Shapley values to interpret results, we explore the contribution of the different weather parameters to interannual yield variability and find a larger influence of precipitation-related variables. Our study demonstrates the usefulness of this method for transferring models of weather impacts on crop yields trained on a data-rich country to one with more limited data. It suggests the potential of applying the transfer learning model to mitigate the need for extensive raw data globally.

Vishwakarma, Srishti [ORNL] (ORCID:000000031674419↗

Enabling modern data discovery for atmospheric measurements

The Atmospheric Radiation Measurement (ARM) user facility is a US Department of Energy Office of Science user facility that is managed and operated through a collaborative effort led by nine US Department of Energy national laboratories. The ARM Data Center, located at Oak Ridge National Laboratory, is responsible for the timely collection, processing, and delivery of data products to the scientific community. The ARM Data Center holds more than 11,000 data products, including metadata collected from field campaigns, instruments, value-added products, and principal investigator–contributed data. These data sets are checked for successful transfer (for most data, this transfer is carried out automatically via the network; however, some of the largest data sets and some of the most remote sites require manual shipping of hard disks) and both the data and metadata are processed to a standard format, which is an ARM-standardized structure, via the Network Common Data Form. The Network Common Data Form is a self-describing binary format with many compatible software tools. Once processed, the data are cataloged, stored in the ARM Data Archive, and made discoverable through association with an array of metadata-characterizing information, such as location and measurement classification. These metadata enable powerful search capabilities through the ARM Data Center Data Discovery interface. This paper discusses the workflow of how the new discovery system has been redesigned from user requirements and how the data are distributed to the scientific community.

54 ENVIRONMENTAL SCIENCES↗

Classification of animal sounds in a hyperdiverse rainforest using convolutional neural networks with data augmentation

To protect tropical forest biodiversity, we need to be able to detect it reliably, cheaply, and at scale. Automated detection of sound producing animals from passively recorded soundscapes via machine-learning approaches is a promising technique towards this goal, but it is constrained by the necessity of large training data sets. Using soundscapes from a tropical forest in Borneo and a Convolutional Neural Network model (CNN), we investigate i) the minimum viable training data set size for accurate prediction of call types (‘sonotypes’), and ii) the extent to which data augmentation and transfer learning can overcome the issue of small and imbalanced training data sets. We found that even relatively high sample sizes (>80 per sonotype) lead to mediocre accuracy, which however improved significantly with data augmentation and transfer learning, including at extremely small sample sizes (3 per sonotype), regardless of taxonomic group or call characteristics. Neither transfer learning nor data augmentation alone achieved high accuracy. Our results suggest that transfer learning and data augmentation could make the use of CNNs to classify species’ vocalizations feasible even for small soundscape-based projects with many rare species. Retraining our open-source model requires only basic programming skills which makes it possible for individual conservation initiatives to match their local context, in order to enable more evidence-informed management of biodiversity.

54 ENVIRONMENTAL SCIENCES↗

Verification of RESRAD-OFFSITE Code (V.4)

This report documents the verification of RESRAD-OFFSITE Version 4.0 and describes, where necessary, the verification of the following: • The data comprising the standard dose and risk coefficient libraries in the RESRAD database files Master_dcf_ICRP07.mdb and Master_dcf_2k.mdb. • The extraction and transfer of the data from the selected database file to the computational code by the RESRAD-OFFSITE 4.0 interface, ResOWin.exe. • The different processes that are modeled by the main computational code in RESRAD OFFSITE 4.0, ResOMain.exe. • The data displayed in the graphical and text reports. Many verifications were performed as part of the quality assurance quality control program associated with the development and release of RESRAD-OFFSITE 4.0, namely: • developer testing, • internal independent testing, and • release testing. Some were also performed in response to questions from users regarding the performance of the code. The main text of the report focuses on summarizing a subset of those tests, both independent and developer tests that verified the computations performed by the code. The verifications included in this report served as the basis for the development of the release tests of the computational executables and provided the quantitative results to be compared with the code output. The input and output interfaces and the data transfers between the various executables of the code were tested while performing the verification testing. They were tested intentionally during release testing. This report also provides some basic information to help in understanding the activities that were verified. The report: • outlines the components of RESRAD-OFFSITE 4.0 and the interconnections between these components, • outlines the processes modeled by the computational code, • provides summary figures and tables to offer confirmation of the verification of the computational components of the code, • reproduces the verifiers’ reports, if available, in individual appendices, • refers to the previous verification report (Yu et al. 2011) for more details about some of the verifications, and • reproduces the test cases and the testers’ reports from the release testing in individual appendices, when possible.

54 ENVIRONMENTAL SCIENCES↗

A General Framework for Progressive Data Compression and Retrieval

In scientific simulations, observations, and experiments, the transfer of data to and from disk and across networks has become a major bottleneck for data analysis and visualization. Compression techniques have been employed to tackle this challenge, but traditional lossy methods often demand conservative error tolerances to meet the numerical accuracy requirements of both anticipated and unknown data analysis tasks. Progressive data compression and retrieval has emerged as a promising solution, where each analysis task dictates its own accuracy needs. However, few analysis algorithms inherently support progressive data processing, and adapting compression techniques, file formats, client/server frameworks, and APIs to support progressivity can be challenging. Here, this paper presents a framework that enables progressive-precision data queries for any data compressor or numerical representation. Our strategy hinges on a multi-component representation that successively reduces the error between the original and compressed field, allowing each field in the progressive sequence to be expressed as a partial sum of components. We have implemented this approach with four established scientific data compressors and assessed its effectiveness using real-world data sets from the SDRBench collection. The results show that our framework competes in accuracy with the standalone compressors it is based upon. Additionally, (de)compression time is proportional to the number of components requested by the user. Finally, our framework allows for fully lossless compression using lossy compressors when a sufficient number of components are employed.

97 MATHEMATICS AND COMPUTING↗

Replication Data for: Oxophilicity Drives Oxygen Transfer at a Palladium−Silver Interface for Increased CO Oxidation Activity

The data underlying this published work have been made publicly available in this repository as part of the IMASC Data Management Plan. This work was supported as part of the Integrated Mesoscale Architectures for Sustainable Catalysis (IMASC), an Energy Frontier Research Center funded by the U.S. Department of Energy, Office of Science, Basic Energy Sciences under Award # DE-SC0012573.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

FTS3: Data Movement Service in containers deployed in OKD

The File Transfer Service (FTS3) is a data movement service developed at CERN which is used to distribute the majority of the Large Hadron Collider's data across the Worldwide LHC Computing Grid (WLCG) infrastructure. At Fermilab, we have deployed FTS3 instances for Intensity Frontier experiments (e.g. DUNE) to transfer data in America and Europe, using a container-based strategy. In this article we summarize our experience building docker images based on work from the SLATE project (slateci.io) and deployed in OKD, the community distribution of Red Hat OpenShift. Additionally, we discuss our method of certificate management and maintenance utilizing Kubernetes CronJobs. Finally, we also report on the two different configurations currently running at Fermilab, comparing and contrasting a Docker-based OKD deployment against a traditional RPM-based deployment.

Lobato Pardavila, Lorena↗