Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data Sharing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

RBDMS, FracFocus, State Support, and Produced Water Initiatives

Award DE-FE-0027702 from the Department of Energy to the Ground Water Protection Council (GWPC) focused on state and federal priorities in the areas of state Risk Based Data Management System (RBDMS) development, connectivity between state systems and FracFocus.org, and data sharing initiatives across agencies. The primary objective was to enhance the RBDMS by adding new components relevant to current environmental topics such as hydraulic fracturing, increasing field inspection capabilities, creating linkages between FracFocus and state programs, upgrading eForm capabilities, and analyzing potential for data sharing. The recipient worked with state agencies developing RBDMS module(s) that meet these needs.

54 ENVIRONMENTAL SCIENCES↗

An automated integrated web-based smart tool for open stope design

The Stability Graph is a widely used tool for the design of open stopes in underground mining. Many users of the Stability Graph still apply this design method manually. Although the manual approach has benefits, using multiple graphs and stability number computation charts for each stope surface is time-consuming, even for the experienced mining engineer. Current practice in the use of the method also limits data sharing. This paper presents a StopeSoft web-based tool for open stope stability prediction that is developed on the basis of the Stability Graph method and is available at openstope.com. StopeSoft incorporates flexibility in terms of Stability Graph options and incorporates additional critical factors often overlooked. As a web-based tool, StopeSoft encourages and makes data sharing possible globally, focused on expanding the database and improving the current limitations of the Stability Graph to provide practical, reliable solutions for mining engineers, consultants, and academics. The StopeSoft automated process facilitates the process of open stope stability prediction, saving time and minimizing potential human errors. Statistical treatment of the data accounts for the variability of input parameters to emphasize the probabilistic nature of the Stability Graph method. The probabilistic interpretation of the stability states of stope surfaces eliminates the false feeling of absolute stope performance based on its location on the Stability Graph , as implied by the deterministic approach.

58 GEOSCIENCES↗

Baylor University Campus-Wide Deep Dive

In January 2020, staff members from the Engagement and Performance Operations Center (EPOC) and the Lonestar Education And Research Network (LEARN) met with researchers and staff at Baylor University for the purpose of a Campus-Wide Deep Dive into research drivers. The goal of this meeting was to help characterize the requirements for five campus research use cases and to enable cyberinfrastructure support staff to better understand the needs of the researchers they support. Profiled scientific use cases included: - Experimental High Energy Physics (HEP) - Proton Computed Tomography (pCT) - Nutrition and Relation to Digestive Microbiome - Baylor University Core Research Facilities - Molecular Quantum-dot Cellular Automata (QCA), and Material Science of Quantum Computing - Modeling and Simulation of Low-Dimensional and Nano-Structured Materials - Computational Fluid Dynamics Material for this event included the written documentation from each of the research areas at Baylor University, documentation about the current state of technology support, and a write-up of the discussion that took place in person. The Case Studies highlighted the ongoing challenges that Baylor University has in supporting a cross-section of established and emerging research use cases. Each Case Study mentioned unique challenges which were summarized into common needs. These included: - Tradeoffs for network/software security, and usability of the resulting infrastructure. Better communication to set expectations and understand realities is required. - Computation use on campus is widespread and healthy. While no major problems were uncovered, upgrades to maintain current usage patterns and encourage growth will be required. - Storage is a critical need for enterprise use cases and research. In particular, a campus wide ‘storage architecture’ to support research use cases (e.g. instruments, data sharing) is required in the 2-5 year time window. - Instrumentation on campus is healthy and expanding. Technology must scale with this in the form of computation and storage. - Working with LEARN to upgrade network capacity (in multiples of 10G, or upgrades to 100G) will be required in the 1-3 year time frame. - Network monitoring and visibility will help to establish external science use cases. - Data sharing via portal systems is not currently a critical need, but growing in scope. EPOC can assist Baylor with options.

99 GENERAL AND MISCELLANEOUS↗

The future low-temperature geochemical data-scape as envisioned by the U.S. geochemical community

Data sharing benefits the researcher, the scientific community, and the public by allowing the impact of data to be generalized beyond one project and by making science more transparent. However, many scientific communities have not developed protocols or standards for publishing, citing, and versioning datasets. One community that lags in data management is that of low-temperature geochemistry (LTG). This paper resulted from an initiative from 2018 through 2020 to convene LTG and data scientists in the U.S. to strategize future management of LTG data. Through webinars, a workshop, a preprint, a townhall, and a community survey, the group of U.S. scientists discussed the landscape of data management for LTG – the data-scape. Currently this data-scape includes a “street bazaar” of data repositories. This was deemed appropriate in the same way that LTG scientists publish articles in many journals. The variety of data repositories and journals reflect that LTG scientists target many different scientific questions, produce data with extremely different structures and volumes, and utilize copious and complex metadata. Nonetheless, the group agreed that publication of LTG science must be accompanied by sharing of data in publicly accessible repositories, and, for sample-based data, registration of samples with globally unique persistent identifiers. LTG scientists should use certified data repositories that are either highly structured databases designed for specialized types of data, or unstructured generalized data systems. Recognizing the need for tools to enable search and cross-referencing across the proliferating data repositories, the group proposed that the overall data informatics paradigm in LTG should shift from “build data repository, data will come” to “publish data online, cybertools will find”. Funding agencies could also provide portals for LTG scientists to register funded projects and datasets, and forge approaches that cross national boundaries. Finally, the needed transformation of the LTG data culture requires emphasis in student education on science and management of data.

58 GEOSCIENCES↗

A novel approach for adaptive skeleton toolpath generation

Industry 4.0 is revolutionizing manufacturing through the integration of automation and real-time data sharing in cyber-physical systems. At the forefront of this revolution is large-format additive manufacturing. In large-format printing, parts are often designed to be an even number of beads wide to produce a completely dense part. However, voids can still arise. This is often due to the part not being an even number of bead widths wide in some areas, or in geometry containing acute angles, as the process of generating closed contours cannot completely fill the space. Voids can be tolerated in smaller models, but in large-format additive manufacturing they may cause mechanical defects. To fill these voids, open loop paths called skeletons are often used, but they are typically limited by the physical constraints defined in the slicing software. To address this, researchers at Oak Ridge National Laboratory have extended skeleton toolpaths via an adaptive methodology. These adaptive skeletons were found to better fill void spaces through manipulation of physical parameters of the build process and were calculated as part of the slicing process.

42 ENGINEERING↗

NGEE Arctic Authorship Guidelines

Authorship Guidelines were developed to help facilitate trust among team members as we span multiple institutions, scientific disciplines, and career stages. NGEE Arctic was built on a foundation of open science, data sharing, and collaboration. In Phase 4 of the project, it was particularly important to keep this foundation in mind as we develop new collaborations across the Arctic. Included in this package is one *.pdf. The Next-Generation Ecosystem Experiments in the Arctic (NGEE Arctic) project is a research effort to reduce uncertainty in the Department of Energy’s Energy Exascale Earth System Model (E3SM) by developing a predictive understanding of Arctic tundra ecosystems underlain by permafrost and to quantify feedbacks from the Arctic tundra to the Earth system. NGEE Arctic is supported by the Department of Energy's Office of Biological and Environmental Research. Over Phases 1–3, observations made by the NGEE Arctic team across a gradient of permafrost landscapes in Arctic Alaska improved the representation of tundra processes in the land surface component of E3SM (the E3SM Land Model, ELM). Model improvements emphasized unique aspects of permafrost environments and explored reductions in model complexity while retaining predictive power. The Arctic-informed ELM developed by NGEE Arctic has been used to make novel predictions on processes ranging from permafrost thaw to soil biogeochemical cycling to Earth system feedbacks associated with the unique characteristics of tundra plants. In Phase 4, the NGEE Arctic team is evaluating our new predictive understanding under novel conditions across the Arctic domain. In collaboration with partners at long-term pan-Arctic research sites we are examining whether an Arctic-informed ELM can faithfully simulate interactions among surface and subsurface processes at site, regional, and pan-Arctic scales. In turn, we are using variety of tools to dynamically extend and evaluate ELM inference, with an emphasis on data synthesis and pan-Arctic model evaluation, reintegration of code with an evolving E3SM, scaling across heterogeneous Arctic landscapes, and the appropriate representation of the impacts of increasingly frequent Arctic disturbances.

Iversen, Colleen [ORNL] (ORCID:0000000182933450)↗

A Data Deposition Platform for Sharing Nuclear Magnetic Resonance Data

Nuclear magnetic resonance (NMR) data are rarely deposited in open databases, leading to loss of critical scientific knowledge. Existing data reporting methods (images, tables, lists of values) contain less information than raw data, and are poorly standardized. Together, these issues limit FAIR (findable, accessible, interoperable, reusable) access to these data, which in turn creates barriers for compound dereplication and the development of new data-driven discovery tools. Existing NMR databases are either not designed for natural products data, or employ complex deposition interfaces that disincentivize deposition. Journals, including the Journal of Natural Products (JNP), are now requiring data submission as part of the publication process, creating the need for a streamlined, user-friendly mechanism to deposit and distribute NMR data. Recently, our team reported the development of the Natural Products Magnetic Resonance Database (NP-MRD; www.np-mrd.org). Here in this paper we present a new data deposition platform for the NP-MRD project that is designed to enable users to deposit NMR data for published or submitted manuscripts in under five minutes. This platform includes a suite of automated data extraction and standardization tools, together with a simple-to-use web-based interface and detailed error reporting to simplify the data deposition process and is available at www.np-mrd.org/submissions.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Challenges of open data in aquatic sciences: issues faced by data users and data providers

Free use and redistribution of data (i.e., Open Data) increases the reproducibility, transparency, and pace of aquatic sciences research. However, barriers to both data users and data providers may limit the adoption of Open Data practices. Here, we describe common Open Data challenges faced by data users and data providers within the aquatic sciences community (i.e., oceanography, limnology, hydrology, and others). These challenges were synthesized from literature, authors’ experiences, and a broad survey of 174 data users and data providers across academia, government agencies, industry, and other sectors. Through this work, we identified seven main challenges: 1) metadata shortcomings, 2) variable data quality and reusability, 3) open data inaccessibility, 4) lack of standardization, 5) authorship and acknowledgement issues 6) lack of funding, and 7) unequal barriers around the globe. Our key recommendation is to improve resources to advance Open Data practices. This includes dedicated funds for capacity building, hiring and maintaining of skilled personnel, and robust digital infrastructures for preparation, storage, and long-term maintenance of Open Data. Further, to incentivize data sharing we reinforce the need for standardized best practices to handle data acknowledgement and citations for both data users and data providers. We also highlight and discuss regional disparities in resources and research practices within a global perspective.

54 ENVIRONMENTAL SCIENCES↗

DYnamic and Asynchronous Data Streamliner

DYAD aims to help sharing data files between producer and consumer job elements, especially between co-scheduled jobs or within an ensemble. DYAD provides the service by two components: a FLUX module and a I/O wraper set. DYAD transparently synchronizes file I/O between producer and consumer, and transfers data from the producer location to the consumer location managed by the service. Users only need to use the file path that is under the directory managed by the service.

Ahn, DongH↗

First Plant Cell Atlas symposium report

The Plant Cell Atlas (PCA) community hosted a virtual symposium on December 9 and 10, 2021 on single cell and spatial omics technologies. The conference gathered almost 500 academic, industry, and government leaders to identify the needs and directions of the PCA community and to explore how establishing a data synthesis center would address these needs and accelerate progress. This report details the presentations and discussions focused on the possibility of a data synthesis center for a PCA and the expected impacts of such a center on advancing science and technology globally. Community discussions focused on topics such as data analysis tools and annotation standards; computational expertise and cyber-infrastructure; modes of community organization and engagement; methods for ensuring a broad reach in the PCA community; recruitment, training, and nurturing of new talent; and the overall impact of the PCA initiative. These targeted discussions facilitated dialogue among the participants to gauge whether PCA might be a vehicle for formulating a data synthesis center. The conversations also explored how online tools can be leveraged to help broaden the reach of the PCA (i.e., online contests, virtual networking, and social media stakeholder engagement) and decrease costs of conducting research (e.g., virtual REU opportunities). Major recommendations for the future of the PCA included establishing standards, creating dashboards for easy and intuitive access to data, and engaging with a broad community of stakeholders. The discussions also identified the following as being essential to the PCA's success: identifying homologous cell-type markers and their biocuration, publishing datasets and computational pipelines, utilizing online tools for communication (such as Slack), and user-friendly data visualization and data sharing. In conclusion, the development of a data synthesis center will help the PCA community achieve these goals by providing a centralized repository for existing and new data, a platform for sharing tools, and new analytical approaches through collaborative, multidisciplinary efforts. A data synthesis center will help the PCA reach milestones, such as community-supported data evaluation metrics, accelerating plant research necessary for human and environmental health.

59 BASIC BIOLOGICAL SCIENCES↗

Federated Learning and Differential Privacy: What might AI-Enhanced co-design of microelectronics learn?

Data is a valuable commodity, and it is often dispersed over multiple entities. Sharing data or models created from the data is not simple due to concerns regarding security, privacy, ownership, and model inversion. This limitation in sharing can hinder model training and development. Federated learning can enable data or model sharing across multiple entities that control local data without having to share or exchange the data themselves. Differential privacy is a conceptual framework that brings strong mathematical guarantee for privacy protection and helps provide a quantifiable privacy guarantee to any data or models shared. The concepts of federated learning and differential privacy are introduced along with possible connections. Lastly, some open discussion topics on how federated learning and differential privacy can tied to AI-Enhanced co-design of microelectronics are highlighted.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

HydroGEN and H2NEW Data Hub

HydroGEN is a multi-lab consortium focused on early-stage R&D in H2 production, supported by the U.S. Department of Energy (DOE), Office of Energy Efficiency and Renewable Energy (EERE), Hydrogen and Fuel Cell Technologies Office (HFTO). The consortium advances research and development (R&D) of innovative materials for advanced water splitting (AWS) technologies to enable clean, sustainable and low-cost ($1/kg H2) hydrogen production and fosters cross-cutting innovation using theory-guided applied materials R&D to advance all emerging water-splitting pathways for hydrogen production. Hydrogen from Next-generation Electrolyzers of Water (H2NEW) is a consortium of nine U.S. Department of Energy (DOE) national laboratories focused on making large-scale electrolyzers, which produce hydrogen from electricity and water, more durable, efficient, and affordable. The presentation introduces the HydroGEN data hub to the H2NEW consortium and encourages H2NEW members to use it to store and share data with team members and eventually the public. The two consortia share this data hub.

advanced water splitting technologies↗

Discussion on “Saving Storage in Climate Ensembles: A Model-Based Stochastic Approach”

Abstract We thank the authors for this interesting paper that highlights important ideas and concepts for the future of climate model ensembles and their storage, as well as future uses of stochastic emulators. Stochastic emulators are particularly relevant because of the statistical nature of climate model ensembles, as discussed in previous work of the authors (Castruccio et al. in J Clim 32:8511–8522, 2019; Hu and Castruccio in J Clim 34:8409–8418, 2021). We thank the authors for sharing of some of their data with us in order to illustrate this discussion. In the following, in Sect. 1 we discuss alternative techniques currently used and studied, namely lossy compression and ideas emerging from the climate modeling community, that could feed the discussion on ensemble and storage. In that section, we also present numerical results of compression performed on the data shared by the authors. In Sect. 2, we discuss the current statistical model proposed by the authors and its context. We discuss other potential uses of stochastic emulators in climate and Earth modeling.

97 MATHEMATICS AND COMPUTING↗

International Radiation Monitoring Information System (IRMIS) – a versatile EPR tool for all member states [Slides]

IRMIS provides Competent Authorities, IEC, and other relevant International Organizations with a data sharing tool that helps to report and share information, evaluate radiation monitoring data to assess if the public is safe, identify protective actions, keep the public informed by Member State Competent Authority, and maintain transparency of data handling and processing.

61 RADIATION PROTECTION AND DOSIMETRY↗

An Integrated Platform for Collaborative Data Analytics

While collaboration among data scientists is a key to organizational productivity, data analysts face significant barriers to achieving this end, including data sharing, accessing and configuring the required computational environment, and a unified method of sharing knowledge. Each of these barriers to collaboration is related to the fundamental question of knowledge management “how can organizations use knowledge more effectively?”. In this paper, we consider the problem of knowledge management in collaborative data analytics and present ShareAL, an integrated knowledge management platform, as a solution to that problem. The ShareAL platform consists of three core components: a full stack web application, a dashboard for analyzing streaming data and a High Performance Computing (HPC) cluster for performing real time analysis. Prior research has not applied knowledge management to collaborative analytics or developed a platform with the same capabilities as ShareAL. ShareAL overcomes the barriers data scientists face to collaboration by providing intuitive sharing of data and analytics via the web application, a shared computing environment via the HPC cluster and knowledge sharing and collaboration via a real time messaging application.

Oesch, T↗

Validation of standardized data formats and tools for ground-level particle-based gamma-ray observatories

Context. Ground-based γ-ray astronomy is still a rather young field of research, with strong historical connections to particle physics. This is why most observations are conducted by experiments with proprietary data and analysis software, as is usual in the particle physics field. However, in recent years, this paradigm has been slowly shifting toward the development and use of open-source data formats and tools, driven by upcoming observatories such as the Cherenkov Telescope Array (CTA). In this context, a community-driven, shared data format (the gamma-astro-data-format, or GADF) and analysis tools such as Gammapy and ctools have been developed. So far, these efforts have been led by the Imaging Atmospheric Cherenkov Telescope community, leaving out other types of ground-based γ-ray instruments. Aims. We aim to show that the data from ground particle arrays, such as the High-Altitude Water Cherenkov (HAWC) observatory, are also compatible with the GADF and can thus be fully analyzed using the related tools, in this case, Gammapy. Methods. We reproduced several published HAWC results using Gammapy and data products compliant with GADF standard. We also illustrate the capabilities of the shared format and tools by producing a joint fit of the Crab spectrum including data from six different γ-ray experiments. Results. We find excellent agreement with the reference results, a powerful confirmation of both the published results and the tools involved. Conclusions. The data from particle detector arrays such as the HAWC observatory can be adapted to the GADF and thus analyzed with Gammapy. A common data format and shared analysis tools allow multi-instrument joint analysis and effective data sharing. To emphasize this, a sample of Crab nebula event lists is made public with this paper. Because of the complementary nature of pointing and wide-field instruments, this synergy will be distinctly beneficial for the joint scientific exploitation of future observatories such as the Southern Wide-field Gamma-ray Observatory and CTA.

79 ASTRONOMY AND ASTROPHYSICS↗

Methods for safely sharing dual-use genetic data

Background: Some genetic data has dual-use potential. Sharing pathogen data has shown tremendous value. For example therapeutic development and lineage tracking during the COVID pandemic. This data sharing is complicated by the fact that these data have the potential to be used for harm. The genome sequence of a pathogen can be used to enable malicious genetic engineering approaches or to recreate the pathogen from synthetic DNA. Standard data security methods can be applied to genetic data, but when data is shared between institutions, ensuring appropriate security can be difficult. Sensitive data that is shared internationally among a wide array of institutions can be especially difficult to control. Methods for securely storing and sharing genetic data with potential for dual-use are needed to mitigate this potential harm.Results: Here we propose new methods that allow genetic data to be shared in a data format that prevents a nefarious actor from accessing sensitive aspects of the data. Our methods obfuscate raw sequence data by pooling reads from different samples. This approach can ensure that data is secure while stored and during electronic transfer. We demonstrate that by pooling raw sequence data from multiple samples of the same organism, the ability to fully reconstruct any individual sample is prevented. In the pooled data, most genomic information remains, but reads or mutations cannot be directly attributed to any individual sample. To further restrict access to information, regions of a genome can be removed from the reads.Conclusion: Our methods obscure genomic information within raw sequence reads. This method can allow genetic data to be stored and shared while preventing a nefarious actor from being able to perfectly reconstruct an organism. Broad-scale sequence information remains, while fine scale details about specific samples are difficult or impossible to reconstruct. Our software is available at https://github.com/Geneinfosec-Inc/ReadMixer.

59 BASIC BIOLOGICAL SCIENCES↗

Livewire User Guide

The Livewire Data Platform houses a catalog of transportation- and mobility-related project data, as well as a publications database, making it easy to search and share data. It allows transportation researchers, industry, and academic partners to increase the visibility of their projects within the research community, securely share and preserve data, and leverage datasets from other projects. Public data on Livewire are open to anyone with a Livewire account. This guide will help Livewire users understand how to store project data as a data steward, as well as access data as a data consumer.

33 ADVANCED PROPULSION SYSTEMS↗