Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data Sharing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Collaborative Data Publication Utilizing the Open Data Repository's Data Publisher

For small communities in multidisciplinary fields such as astrobiology, publishing and sharing data can be challenging. While large, homogenous fields often have repositories and existing data standards, small groups of independent researchers have few options for publishing data that can be utilized within their community. In conjunction with teams at NASA Ames and the University of Arizona, a number of pilot studies are being conducted to assess the needs of these research groups and to guide the software development so that it allows them to publish and share their data collaboratively.

Human-readable interfaces↗

DYnamic and Asynchronous Data Streamliner

DYAD aims to help sharing data files between producer and consumer job elements, especially between co-scheduled jobs or within an ensemble. DYAD provides the service by two components: a FLUX module and a I/O wraper set. DYAD transparently synchronizes file I/O between producer and consumer, and transfers data from the producer location to the consumer location managed by the service. Users only need to use the file path that is under the directory managed by the service.

Ahn, DongH↗

First Plant Cell Atlas symposium report

The Plant Cell Atlas (PCA) community hosted a virtual symposium on December 9 and 10, 2021 on single cell and spatial omics technologies. The conference gathered almost 500 academic, industry, and government leaders to identify the needs and directions of the PCA community and to explore how establishing a data synthesis center would address these needs and accelerate progress. This report details the presentations and discussions focused on the possibility of a data synthesis center for a PCA and the expected impacts of such a center on advancing science and technology globally. Community discussions focused on topics such as data analysis tools and annotation standards; computational expertise and cyber-infrastructure; modes of community organization and engagement; methods for ensuring a broad reach in the PCA community; recruitment, training, and nurturing of new talent; and the overall impact of the PCA initiative. These targeted discussions facilitated dialogue among the participants to gauge whether PCA might be a vehicle for formulating a data synthesis center. The conversations also explored how online tools can be leveraged to help broaden the reach of the PCA (i.e., online contests, virtual networking, and social media stakeholder engagement) and decrease costs of conducting research (e.g., virtual REU opportunities). Major recommendations for the future of the PCA included establishing standards, creating dashboards for easy and intuitive access to data, and engaging with a broad community of stakeholders. The discussions also identified the following as being essential to the PCA's success: identifying homologous cell-type markers and their biocuration, publishing datasets and computational pipelines, utilizing online tools for communication (such as Slack), and user-friendly data visualization and data sharing. In conclusion, the development of a data synthesis center will help the PCA community achieve these goals by providing a centralized repository for existing and new data, a platform for sharing tools, and new analytical approaches through collaborative, multidisciplinary efforts. A data synthesis center will help the PCA reach milestones, such as community-supported data evaluation metrics, accelerating plant research necessary for human and environmental health.

59 BASIC BIOLOGICAL SCIENCES↗

Federated Learning and Differential Privacy: What might AI-Enhanced co-design of microelectronics learn?

Data is a valuable commodity, and it is often dispersed over multiple entities. Sharing data or models created from the data is not simple due to concerns regarding security, privacy, ownership, and model inversion. This limitation in sharing can hinder model training and development. Federated learning can enable data or model sharing across multiple entities that control local data without having to share or exchange the data themselves. Differential privacy is a conceptual framework that brings strong mathematical guarantee for privacy protection and helps provide a quantifiable privacy guarantee to any data or models shared. The concepts of federated learning and differential privacy are introduced along with possible connections. Lastly, some open discussion topics on how federated learning and differential privacy can tied to AI-Enhanced co-design of microelectronics are highlighted.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

HydroGEN and H2NEW Data Hub

HydroGEN is a multi-lab consortium focused on early-stage R&D in H2 production, supported by the U.S. Department of Energy (DOE), Office of Energy Efficiency and Renewable Energy (EERE), Hydrogen and Fuel Cell Technologies Office (HFTO). The consortium advances research and development (R&D) of innovative materials for advanced water splitting (AWS) technologies to enable clean, sustainable and low-cost ($1/kg H2) hydrogen production and fosters cross-cutting innovation using theory-guided applied materials R&D to advance all emerging water-splitting pathways for hydrogen production. Hydrogen from Next-generation Electrolyzers of Water (H2NEW) is a consortium of nine U.S. Department of Energy (DOE) national laboratories focused on making large-scale electrolyzers, which produce hydrogen from electricity and water, more durable, efficient, and affordable. The presentation introduces the HydroGEN data hub to the H2NEW consortium and encourages H2NEW members to use it to store and share data with team members and eventually the public. The two consortia share this data hub.

advanced water splitting technologies↗

Priority Queuing On A Parallel Data Bus

Queuing strategy for communications along shared data bus minimizes number of data lines while always assuring user of highest priority given access to bus. New system handles up to 32 user demands on 17 data lines that previously serviced only 17 demands.

Wallis, D. E.↗

CubeSat Array for Detection of RF Emissions from Exoplanets using Inter-Satellite Optical Communicators

We are presenting a CubeSat mission concept wherein a large number of small spacecraft furnished with Ultra-Long Wavelength (ULW) observation antennas is to be implemented to synthesize a large aperture for detecting ULW emissions from selected exoplanets. The CubeSat Array for the Detection of RF Emissions from Exoplanets (CADRE) concept involves the placement of a swarm arrangement of CubeSats into small amplitude Lissajous orbits around the Earth-Moon L2 point. In addition to the ULW receivers, each CADRE CubeSat will be furnished with a novel omnidirectional optical communicator which should allow data sharing at gigabit per second data rates. In this paper we discuss the science goals, potential architecture and key technical challenges of an eventual CADRE implementation. We present the estimated interferometer specifications needed to detect Exoplanet RF emissions, CubeSat orbits and formations as well as CubeSat design considerations. Special focus is placed on the design of the mission-enabling inter-CubeSat optical cross link which allows for fast sharing of the detected RF signals for subsequent correlation.

Bentum, Mark↗

Discussion on “Saving Storage in Climate Ensembles: A Model-Based Stochastic Approach”

Abstract We thank the authors for this interesting paper that highlights important ideas and concepts for the future of climate model ensembles and their storage, as well as future uses of stochastic emulators. Stochastic emulators are particularly relevant because of the statistical nature of climate model ensembles, as discussed in previous work of the authors (Castruccio et al. in J Clim 32:8511–8522, 2019; Hu and Castruccio in J Clim 34:8409–8418, 2021). We thank the authors for sharing of some of their data with us in order to illustrate this discussion. In the following, in Sect. 1 we discuss alternative techniques currently used and studied, namely lossy compression and ideas emerging from the climate modeling community, that could feed the discussion on ensemble and storage. In that section, we also present numerical results of compression performed on the data shared by the authors. In Sect. 2, we discuss the current statistical model proposed by the authors and its context. We discuss other potential uses of stochastic emulators in climate and Earth modeling.

97 MATHEMATICS AND COMPUTING↗

International Radiation Monitoring Information System (IRMIS) – a versatile EPR tool for all member states [Slides]

IRMIS provides Competent Authorities, IEC, and other relevant International Organizations with a data sharing tool that helps to report and share information, evaluate radiation monitoring data to assess if the public is safe, identify protective actions, keep the public informed by Member State Competent Authority, and maintain transparency of data handling and processing.

61 RADIATION PROTECTION AND DOSIMETRY↗

An overview of the Opus language and runtime system

We have recently introduced a new language, called Opus, which provides a set of Fortran language extensions that allow for integrated support of task and data parallelism. lt also provides shared data abstractions (SDA's) as a method for communication and synchronization among these tasks. In this paper, we first provide a brief description of the language features and then focus on both the language-dependent and language-independent parts of the runtime system that support the language. The language-independent portion of the runtime system supports lightweight threads across multiple address spaces, and is built upon existing lightweight thread and communication systems. The language-dependent portion of the runtime system supports conditional invocation of SDA methods and distributed SDA argument handling.

Mehrotra, Piyush↗

Runtime support for data parallel tasks

We have recently introduced a set of Fortran language extensions that allow for integrated support of task and data parallelism, and provide for shared data abstractions (SDA's) as a method for communications and synchronization among these tasks. In this paper we discuss the design and implementation issues of the runtime system necessary to support these extensions, and discuss the underlying requirements for such a system. To test the feasibility of this approach, we implement a prototype of the runtime system and use this to support an abstract multidisciplinary optimization (MDO) problem for aircraft design. We give initial results and discuss future plans.

Haines, Matthew↗

A Virtual Bioinformatics Knowledge Environment for Early Cancer Detection

Discovery of disease biomarkers for cancer is a leading focus of early detection. The National Cancer Institute created a network of collaborating institutions focused on the discovery and validation of cancer biomarkers called the Early Detection Research Network (EDRN). Informatics plays a key role in enabling a virtual knowledge environment that provides scientists real time access to distributed data sets located at research institutions across the nation. The distributed and heterogeneous nature of the collaboration makes data sharing across institutions very difficult. EDRN has developed a comprehensive informatics effort focused on developing a national infrastructure enabling seamless access, sharing and discovery of science data resources across all EDRN sites. This paper will discuss the EDRN knowledge system architecture, its objectives and its accomplishments.

knowledge systems↗

An Integrated Platform for Collaborative Data Analytics

While collaboration among data scientists is a key to organizational productivity, data analysts face significant barriers to achieving this end, including data sharing, accessing and configuring the required computational environment, and a unified method of sharing knowledge. Each of these barriers to collaboration is related to the fundamental question of knowledge management “how can organizations use knowledge more effectively?”. In this paper, we consider the problem of knowledge management in collaborative data analytics and present ShareAL, an integrated knowledge management platform, as a solution to that problem. The ShareAL platform consists of three core components: a full stack web application, a dashboard for analyzing streaming data and a High Performance Computing (HPC) cluster for performing real time analysis. Prior research has not applied knowledge management to collaborative analytics or developed a platform with the same capabilities as ShareAL. ShareAL overcomes the barriers data scientists face to collaboration by providing intuitive sharing of data and analytics via the web application, a shared computing environment via the HPC cluster and knowledge sharing and collaboration via a real time messaging application.

Oesch, T↗

We Need A Better Way to Share Earth Observations

A more accessible, open data-sharing infrastructure will engage a broader community of contributors, helping to develop satellite data products that benefit Earth science research and applications.

Apps & Software↗

Validation of standardized data formats and tools for ground-level particle-based gamma-ray observatories

Context. Ground-based γ-ray astronomy is still a rather young field of research, with strong historical connections to particle physics. This is why most observations are conducted by experiments with proprietary data and analysis software, as is usual in the particle physics field. However, in recent years, this paradigm has been slowly shifting toward the development and use of open-source data formats and tools, driven by upcoming observatories such as the Cherenkov Telescope Array (CTA). In this context, a community-driven, shared data format (the gamma-astro-data-format, or GADF) and analysis tools such as Gammapy and ctools have been developed. So far, these efforts have been led by the Imaging Atmospheric Cherenkov Telescope community, leaving out other types of ground-based γ-ray instruments. Aims. We aim to show that the data from ground particle arrays, such as the High-Altitude Water Cherenkov (HAWC) observatory, are also compatible with the GADF and can thus be fully analyzed using the related tools, in this case, Gammapy. Methods. We reproduced several published HAWC results using Gammapy and data products compliant with GADF standard. We also illustrate the capabilities of the shared format and tools by producing a joint fit of the Crab spectrum including data from six different γ-ray experiments. Results. We find excellent agreement with the reference results, a powerful confirmation of both the published results and the tools involved. Conclusions. The data from particle detector arrays such as the HAWC observatory can be adapted to the GADF and thus analyzed with Gammapy. A common data format and shared analysis tools allow multi-instrument joint analysis and effective data sharing. To emphasize this, a sample of Crab nebula event lists is made public with this paper. Because of the complementary nature of pointing and wide-field instruments, this synergy will be distinctly beneficial for the joint scientific exploitation of future observatories such as the Southern Wide-field Gamma-ray Observatory and CTA.

79 ASTRONOMY AND ASTROPHYSICS↗

Methods for safely sharing dual-use genetic data

Background: Some genetic data has dual-use potential. Sharing pathogen data has shown tremendous value. For example therapeutic development and lineage tracking during the COVID pandemic. This data sharing is complicated by the fact that these data have the potential to be used for harm. The genome sequence of a pathogen can be used to enable malicious genetic engineering approaches or to recreate the pathogen from synthetic DNA. Standard data security methods can be applied to genetic data, but when data is shared between institutions, ensuring appropriate security can be difficult. Sensitive data that is shared internationally among a wide array of institutions can be especially difficult to control. Methods for securely storing and sharing genetic data with potential for dual-use are needed to mitigate this potential harm.Results: Here we propose new methods that allow genetic data to be shared in a data format that prevents a nefarious actor from accessing sensitive aspects of the data. Our methods obfuscate raw sequence data by pooling reads from different samples. This approach can ensure that data is secure while stored and during electronic transfer. We demonstrate that by pooling raw sequence data from multiple samples of the same organism, the ability to fully reconstruct any individual sample is prevented. In the pooled data, most genomic information remains, but reads or mutations cannot be directly attributed to any individual sample. To further restrict access to information, regions of a genome can be removed from the reads.Conclusion: Our methods obscure genomic information within raw sequence reads. This method can allow genetic data to be stored and shared while preventing a nefarious actor from being able to perfectly reconstruct an organism. Broad-scale sequence information remains, while fine scale details about specific samples are difficult or impossible to reconstruct. Our software is available at https://github.com/Geneinfosec-Inc/ReadMixer.

59 BASIC BIOLOGICAL SCIENCES↗

Livewire User Guide

The Livewire Data Platform houses a catalog of transportation- and mobility-related project data, as well as a publications database, making it easy to search and share data. It allows transportation researchers, industry, and academic partners to increase the visibility of their projects within the research community, securely share and preserve data, and leverage datasets from other projects. Public data on Livewire are open to anyone with a Livewire account. This guide will help Livewire users understand how to store project data as a data steward, as well as access data as a data consumer.

33 ADVANCED PROPULSION SYSTEMS↗

Clinical knowledge extraction via sparse embedding regression (KESER) with multi-center large scale electronic health record data

The increasing availability of electronic health record (EHR) systems has created enormous potential for translational research. However, it is difficult to know all the relevant codes related to a phenotype due to the large number of codes available. Traditional data mining approaches often require the use of patient-level data, which hinders the ability to share data across institutions. In this project, we demonstrate that multi-center large-scale code embeddings can be used to efficiently identify relevant features related to a disease of interest. We constructed large-scale code embeddings for a wide range of codified concepts from EHRs from two large medical centers. We developed knowledge extraction via sparse embedding regression (KESER) for feature selection and integrative network analysis. We evaluated the quality of the code embeddings and assessed the performance of KESER in feature selection for eight diseases. Besides, we developed an integrated clinical knowledge map combining embedding data from both institutions. The features selected by KESER were comprehensive compared to lists of codified data generated by domain experts. Features identified via KESER resulted in comparable performance to those built upon features selected manually or with patient-level data. The knowledge map created using an integrative analysis identified disease-disease and disease-drug pairs more accurately compared to those identified using single institution data. Analysis of code embeddings via KESER can effectively reveal clinical knowledge and infer relatedness among codified concepts. KESER bypasses the need for patient-level data in individual analyses providing a significant advance in enabling multi-center studies using EHR data.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗