Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data Sharing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Livewire User Guide

The Livewire Data Platform houses a catalog of transportation- and mobility-related project data, as well as a publications database, making it easy to search and share data. It allows transportation researchers, industry, and academic partners to increase the visibility of their projects within the research community, securely share and preserve data, and leverage datasets from other projects. Public data on Livewire are open to anyone with a Livewire account. This guide will help Livewire users understand how to store project data as a data steward, as well as access data as a data consumer.

33 ADVANCED PROPULSION SYSTEMS↗

Clinical knowledge extraction via sparse embedding regression (KESER) with multi-center large scale electronic health record data

The increasing availability of electronic health record (EHR) systems has created enormous potential for translational research. However, it is difficult to know all the relevant codes related to a phenotype due to the large number of codes available. Traditional data mining approaches often require the use of patient-level data, which hinders the ability to share data across institutions. In this project, we demonstrate that multi-center large-scale code embeddings can be used to efficiently identify relevant features related to a disease of interest. We constructed large-scale code embeddings for a wide range of codified concepts from EHRs from two large medical centers. We developed knowledge extraction via sparse embedding regression (KESER) for feature selection and integrative network analysis. We evaluated the quality of the code embeddings and assessed the performance of KESER in feature selection for eight diseases. Besides, we developed an integrated clinical knowledge map combining embedding data from both institutions. The features selected by KESER were comprehensive compared to lists of codified data generated by domain experts. Features identified via KESER resulted in comparable performance to those built upon features selected manually or with patient-level data. The knowledge map created using an integrative analysis identified disease-disease and disease-drug pairs more accurately compared to those identified using single institution data. Analysis of code embeddings via KESER can effectively reveal clinical knowledge and infer relatedness among codified concepts. KESER bypasses the need for patient-level data in individual analyses providing a significant advance in enabling multi-center studies using EHR data.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

The Open Data Repositorys Data Publisher

Data management and data publication are becoming increasingly important components of researcher's workflows. The complexity of managing data, publishing data online, and archiving data has not decreased significantly even as computing access and power has greatly increased. The Open Data Repository's Data Publisher software strives to make data archiving, management, and publication a standard part of a researcher's workflow using simple, web-based tools and commodity server hardware. The publication engine allows for uploading, searching, and display of data with graphing capabilities and downloadable files. Access is controlled through a robust permissions system that can control publication at the field level and can be granted to the general public or protected so that only registered users at various permission levels receive access. Data Publisher also allows researchers to subscribe to meta-data standards through a plugin system, embargo data publication at their discretion, and collaborate with other researchers through various levels of data sharing. As the software matures, semantic data standards will be implemented to facilitate machine reading of data and each database will provide a REST application programming interface for programmatic access. Additionally, a citation system will allow snapshots of any data set to be archived and cited for publication while the data itself can remain living and continuously evolve beyond the snapshot date. The software runs on a traditional LAMP (Linux, Apache, MySQL, PHP) server and is available on GitHub (http://github.com/opendatarepository) under a GPLv2 open source license. The goal of the Open Data Repository is to lower the cost and training barrier to entry so that any researcher can easily publish their data and ensure it is archived for posterity.

Astrobiology data↗

International Data Collaboration for Risk-Based Safety and Mission Assurance

Some of the topics being discussed during this presentation are How things are being collaborating now, the Benefits from more data collaboration, the Opportunities for data collaboration and How can they share data. Additional information will be discussed through out the presentation slide deck.

Fischer, Gerd M.↗

Advanced Computing, Data Science, and Artificial Intelligence Research Opportunities for Energy-Focused Transportation Science

The Energy Efficient Mobility Systems (EEMS) technology landscape is complex and rapidly evolving, which provides both tremendous opportunities and formidable challenges. Significant alterations to the mobility landscape are underway due to the advent of vehicle and infrastructure connectivity, autonomous driving, and rapid passenger- and freight-vehicle electrification. Advanced computing will play an increasingly important role in enabling the EEMS program to understand and identify the most important levers to improve the energy productivity of future integrated mobility systems. It is also driving new approaches to mobility and the research to unlock an affordable, efficient, safe, and accessible transportation future. Driving much of this change is the collection, analysis, and strategic use of massive amounts of diverse, complex data from infrastructure and vehicles with on-board sensors and data storage and transmission capabilities. Diverse and representative data are key to implementing approaches to maximize mobility energy productivity. While high-fidelity modeling of integrated transportation networks has strengthened our understanding of dynamic movement and behavior patterns, existing tools must be expanded beyond their current focus. This work necessitates data infrastructure investments (e.g., secure-streaming data platforms driven by ubiquitous sensors and video analytics) as well as investments in critical capabilities for large-scale automated analysis and organization using modern machine learning, statistics, and artificial intelligence. Other chief needs include agile, large-scale storage that can be quickly searched and queried for relevant data to support validation and model development, data-sharing agreements, and formatting standards for key data types. The future of public transit must be explored in greater detail, research must inform design, and opportunities must be identified for improving the mobility productivity of public transit in both urban and rural America.

33 ADVANCED PROPULSION SYSTEMS↗

Tracking and data relay satellite system - NASA's new spacecraft data acquisition system

This paper describes NASA's new spacecraft acquisition system provided by the Tracking and Data Relay Satellite System (TDRSS). Four satellites in geostationary orbit and a ground terminal will provide complete tracking, telemetry, and command service for all of NASA's orbital satellites below a 12,000 km altitude. Western Union will lease the system, operate the ground terminal and provide operational satellite control. NASA's network control center will be the focal point for scheduling user services and controlling the interface between TDRSS and the NASA communications network, project control centers, and data processing. TDRSS single access user spacecraft data systems will be designed for time shared data relay support, and reimbursement policy and rate structure for non-NASA users are being developed.

Schneider, W. C.↗

The Deep-Time Digital Earth program: data-driven discovery in geosciences

Current barriers hindering data-driven discoveries in deep-time Earth (DE) include: substantial volumes of DE data are not digitized; many DE databases do not adhere to FAIR (findable, accessible, interoperable and reusable) principles; we lack a systematic knowledge graph for DE; existing DE databases are geographically heterogeneous; a significant fraction of DE data is not in open-access formats; tailored tools are needed. These challenges motivate the Deep-Time Digital Earth (DDE) program initiated by the International Union of Geological Sciences and developed in cooperation with national geological surveys, professional associations, academic institutions and scientists around the world. DDE’s mission is to build on previous research to develop a systematic DE knowledge graph, a FAIR data infrastructure that links existing databases and makes dark data visible, and tailored tools for DE data, which are universally accessible. DDE aims to harmonize DE data, share global geoscience knowledge and facilitate data-driven discovery in the understanding of Earth’s evolution.

Chengshan Wang↗

PAS: Privacy Algorithms in Systems

Today we face an explosion of data generation, ranging from health monitoring to national security infrastructure systems. More and more systems are connected to the Internet that collects data at regular time intervals. These systems share data and use machine learning methods for intelligent decisions, which resulted in numerous real-world applications (e.g., autonomous vehicles, recommendation systems, and heart-rate monitoring) that have benefited from it. However, these approaches are prone to identity thief and other privacy related cyber-security attacks. So, how can data privacy be protected efficiently in these scenarios? More dedicated efforts are needed to propose the integration of privacy techniques into existing systems and develop more advanced privacy techniques to address the complex challenges of multi-system connectivity and data fusion. Therefore, we have introduced Privacy Algorithms in Systems (PAS) at CIKM which provides a venue to gather academic researchers and industry researchers/practitioners to present their research in an effort to advance the frontier of this critical direction of privacy algorithms in systems.

Kotevska, Olivera↗

Metadata Standards for the NSE: Extended Field Standards

This standard presents a set of optional metadata fields for managed digital objects within the Nuclear Security Enterprise (NSE) and provides a deeper look at data representation in metadata by looking at the representation of 1) Records Management required metadata, and 2) common representations of technical/scientific data. Metadata standardization is a critical enabler for effectively sharing data, documents, and other digital objects between NSE sites, and for tracing the digital thread at the object level. Standardization is necessary for both schemas and vocabularies, meaning that both field standards and value standards must be specified. This document serves as a complementary field standard, recommending an optional set of fields that should be uniformly built for all managed digital objects within the NSE. This document specifically focuses on extending the shared discovery layer defined in the first white paper by introducing additional descriptive and data representation fields that improve cross-site search and interpretation.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Satellites as Shared Resources for Caribbean Climate and Health Studies

Remotely-sensed data and observations are providing powerful new tools for addressing climate and environment-related human health problems through increased capabilities for monitoring, risk mapping, and surveillance of parameters useful to such problems as vector-borne and infectious diseases, air and water quality, harmful algal blooms, UV (ultraviolet) radiation, contaminant and pathogen transport in air and water, and thermal stress. Remote sensing, geographic information systems (GIS), global positioning systems (GPS), improved computational capabilities, and interdisciplinary research between the Earth and health science communities are being combined in rich collaborative efforts resulting in more rapid problem-solving, early warning, and prevention in global health issues. Collaborative efforts among scientists from health and Earth sciences together with local decision-makers are enabling increased understanding of the relationships between changes in temperature, rainfall, wind, soil moisture, solar radiation, vegetation, and the patterns of extreme weather events and the occurrence and patterns of diseases (especially, infectious and vector-borne diseases) and other health problems. This increased understanding through improved information and data sharing, in turn, empowers local health and environmental officials to better predict health problems, take preventive measure, and improve response actions. This paper summarizes the remote sensing systems most useful for climate, environment and health studies of the Caribbean region and provides several examples of interdisciplinary research projects in the Caribbean currently using remote sensing technologies. These summaries include the use of remote sensing of algal blooms, pollution transport, coral reef monitoring, vectorborne disease studies, and potential health effects of African dust on Trinidad and Barbados.

Maynard, Nancy G.↗

Source Term Analysis of Xenon (STAX): An effort focused on differentiating man-made isotope production from nuclear explosions via stack monitoring

An overview of the hardware and software developed for the Source Term Analysis of Xenon (STAX) project is presented which includes the data collection from two stack monitoring systems installed at medical isotope production facilities, infrastructure to transfer data to a central repository, and methods for sharing data from the repository with users. STAX is an experiment to collect radioxenon emission data from industrial nuclear facilities with the goal of developing a better understanding of the global radioxenon background and the effect industrial radioxenon releases have on nuclear explosion monitoring. The final goal of this work is to utilize collected data along with atmospheric transport modeling to calculate the contribution of a peak or set of peaks detected by the International Monitoring System (IMS) to provide desired discriminating information to the International Data Centre (IDC) and National Data Centers (NDCs). Types of data received from the STAX equipment are shown and collected data was used for a case study to predict radioxenon concentrations at two IMS stations closest to the Institute for RadioElements (IRE) in Belgium. The initial evaluation of results indicate that the data is very valuable to the nuclear explosion monitoring community.

07 ISOTOPE AND RADIATION SOURCES↗

Leveraging History to Predict Infrequent Abnormal Transfers in Distributed Workflows

Scientific computing heavily relies on data shared by the community, especially in distributed data-intensive applications. This research focuses on predicting slow connections that create bottlenecks in distributed workflows. In this study, we analyze network traffic logs collected between January 2021 and August 2022 at the National Energy Research Scientific Computing Center (NERSC). Based on the observed patterns, we define a set of features primarily based on history for identifying low-performing data transfers. Typically, there are far fewer slow connections on well-maintained networks, which creates difficulty in learning to identify these abnormally slow connections from the normal ones. We devise several stratified sampling techniques to address the class-imbalance challenge and study how they affect the machine learning approaches. Our tests show that a relatively simple technique that undersamples the normal cases to balance the number of samples in two classes (normal and slow) is very effective for model training. This model predicts slow connections with an F1 score of 0.926.

97 MATHEMATICS AND COMPUTING↗

Using Block-local Atomicity to Detect Stale-value Concurrency Errors

Data races do not cover all kinds of concurrency errors. This paper presents a data-flow-based technique to find stale-value errors, which are not found by low-level and high-level data race algorithms. Stale values denote copies of shared data where the copy is no longer synchronized. The algorithm to detect such values works as a consistency check that does not require any assumptions or annotations of the program. It has been implemented as a static analysis in JNuke. The analysis is sound and requires only a single execution trace if implemented as a run-time checking algorithm. Being based on an analysis of Java bytecode, it encompasses the full program semantics, including arbitrarily complex expressions. Related techniques are more complex and more prone to over-reporting.

Artho, Cyrille↗

Influence Network: Network visualization of influence between stories for Earth Science data and information exploration

Using storytelling to present data has been demonstrated as an effective way to help data users gain deeper insight of information. For this purpose, the “Data in Action” story concept has been adopted by NASA Commercial Smallsat Data Acquisition (CSDA) Program to encourage creating and sharing data information in story form. To aid researchers in exploring similar stories and data in their fields of interest, we created an “Influence Network” within the “Data in Action” framework. The Influence Network is a data visualization system component which gathers information of story relationships using a concept called “influence,” which we define based on the number of visits and keyword similarity between stories. The visualization then provides insights into the influence flows between different stories within the system. The visualization of this “Influence Network” focuses on one story at a time, introducing a time series of neighbor stories that have the most influence to the currently focused story. By focusing on views over time and visualizing influence flows between stories, we aim to assist authors in understanding how their readers perceive their stories as well as advancing their methods for delivering more meaningful stories to expand and lower the barrier to use of CSDA datasets.

Dan Pham↗

Integration of evidence across human and model organism studies: A meeting report

The National Institute on Drug Abuse and Joint Institute for Biological Sciences at the Oak Ridge National Laboratory hosted a meeting attended by a diverse group of scientists with expertise in substance use disorders (SUDs), computational biology, and FAIR (Findability, Accessibility, Interoperability, and Reusability) data sharing. The meeting's objective was to discuss and evaluate better strategies to integrate genetic, epigenetic, and 'omics data across human and model organisms to achieve deeper mechanistic insight into SUDs. Specific topics were to (a) evaluate the current state of substance use genetics and genomics research and fundamental gaps, (b) identify opportunities and challenges of integration and sharing across species and data types, (c) identify current tools and resources for integration of genetic, epigenetic, and phenotypic data, (d) discuss steps and impediment related to data integration, and (e) outline future steps to support more effective collaboration—particularly between animal model research communities and human genetics and clinical research teams. This review summarizes key facets of this catalytic discussion with a focus on new opportunities and gaps in resources and knowledge on SUDs.

59 BASIC BIOLOGICAL SCIENCES↗

Enabling Model Organism and Commercial Astronaut Data Access Through the NASA Open Science Data Repository

NASA’s Open Science Data Repository (OSDR) brings together omics data from NASA’s GeneLab project and non-omics data, including physiological, phenotypic, imaging, and behavioral data from NASA’s Ames Life Sciences Data Archive (ALSDA) collected from decades of space biology research, providing open and FAIR (findable, accessible, interoperable, and reusable) access of these precious data to scientists world-wide. This rich source of meticulously curated metadata and data from spaceflight and analog studies has been mined by the scientific community resulting in dozens of high impact scientific publications that reveals a complex network of molecular and physiological effects of spaceflight across living systems, from microbes to plants, to mammals. Understanding how these effects translate to the human condition is critical as we move deeper into the era of commercial space travel. However, the integration of data, specifically omics data, from astronauts is particularly challenging due to their sensitive nature. OSDR has risen to this challenge by developing a mechanism to control access to identifiable levels of omics data, such as raw sequence data, while enabling public access to processed, unidentifiable, data and associated metadata that will allow the scientific community to interrogate human astronaut data alongside data from model organisms to begin answering these critical questions. The 2021 SpaceX Inspiration4 (I4) mission collected a comprehensive atlas of biological measurements from four civilian astronauts, providing a wealth of data to characterize the effects of spaceflight on the human body. These data include both non-omics and omics assays such as direct RNA sequencing (RNA-seq), single nuclei ATAC-seq and RNA-seq, metagenomics, proteomics, and comprehensive metabolic and cytokine panels, all of which have been integrated into the OSDR system across no less than 9 studies. Each study has been carefully curated using community-backed OSDR standards for sample and assay level metadata ensuring these data are findable and accessible. In addition to hosting both raw and processed data from the principal investigator team for each assay type, the GeneLab team plans to re-process the I4 omics data using GeneLab’s standard processing pipelines. The GeneLab processed data outputs will allow for comparisons across studies on OSDR and enable visualization of these data through the OSDR data visualization platform thereby enabling data reusability and interoperability. Here we describe the robust privacy and security protocols implemented by OSDR to safeguard sensitive health data from astronauts while facilitating metadata and processed data sharing for research purposes. We further provide a road map for navigating the vast amount of data provided for each I4 study on the OSDR, including experimental design, associated experiments, payloads, and missions, data generation and analysis protocols, and associated scientific articles. Additionally, we illustrate how to interrogate the standardized metadata provided in the sample and assay tables as well as various means to download and access the data including programmatically through the GeneLab Open API (GLOpenAPI). The open access of datasets in NASA’s OSDR provides a unique opportunity for the scientific community, as well as citizen scientists and students, to continue using OSDR resources to further unlock profound insights into the consequences of space travel on the human body. Through implementation of security measures to protect sensitive human data, the OSDR seeks to strengthen the science exchange between the Biological and Physical Sciences Program and the Human Research Program, per recommendation 4-1 of the 2023-2032 Decadal Survey, and encourage further sharing and dissemination of astronaut data to provide the scientific community with the resources needed to lay the groundwork for developing targeted mitigation strategies to help withstand the rigors of long-duration spaceflight.

Amanda Marie Saravia-butler↗

Positron emission tomography harmonization in the Alzheimer's Disease Neuroimaging Initiative: A scalable and rigorous approach to multisite amyloid and tau quantification

Abstract INTRODUCTION A key goal of the Alzheimer's Disease NeuroImaging Initiative (ADNI) positron emission tomography (PET) Core is to harmonize quantification of β‐amyloid (Aβ) and tau PET image data across multiple scanners and tracers. METHODS We developed an analysis pipeline (Berkeley PET Imaging Pipeline, B‐PIP) for ADNI Aβ and tau PET images and applied it to PET data from other multisite studies. Steps include image pre‐processing, refacing, magnetic resonance imaging (MRI)/PET co‐registration, visual quality control (QC), quantification of tracer uptake, and standardization of Aβ and tau standardized uptake value ratios (SUVrs) across tracers. RESULTS Measurements from 10,105 cross‐sectional and longitudinal Aβ and tau PET scans acquired in several studies between 2010 and 2024 can be processed, harmonized, and directly merged across tracers and cohorts. DISCUSSION The B‐PIP developed in ADNI is a scalable image harmonization approach used in several observational studies and clinical trials that facilitates rigorous Aβ and tau PET quantification and data sharing. Highlights Quantitative results from ADNI Aβ and tau PET data are generated using a rigorous, scalable image processing pipeline This pipeline has been applied to PET data from several other large, multisite studies and trials Quantitative outcomes are harmonizable across studies and are shared with the scientific community

Neurosciences & Neurology↗