Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data Curation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

NASA Life Sciences Portal (NLSP): Supporting Scientific Transparency and Reproducibility

NASA’s Life Sciences Ports (NLSP) serves the scientific community by providing curated data from space life science experiment. The Human Research Program (HRP) with the help of NLSP is currently transforming their life sciences data archive systems and processes to improve compliance with the FAIR principles [1]. Some of these improvements will at the same time support the twin pillars of Open Science [2]: transparency of methods and reproducibility of results. Scientific transparency is marked by the easily intelligible communication of what has been investigated: what were the procedures for collecting sample and the characteristics of samples collected? what kinds of measurements were made, what were the environmental conditions of the measurements? What were the analysis techniques of the collected data? Reproducibility of the results and findings from the investigation requires a high level of transparency for all but the simplest investigations; the slightest deviation in communicating and replicating complex experimental procedures or data analyses can often yield quite different data and even findings, thwarting their validation. One of the ways the NLSP is aiming to improve the communication of scientific information is through the use of ontology-driven metadata. Ontologies are powerful, graph-based knowledge representation structures, which can be leveraged to increase data interoperability, the area of the FAIR principles in which many data systems most lack compliance. Over the past decade, there has been a concerted effort in the biomedical community to develop modular and narrowly focused domain and application-specific ontologies in a common, open-source framework, the Open Biological and Biomedical Ontology (OBO) Foundry [3]. The open sharing and modular nature of this effort promises huge increases in harmonized data sharing for systems that leverage these models. Which is in line with the FAIR Data Principles of Findability, Accessibility, Interoperability, and Reuse for scientific data management and stewardship.

Life Sciences data↗

NLSP: NASA Life Sciences Portal

NASA’s Life Sciences Ports (NLSP) serves the scientific community by providing curated data from space life science experiment. The Human Research Program (HRP) with the help of NLSP is currently transforming their life sciences data archive systems and processes to improve compliance with the FAIR principles. Some of these improvements will at the same time support the twin pillars of Open Science: transparency of methods and reproducibility of results. This video is a high level overview of the NLSP for existing and new users.

Life Sciences data↗

Technical Tension Between Achieving Particulate and Molecular Organic Environmental Cleanliness: Data from Astromaterial Curation Laboratories

NASA Johnson Space Center operates clean curation facilities for Apollo lunar, Antarctic meteorite, stratospheric cosmic dust, Stardust comet and Genesis solar wind samples. Each of these collections is curated separately due unique requirements. The purpose of this abstract is to highlight the technical tensions between providing particulate cleanliness and molecular cleanliness, illustrated using data from curation laboratories. Strict control of three components are required for curating samples cleanly: a clean environment; clean containers and tools that touch samples; and use of non-shedding materials of cleanable chemistry and smooth surface finish. This abstract focuses on environmental cleanliness and the technical tension between achieving particulate and molecular cleanliness. An environment in which a sample is manipulated or stored can be a room, an enclosed glovebox (or robotic isolation chamber) or an individual sample container.

Allton, J. H.↗

AI Curation Methods for NASA Scientific Data

The NASA Open Science Data Repository (OSDR) serves as a central hub for sharing and accessing NASA's vast collection of scientific data, supporting researchers across diverse fields. To enhance the efficiency, accuracy, and accessibility of this data, we are leveraging advanced artificial intelligence (AI) techniques as part of the AI for Curation project. By integrating large language models (LLMs) into our data curation workflow, we aim to streamline the entire process—from data submission to user interaction. This initiative focuses on improving key areas, including data ingestion, curation, and user engagement with curated datasets, impacting multiple domains and a wide user base. First, we are developing tools that can automatically parse data in various formats, using LLMs to convert unstructured data into structured, standardized formats. This reduces the manual effort required for curation, allowing curators to focus on more critical scientific analyses. Additionally, AI and machine learning (ML) models are being implemented to automate data validation and verification, ensuring the highest standards of data quality and reliability. Finally, we are creating a conversational AI agent to interact with the curated scientific studies in OSDR, helping users easily navigate the repository and access relevant data. By enhancing data discoverability and accessibility, these advancements will foster new research opportunities and promote the principles of open science.

Walter Alvarado↗

Auto-Curation of Seismic Event Data for Signal Denoising

Denoising contaminated seismic signals for later processing is a fundamental problem in seismic signals analysis. Neural network approaches have shown success denoising local signals when trained on short-time Fourier transform spectrograms. One challenge of this approach is the onerous process of hand-labeling event signals for training. By leveraging the SCALODEEP seismic event detector, we develop an automated set of techniques for labeling event data. Despite region specific challenges, training the neural network denoiser on machine curated events shows comparable performance to the neural network trained on hand curated events. We showcase our technique with two experiments, one using Utah regional data and one using regional data from the Korean peninsula.

58 GEOSCIENCES↗

Configuration and Implementation

The main objectives of a Planetary Data System (PDS) are to curate planetary data - that is to store and maintain complete planetary data sets and the necessary supporting documentation to make them useful - and to facilitate scientific study of these data through improved organization and access. A hardware system, no matter how sophisticated and supportive, cannot replace the live colleague interaction needed for good science.System users with only a general knowledge of a data file or data set will need guidance in determining whether that data is most appropriate to support a theory. The PDS is meant to allow easy access to data, to provide useful presentations of that data, and to facilitate data analysis; however it is not designed to supercede the critical examination and interactions that data analysis requires. The Planetary Data System will be a powerful research tool which allows the entire planetary science community access to a complete data and will facilitate the archiving of pass, current, and future mission data sets.

Source record↗

WormBase in 2022—data, processes, and tools for analyzing Caenorhabditis elegans

WormBase (www.wormbase.org) is the central repository for the genetics and genomics of the nematode Caenorhabditis elegans. We provide the research community with data and tools to facilitate the use of C. elegans and related nematodes as model organisms for studying human health, development, and many aspects of fundamental biology. Throughout our 22-year history, we have continued to evolve to reflect progress and innovation in the science and technologies involved in the study of C. elegans. We strive to incorporate new data types and richer data sets, and to provide integrated displays and services that avail the knowledge generated by the published nematode genetics literature. Here, we provide a broad overview of the current state of WormBase in terms of data type, curation workflows, analysis, and tools, including exciting new advances for analysis of single-cell data, text mining and visualization, and the new community collaboration forum. Concurrently, we continue the integration and harmonization of infrastructure, processes, and tools with the Alliance of Genome Resources, of which WormBase is a founding member.

59 BASIC BIOLOGICAL SCIENCES↗

Curation and Dissemination of Complex Multi-Modal Datasets for Radiation Detection, Localization, and Tracking

The PANDAWN sensor network in Chicago, IL, is a state-of-the-art testbed for networked, multi-modal sensing. It integrates AI/data science methods into its operation, from data acquisition to automated data labeling and curation workflows. The curation and dissemination of diverse multi-modal datasets will enable the development of new radiological/nuclear (R/N) detection, localization, and tracking algorithms and methods relevant across the nonproliferation mission space. This article first introduces the PANDAWN sensor network and the features that make it stand out from previous multi-modal data acquisition efforts. We then review the various data streams acquired on the PANDAWN nodes and present the implementation of an automated data curation pipeline that includes the labeling of radiation and contextual data streams. Here, we finally provide a short overview of different studies that leveraged the curated datasets.

Data curation↗

Managing Multi-Instrument Data Streams in Secure Environments

The capture and curation of all primary instrument data is a potentially valuable source of added insight into experiments or diagnostics in laboratory experiments. The data can, when properly curated, enable analysis beyond the current practice that uses just a subset of the as-measured data. Complete curated data can also be input for machine learning and other data exploration tools. Conveniently storing and accessing instrument data requires that the instruments are connected to databases and users through a networking infrastructure. This infrastructure needs to accommodate a wide array of instruments which can range from single laboratory mounted probes for environment monitoring to computers managing multiple instruments. These resources may also include mobile devices on which researchers record instrument and experiment state related notes. These varied data sources bring with them the challenges of different communications capabilities and protocols as well as the primary data typically being produced in proprietary formats. These challenges are further compounded when the instruments need to operate in secure environments such as required in national laboratories. We will discuss the SmartLab, an ongoing effort to set up a system for instrument and simulation data curation at NASA Langley Research Center. We will outline the challenges faced in managing the data sources required for ongoing research activities and the solutions that are being considered and implemented to address those challenges.

instrument data management↗

Carbon Storage Open Database

The Carbon Storage Open Database is a collection of spatial data obtained from publicly available sources published by several NATCARB Partnerships and other organizations. The carbon storage open database was collected from open-source data on ArcREST servers and websites in 2018, 2019, 2021, and 2022. The original database was published on the former GeoCube, which is now EDX Spatial, in July 2020, and has since been updated with additional data resources from the Energy Data eXchange (EDX) and external public data resources. The shapefile geodatabase is available in total, and has also been split up into multiple databases based on the maps produced for EDX spatial. These are topical map categories that describe the type of data, and sometimes the region for which the data relates. The data is separated in case there is only a specific area or data type that is of interest for download. In addition to the geodatabases, this submission contains: 1. A ReadMe file describing the processing steps completed to collect and curate the data. 2. A data catalog of all feature layers within the database. Additional published resources are available that describe the work done to produce the geodatabase: Morkner, P., Bauer, J., Creason, C., Sabbatino, M., Wingo, P., Greenburg, R., Walker, S., Yeates, D., Rose, K. 2022. Distilling Data to Drive Carbon Storage Insights. Computers & Geosciences. https://doi.org/10.1016/j.cageo.2021.104945 Morkner, P., Bauer, J., Shay, J., Sabbatino, M., and Rose, K. An Updated Carbon Storage Open Database - Geospatial Data Aggregation to Support Scaling -Up Carbon Capture and Storage. United States: N. p., 2022. Web. https://www.osti.gov/biblio/1890730 Morkner, P., Rose, K., Bauer, J., Rowan, C., Barkhurst, A., Baker, D.V., Sabbatino, M., Bean, A., Creason, C.G., Wingo, P., and Greenburg, R. Tools for Data Collection, Curation, and Discovery to Support Carbon Sequestration Insights. United States: N. p., 2020. Web. https://www.osti.gov/biblio/1777195 Disclaimer: This project was funded by the United States Department of Energy, National Energy Technology Laboratory, in part, through a site support contract. Neither the United States Government nor any agency thereof, nor any of their employees, nor the support contractor, nor any of their employees, makes any warranty, express or implied, or assumes any legal liability or responsibility for the accuracy, completeness, or usefulness of any information, apparatus, product, or process disclosed, or represents that its use would not infringe privately owned rights. Reference herein to any specific commercial product, process, or service by trade name, trademark, manufacturer, or otherwise does not necessarily constitute or imply its endorsement, recommendation, or favoring by the United States Government or any agency thereof. The views and opinions of authors expressed herein do not necessarily state or reflect those of the United States Government or any agency thereof.

carbon storage↗

Data Albums: An Event Driven Search, Aggregation and Curation Tool for Earth Science

Approaches used in Earth science research such as case study analysis and climatology studies involve discovering and gathering diverse data sets and information to support the research goals. To gather relevant data and information for case studies and climatology analysis is both tedious and time consuming. Current Earth science data systems are designed with the assumption that researchers access data primarily by instrument or geophysical parameter. In cases where researchers are interested in studying a significant event, they have to manually assemble a variety of datasets relevant to it by searching the different distributed data systems. This paper presents a specialized search, aggregation and curation tool for Earth science to address these challenges. The search rool automatically creates curated 'Data Albums', aggregated collections of information related to a specific event, containing links to relevant data files [granules] from different instruments, tools and services for visualization and analysis, and information about the event contained in news reports, images or videos to supplement research analysis. Curation in the tool is driven via an ontology based relevancy ranking algorithm to filter out non relevant information and data.

Ramachandran, Rahul↗

Data Albums: An Event Driven Search, Aggregation and Curation Tool for Earth Science

One of the largest continuing challenges in any Earth science investigation is the discovery and access of useful science content from the increasingly large volumes of Earth science data and related information available. Approaches used in Earth science research such as case study analysis and climatology studies involve gathering discovering and gathering diverse data sets and information to support the research goals. Research based on case studies involves a detailed description of specific weather events using data from different sources, to characterize physical processes in play for a specific event. Climatology-based research tends to focus on the representativeness of a given event, by studying the characteristics and distribution of a large number of events. This allows researchers to generalize characteristics such as spatio-temporal distribution, intensity, annual cycle, duration, etc. To gather relevant data and information for case studies and climatology analysis is both tedious and time consuming. Current Earth science data systems are designed with the assumption that researchers access data primarily by instrument or geophysical parameter. Those who know exactly the datasets of interest can obtain the specific files they need using these systems. However, in cases where researchers are interested in studying a significant event, they have to manually assemble a variety of datasets relevant to it by searching the different distributed data systems. In these cases, a search process needs to be organized around the event rather than observing instruments. In addition, the existing data systems assume users have sufficient knowledge regarding the domain vocabulary to be able to effectively utilize their catalogs. These systems do not support new or interdisciplinary researchers who may be unfamiliar with the domain terminology. This paper presents a specialized search, aggregation and curation tool for Earth science to address these existing challenges. The search tool automatically creates curated "Data Albums", aggregated collections of information related to a specific science topic or event, containing links to relevant data files (granules) from different instruments; tools and services for visualization and analysis; and information about the event contained in news reports, images or videos to supplement research analysis. Curation in the tool is driven via an ontology based relevancy ranking algorithm to filter out non-relevant information and data.

Ramachandran, Rahul↗

Importance of Engineered and Learned Molecular Representations in Predicting Organic Reactivity, Selectivity, and Chemical Properties

Machine-readable chemical structure representations are foundational in all attempts to harness machine learning for the prediction of reactivities, selectivities, and chemical properties directly from molecular structure. The featurization of discrete chemical structures into a continuous vector space is a critical phase undertaken before model selection, and the development of new ways to quantitatively encode molecules is an active area of research. Here, we highlight the application and suitability of different representations, from expert-guided “engineered” descriptors to automatically “learned” features, in different prediction tasks relevant to organic and organometallic chemistry, where differing amounts of training data are available. These tasks include statistical models of stereo- and enantioselectivity, thermochemistry, and kinetics developed using experimental and quantum chemical data. The use of expert-guided molecular descriptors provides an opportunity to incorporate chemical knowledge, domain expertise, and physical constraints into statistical modeling. In applications to stereoselective organic and organometallic catalysis, where data sets may be relatively small and 3D-geometries and conformations play an important role, mechanistically informed features can be used successfully to obtain predictive statistical models that are also chemically interpretable. We provide an overview of several recent applications of this approach to obtain quantitative models for reactivity and selectivity, where topological descriptors, quantum mechanical calculations of electronic and steric properties, along with conformational ensembles, all feature as essential ingredients of the molecular representations used. Alternatively, more flexible, general-purpose molecular representations such as attributed molecular graphs can be used with machine learning approaches to learn the complex relationship between a structure and prediction target. This approach has the potential to out-perform more traditional representation methods such as “hand-crafted” molecular descriptors, particularly as data set sizes grow. One area where this is particularly relevant is in the use of large sets of quantum mechanical data to train quantitative structure–property relationships. A general approach toward curating useful data sets and training highly accurate graph neural network models is discussed in the context of organic bond dissociation enthalpies, where this strategy outperforms regression using precomputed descriptors. Finally, we describe how graph neural network predictions can be incorporated into mechanistically informed statistical models of chemical reactivity and selectivity. Once trained, this approach avoids the expensive computational overhead associated with quantum mechanical calculations, while maintaining chemical interpretability. We illustrate examples for which fast predictions of bond dissociation enthalpy and of the identities of radicals formed through cleavage of a molecule’s weakest bond are used in simple physical models of site-selectivity and reactivity.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗