Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “community data standard”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Recent trends in geographic information system research

This paper reviews recent contributions to the body of published research on Geographic Information Systems (GISs). Increased usages of GISs have placed a new demand upon the academic and research community and despite some lack of formalized definitions, categorizations, terminologies, and standard data structures, the community has risen to the challenge. Examinations of published GIS research, in particular on GIS data structures, reveal a healthy, active research community which is using a truly interdisciplinary approach. Future work will undoubtably lead to a clearer understanding of the problems of handling spatial data, while producing a new generation of highly sophisticated GISs.

Clarke, K. C.↗

A Guide to Using GitHub for Developing and Versioning Data Standards and Reporting Formats

Abstract Data standardization combined with descriptive metadata facilitate data reuse, which is the ultimate goal of the Findable, Accessible, Interoperable, and Reusable (FAIR) principles. Community data or metadata standards are increasingly created through an approach that emphasizes collaboration between various stakeholders. Such an approach requires platforms for collaboration on the development process that centers on sharing information and receiving feedback. Our objective in this study was to conduct a systematic review to identify data standards and reporting formats that use version control for developing data standards and to summarize common practices, particularly in earth and environmental sciences. Out of 108 data standards and reporting formats identified in our review, 32 used GitHub as the version control platform, and no other platforms were used. We found no universally accepted methodology for developing and publishing data standards. Many GitHub repositories did not use key features that could help developers to gather user feedback, or to create and revise standards that build on previous work. We provide guidance for community‐driven standard development and associated documentation on GitHub based on a systematic review of existing practices.

54 ENVIRONMENTAL SCIENCES↗

Data Format Standardization of Space Weather Model Output at the Community Coordinated Modeling Center

The disparate nature of space weather model output provides many challenges with regards to the portability and reuse of not only the data itself, but also any tools that are developed for analysis and visualization. We are developing and implementing a comprehensive data format standardization methodology that allows heterogeneous model output data to be stored uniformly in any common science data format. We will discuss our approach to identifying core meta-data elements that can be used to supplement raw model output data, thus creating self-descriptive files. The meta-data should also contain information describing the simulation grid. This will ultimately assists in the development of efficient data access tools capable of extracting data at any given point and time. We will also discuss our experiences standardizing the output of two global magnetospheric models, and how we plan to apply similar procedures when standardizing the output of the solar, heliospheric, and ionospheric models that are also currently hosted at the Community Coordinated Modeling Center.

Maddox, M.↗

ESIP Information Quality Cluster (IQC)

The Information Quality Cluster (IQC) within the Federation of Earth Science Information Partners (ESIP) was initially formed in 2011 and has evolved significantly over time. The current objectives of the IQC are to: 1. Actively evaluate community data quality best practices and standards; 2. Improve capture, description, discovery, and usability of information about data quality in Earth science data products; 3. Ensure producers of data products are aware of standards and best practices for conveying data quality, and data providers distributors intermediaries establish, improve and evolve mechanisms to assist users in discovering and understanding data quality information; and 4. Consistently provide guidance to data managers and stewards on how best to implement data quality standards and best practices to ensure and improve maturity of their data products. The activities of the IQC include: 1. Identification of additional needs for consistently capturing, describing, and conveying quality information through use case studies with broad and diverse applications; 2. Establishing and providing community-wide guidance on roles and responsibilities of key players and stakeholders including users and management; 3. Prototyping of conveying quality information to users in a more consistent, transparent, and digestible manner; 4. Establishing a baseline of standards and best practices for data quality; 5. Evaluating recommendations from NASA's DQWG in a broader context and proposing possible implementations; and 6. Engaging data providers, data managers, and data user communities as resources to improve our standards and best practices. Following the principles of openness of the ESIP Federation, IQC invites all individuals interested in improving capture, description, discovery, and usability of information about data quality in Earth science data products to participate in its activities.

data products↗

Science Data Center concepts for moderate-sized NASA missions

The paper describes the approaches taken by the NASA Science Data Operations Center to the concepts for two future NASA moderate-sized missions, the Orbiting Solar Laboratory (OSL) and the Tropical Rainfall Measuring Mission (TRMM). The OSL space science mission will be a free-flying spacecraft with a complement of science instruments, placed in a high-inclination, sun synchronous orbit to allow continuous study of the sun for extended periods. The TRMM is planned to be a free-flying satellite for measuring tropical rainfall and its variations. Both missions will produce 'standard' data products for the benefit of their communities, and both depend upon their own scientific community to provide algorithms for generating the standard data products.

Price, R.↗

py4DSTEM: A Software Package for Four-Dimensional Scanning Transmission Electron Microscopy Data Analysis

Scanning transmission electron microscopy (STEM) allows for imaging, diffraction, and spectroscopy of materials on length scales ranging from microns to atoms. By using a high-speed, direct electron detector, it is now possible to record a full two-dimensional (2D) image of the diffracted electron beam at each probe position, typically a 2D grid of probe positions. These 4D-STEM datasets are rich in information, including signatures of the local structure, orientation, deformation, electromagnetic fields, and other sample-dependent properties. However, extracting this information requires complex analysis pipelines that include data wrangling, calibration, analysis, and visualization, all while maintaining robustness against imaging distortions and artifacts. In this paper, we present py4DSTEM, an analysis toolkit for measuring material properties from 4D-STEM datasets, written in the Python language and released with an open-source license. We describe the algorithmic steps for dataset calibration and various 4D-STEM property measurements in detail and present results from several experimental datasets. We also implement a simple and universal file format appropriate for electron microscopy data in py4DSTEM, which uses the open-source HDF5 standard. We hope this tool will benefit the research community and help improve the standards for data and computational methods in electron microscopy, and we invite the community to contribute to this ongoing project.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

PDBx/mmCIF Ecosystem: Foundational Semantic Tools for Structural Biology

PDBx/mmCIF, Protein Data Bank Exchange (PDBx) macromolecular Crystallographic Information Framework (mmCIF), has become the data standard for structural biology. With its early roots in the domain of small-molecule crystallography, PDBx/mmCIF provides an extensible data representation that is used for deposition, archiving, remediation, and public dissemination of experimentally determined three-dimensional (3D) structures of biological macromolecules by the Worldwide Protein Data Bank (wwPDB, wwpdb.org). Extensions of PDBx/mmCIF are similarly used for computed structure models by ModelArchive (modelarchive.org), integrative/hybrid structures by PDB-Dev (pdb-dev.wwpdb.org), small angle scattering data by Small Angle Scattering Biological Data Bank SASBDB (sasbdb.org), and for models computed generated with the AlphaFold 2.0 deep learning software suite (alphafold.ebi.ac.uk). Community-driven development of PDBx/mmCIF spans three decades, involving contributions from researchers, software and methods developers in structural sciences, data repository providers, scientific publishers, and professional societies. Having a semantically rich and extensible data framework for representing a wide range of structural biology experimental and computational results, combined with expertly curated 3D biostructure data sets in public repositories, accelerates the pace of scientific discovery. Herein, we describe the architecture of the PDBx/mmCIF data standard, tools used to maintain representations of the data standard, governance, and processes by which data content standards are extended, plus community tools/software libraries available for processing and checking the integrity of PDBx/mmCIF data. Use cases exemplify how the members of the Worldwide Protein Data Bank have used PDBx/mmCIF as the foundation for its pipeline for delivering Findable, Accessible, Interoperable, and Reusable (FAIR) data to many millions of users worldwide.

59 BASIC BIOLOGICAL SCIENCES↗

ISAIA: Interoperable Systems for Archival Information Access

The ISAIA project was originally proposed in 1999 as a successor to the informal AstroBrowse project. AstroBrowse, which provided a data location service for astronomical archives and catalogs, was a first step toward data system integration and interoperability. The goals of ISAIA were ambitious: '...To develop an interdisciplinary data location and integration service for space science. Building upon existing data services and communications protocols, this service will allow users to transparently query hundreds or thousands of WWW-based resources (catalogs, data, computational resources, bibliographic references, etc.) from a single interface. The service will collect responses from various resources and integrate them in a seamless fashion for display and manipulation by the user.' Funding was approved only for a one-year pilot study, a decision that in retrospect was wise given the rapid changes in information technology in the past few years and the emergence of the Virtual Observatory initiatives in the US and worldwide. Indeed, the ISAIA pilot study was influential in shaping the science goals, system design, metadata standards, and technology choices for the virtual observatory. The ISAIA pilot project also helped to cement working relationships among the NASA data centers, US ground-based observatories, and international data centers. The ISAIA project was formed as a collaborative effort between thirteen institutions that provided data to astronomers, space physicists, and planetary scientists. Among the fruits we ultimately hoped would come from this project would be a central site on the Web that any space scientist could use to efficiently locate existing data relevant to a particular scientific question. Furthermore, we hoped that the needed technology would be general enough to allow smaller, more-focused community within space science could use the same technologies and standards to provide more specialized services. A major challenge to searching for data across a broad community is that information that describe some data products are either not relevant to other data or not applicable in the same way. Some previous metadata standard development efforts (e.g., in the earth science and library communities) have produced standards that are very large and difficult to support. To address this problem, we studied how a standard may be divided into separable pieces. Data providers that wish to participate in interoperable searches can support only those parts of the standard that are relevant to them. We prototyped a top-level metadata standard that was small and applicable to all space science data.

Hanisch, Robert J.↗

Enabling Open and Interoperable Science: Multi-Omics Data Processing Platform with NASA GeneLab Standardized Bioinformatics Workflows for Space and Earth Research

Multi-omics biological data continues to be generated at an astounding pace. Genomics, transcriptomics, metabolomics, and proteomics, or collectively known as multi-omics data, are used to assess biological functions, and provide invaluable insights into human, animal, plant, and environmental health both on Earth and in Space. Despite the abundance of these valuable data, the need for bioinformatics expertise, particularly as it relates to the niche filed of space biology, and a lack of accessible resources for processing these data limit their usefulness in deriving biological insights. The NASA Open Science Data Repository (OSDR) provides access to omics data from various spaceflight and analog studies. To enhance the accessibility and reusability of these data, GeneLab (part of OSDR) designs and implements standardized, community-driven, open-source bioinformatics workflows to transform raw omics data into standardized processed data. Currently, GeneLab-processed data from hundreds of space studies have been reused for meta-analyses. This has led to new insights and scientific publications that extend beyond the initial research, thereby enriching our understanding of molecular-scale biological responses to the space environment. To make these bioinformatics workflows open and accessible, GeneLab teamed up with DOE-funded initiatives, including the National Microbiome Data Collaborative (NMDC), to create the NASA EDGE [Empowering the Development of Genomics Expertise] Bioinformatics web-based platform. NASA EDGE utilizes shared compute resources to run the GeneLab standardized bioinformatics workflows, which eliminates the need for researchers to have their own high performance computing cluster. The web-based platform makes complicated biological analyses incredibly easy to perform, thus expanding the reach of these analyses to bioinformatics novices, students, and even citizen scientists enabling them to contribute to scientific discoveries and progress. The authors will demonstrate how the NASA EDGE platform can be used to process microbial omics data hosted on OSDR as well as user-generated omics datasets using GeneLab’s standard workflows.

Amanda M. Saravia-Butler↗

MIBiG 4.0: advancing biosynthetic gene cluster curation through global collaboration

Specialized or secondary metabolites are small molecules of biological origin, often showing potent biological activities with applications in agriculture, engineering and medicine. Usually, the biosynthesis of these natural products is governed by sets of co-regulated and physically clustered genes known as biosynthetic gene clusters (BGCs). To share information about BGCs in a standardized and machine-readable way, the Minimum Information about a Biosynthetic Gene cluster (MIBiG) data standard and repository was initiated in 2015. Since its conception, MIBiG has been regularly updated to expand data coverage and remain up to date with innovations in natural product research. Here, we describe MIBiG version 4.0, an extensive update to the data repository and the underlying data standard. In a massive community annotation effort, 267 contributors performed 8304 edits, creating 557 new entries and modifying 590 existing entries, resulting in a new total of 3059 curated entries in MIBiG. Particular attention was paid to ensuring high data quality, with automated data validation using a newly developed custom submission portal prototype, paired with a novel peer-reviewing model. MIBiG 4.0 also takes steps towards a rolling release model and a broader involvement of the scientific community. MIBiG 4.0 is accessible online at https://mibig.secondarymetabolites.org/.

59 BASIC BIOLOGICAL SCIENCES↗

Reproducibility and Repeatability of Tensile and Low-Cycle Fatigue Properties in Propulsion Grade Hydrogen

Hydrogen has the potential of increased use in the future as an environmentally friendly fuel. It has, however, shown a tendency to embrittle some materials. To be used in a safe manner and to exploit its full potential, it will be necessary to develop a database of material properties in hydrogen environment. The tests needed to produce this data are costly to perform (tensile test cost 25 times more and low cycle fatigue test are 55 times as expensive). Moreover, there is presently a lack of universal test methods to ensure standardized data within the hydrogen community. Each of the industries that work with hydrogen (aerospace, petroleum, fuel cells, etc.) performs tests by their own laboratory-developed methods, thus rendering cross- comparisons of material property data highly questionable. It is extremely important that data generated in a hydrogen environment be done to a standard that reduces variance to a minimum and allows direct comparison of test results from different laboratories. Doing so will assure that all data generated can be used to further our understanding of the hydrogen effects and to make sure components/products designed for hydrogen are the safest and most reliable possible. This paper reviews the results of two 'round-robin' programs conducted by NASA-MSFC. These two programs examined the reproducibility and repeatability of tensile and low-cycle fatigue test results in high-pressure hydrogen environments. The studies indicated that even with the tightest controls available from current commercial standards, the reproducibility (between different laboratories) and repeatability (within a laboratory) results of the tensile tests exhibited five times the variance as in standard ambient air tests. The variance with the LCF tests were on the same order as with air tests, but that was due to the large variation present in the last Interlaboratory air program. The paper concludes with a recommendation for a program that would allow the development of improved test methods, leading to lower variance in the generation of mechanical property data in the future.

Vesely, E. J.↗

Building a FAIR data ecosystem for incorporating single-cell transcriptomics data into agricultural genome to phenome research

Introduction The agriculture genomics community has numerous data submission standards available, but the standards for describing and storing single-cell (SC, e.g., scRNA- seq) data are comparatively underdeveloped. Methods To bridge this gap, we leveraged recent advancements in human genomics infrastructure, such as the integration of the Human Cell Atlas Data Portal with Terra, a secure, scalable, open-source platform for biomedical researchers to access data, run analysis tools, and collaborate. In parallel, the Single Cell Expression Atlas at EMBL-EBI offers a comprehensive data ingestion portal for high-throughput sequencing datasets, including plants, protists, and animals (including humans). Developing data tools connecting these resources would offer significant advantages to the agricultural genomics community. The FAANG data portal at EMBL-EBI emphasizes delivering rich metadata and highly accurate and reliable annotation of farmed animals but is not computationally linked to either of these resources. Results Herein, we describe a pilot-scale project that determines whether the current FAANG metadata standards for livestock can be used to ingest scRNA-seq datasets into Terra in a manner consistent with HCA Data Portal standards. Importantly, rich scRNA-seq metadata can now be brokered through the FAANG data portal using a semi-automated process, thereby avoiding the need for substantial expert curation. We have further extended the functionality of this tool so that validated and ingested SC files within the HCA Data Portal are transferred to Terra for further analysis. In addition, we verified data ingestion into Terra, hosted on Azure, and demonstrated the use of a workflow to analyze the first ingested porcine scRNA-seq dataset. Additionally, we have also developed prototype tools to visualize the output of scRNA-seq analyses on genome browsers to compare gene expression patterns across tissues and cell populations. This JBrowse tool now features distinct tracks, showcasing PBMC scRNA-seq alongside two bulk RNA-seq experiments. Discussion We intend to further build upon these existing tools to construct a scientist-friendly data resource and analytical ecosystem based on Findable, Accessible, Interoperable, and Reusable (FAIR) SC principles to facilitate SC-level genomic analysis through data ingestion, storage, retrieval, re-use, visualization, and comparative annotation across agricultural species.

Genetics & Heredity↗

First Plant Cell Atlas symposium report

The Plant Cell Atlas (PCA) community hosted a virtual symposium on December 9 and 10, 2021 on single cell and spatial omics technologies. The conference gathered almost 500 academic, industry, and government leaders to identify the needs and directions of the PCA community and to explore how establishing a data synthesis center would address these needs and accelerate progress. This report details the presentations and discussions focused on the possibility of a data synthesis center for a PCA and the expected impacts of such a center on advancing science and technology globally. Community discussions focused on topics such as data analysis tools and annotation standards; computational expertise and cyber-infrastructure; modes of community organization and engagement; methods for ensuring a broad reach in the PCA community; recruitment, training, and nurturing of new talent; and the overall impact of the PCA initiative. These targeted discussions facilitated dialogue among the participants to gauge whether PCA might be a vehicle for formulating a data synthesis center. The conversations also explored how online tools can be leveraged to help broaden the reach of the PCA (i.e., online contests, virtual networking, and social media stakeholder engagement) and decrease costs of conducting research (e.g., virtual REU opportunities). Major recommendations for the future of the PCA included establishing standards, creating dashboards for easy and intuitive access to data, and engaging with a broad community of stakeholders. The discussions also identified the following as being essential to the PCA's success: identifying homologous cell-type markers and their biocuration, publishing datasets and computational pipelines, utilizing online tools for communication (such as Slack), and user-friendly data visualization and data sharing. In conclusion, the development of a data synthesis center will help the PCA community achieve these goals by providing a centralized repository for existing and new data, a platform for sharing tools, and new analytical approaches through collaborative, multidisciplinary efforts. A data synthesis center will help the PCA reach milestones, such as community-supported data evaluation metrics, accelerating plant research necessary for human and environmental health.

59 BASIC BIOLOGICAL SCIENCES↗

Perspectives for artificial intelligence in bioprocess automation

Recent advances in artificial intelligence (AI) have rapidly changed the lab automation landscape, promoting self-driving laboratories (SDLs) that enable autonomous scientific discovery. These trends are increasingly applied in bioprocess development, yet bioprocessing faces unique challenges - biological complexity, regulatory and safety requirements, and multiscale experimentation - that distinguish it from other automation domains. Rather than pursuing full autonomy, we foresee that hybrid SDLs, combining AI-driven decision-making with sustained human oversight, represent the most practical near-term trajectory. This review examines three interconnected perspectives: (i) hybrid human-machine decision-making for bioprocessing; (ii) laboratory design considerations in the era of AI; and (iii) scale-up challenges when transitioning from screening to manufacturing. We highlight critical gaps in data standardization and the required community efforts necessary to realize autonomous bioprocess innovation.

Helleckes, Laura Marie↗

Principles of the Battery Data Genome

Batteries are central to modern society. They are no longer just a convenience but a critical enabler of the transition to a resilient, low-carbon economy. Battery development capabilities are provided by communities spanning materials discovery, battery chemistry and electrochemistry, cell and pack design, scale-up, manufacturing, and deployments. Despite their relative maturity, data-science practices among these diverse groups are far behind the state of the art in other fields, which have demonstrated an ability to significantly improve innovation and economic impact. The negative consequences of the present paradigm include incremental improvements but few breakthroughs, significant manufacturing uncertainties, and cascading investment risks that collectively slow deployments. The primary roadblock to a battery-data-science renaissance is the requirement for large amounts of high-quality data, which are not available in the current fragmented ecosystem. Here, in this study, we identify gaps and propose principles that enable the solution by building a robust community of data hubs with standardized practices and flexible sharing options that will seed advanced tools spanning innovation to deployment. Precedents are offered that demonstrate that both public good and immense economic gains will arise from sharing valuable battery data. The proposed Battery Data Genome looks to broadly transform innovations and revolutionize their translation from research to societal impact.

25 ENERGY STORAGE↗

Community standards and future opportunities for synthetic communities in plant–microbiota research

Harnessing beneficial microorganisms is seen as a promising approach to enhance sustainable agriculture production. Synthetic communities (SynComs) are increasingly being used to study relevant microbial activities and interactions with the plant host. Yet, the lack of community standards limits the efficiency and progress in this important area of research. Here, to address this gap, we recommend three actions: (1) defining reference SynComs; (2) establishing community standards, protocols and benchmark data for constructing and using SynComs; and (3) creating an infrastructure for sharing strains and data. We also outline opportunities to develop SynCom research through technical advances, linking to field studies, and filling taxonomic blind spots to move towards fully representative SynComs.

59 BASIC BIOLOGICAL SCIENCES↗

A customizable data management framework for high-repetition-rate high-energy-density science

The high-energy-density (HED) physics community is moving toward a new paradigm of high-repetition-rate (HRR) operation. To fully leverage the scientific power of HRR HED facilities, all of the components of each subsystem (laser, targetry, and performance diagnostics) must be connected and synchronized in a reliable and robust manner while the data acquired are tagged and archived in real time. To this end, GA has begun developing a generalized NoSQL-database framework, the MongoDB repository for information and archiving. An organizational strategy has been developed that shifts HED data organization from a shot-based to a diagnostic-based approach in order to increase archival and retrieval efficiency that lends itself to optimization applications. This work is a first step in pushing HRR HED science toward data management solutions that emphasize machine actionability and aim to stimulate community engagement to define data standards in HED science.

Instruments & Instrumentation↗