Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “community data standard”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

PDBx/mmCIF Ecosystem: Foundational Semantic Tools for Structural Biology

PDBx/mmCIF, Protein Data Bank Exchange (PDBx) macromolecular Crystallographic Information Framework (mmCIF), has become the data standard for structural biology. With its early roots in the domain of small-molecule crystallography, PDBx/mmCIF provides an extensible data representation that is used for deposition, archiving, remediation, and public dissemination of experimentally determined three-dimensional (3D) structures of biological macromolecules by the Worldwide Protein Data Bank (wwPDB, wwpdb.org). Extensions of PDBx/mmCIF are similarly used for computed structure models by ModelArchive (modelarchive.org), integrative/hybrid structures by PDB-Dev (pdb-dev.wwpdb.org), small angle scattering data by Small Angle Scattering Biological Data Bank SASBDB (sasbdb.org), and for models computed generated with the AlphaFold 2.0 deep learning software suite (alphafold.ebi.ac.uk). Community-driven development of PDBx/mmCIF spans three decades, involving contributions from researchers, software and methods developers in structural sciences, data repository providers, scientific publishers, and professional societies. Having a semantically rich and extensible data framework for representing a wide range of structural biology experimental and computational results, combined with expertly curated 3D biostructure data sets in public repositories, accelerates the pace of scientific discovery. Herein, we describe the architecture of the PDBx/mmCIF data standard, tools used to maintain representations of the data standard, governance, and processes by which data content standards are extended, plus community tools/software libraries available for processing and checking the integrity of PDBx/mmCIF data. Use cases exemplify how the members of the Worldwide Protein Data Bank have used PDBx/mmCIF as the foundation for its pipeline for delivering Findable, Accessible, Interoperable, and Reusable (FAIR) data to many millions of users worldwide.

59 BASIC BIOLOGICAL SCIENCES↗

MIBiG 4.0: advancing biosynthetic gene cluster curation through global collaboration

Specialized or secondary metabolites are small molecules of biological origin, often showing potent biological activities with applications in agriculture, engineering and medicine. Usually, the biosynthesis of these natural products is governed by sets of co-regulated and physically clustered genes known as biosynthetic gene clusters (BGCs). To share information about BGCs in a standardized and machine-readable way, the Minimum Information about a Biosynthetic Gene cluster (MIBiG) data standard and repository was initiated in 2015. Since its conception, MIBiG has been regularly updated to expand data coverage and remain up to date with innovations in natural product research. Here, we describe MIBiG version 4.0, an extensive update to the data repository and the underlying data standard. In a massive community annotation effort, 267 contributors performed 8304 edits, creating 557 new entries and modifying 590 existing entries, resulting in a new total of 3059 curated entries in MIBiG. Particular attention was paid to ensuring high data quality, with automated data validation using a newly developed custom submission portal prototype, paired with a novel peer-reviewing model. MIBiG 4.0 also takes steps towards a rolling release model and a broader involvement of the scientific community. MIBiG 4.0 is accessible online at https://mibig.secondarymetabolites.org/.

59 BASIC BIOLOGICAL SCIENCES↗

Building a FAIR data ecosystem for incorporating single-cell transcriptomics data into agricultural genome to phenome research

Introduction The agriculture genomics community has numerous data submission standards available, but the standards for describing and storing single-cell (SC, e.g., scRNA- seq) data are comparatively underdeveloped. Methods To bridge this gap, we leveraged recent advancements in human genomics infrastructure, such as the integration of the Human Cell Atlas Data Portal with Terra, a secure, scalable, open-source platform for biomedical researchers to access data, run analysis tools, and collaborate. In parallel, the Single Cell Expression Atlas at EMBL-EBI offers a comprehensive data ingestion portal for high-throughput sequencing datasets, including plants, protists, and animals (including humans). Developing data tools connecting these resources would offer significant advantages to the agricultural genomics community. The FAANG data portal at EMBL-EBI emphasizes delivering rich metadata and highly accurate and reliable annotation of farmed animals but is not computationally linked to either of these resources. Results Herein, we describe a pilot-scale project that determines whether the current FAANG metadata standards for livestock can be used to ingest scRNA-seq datasets into Terra in a manner consistent with HCA Data Portal standards. Importantly, rich scRNA-seq metadata can now be brokered through the FAANG data portal using a semi-automated process, thereby avoiding the need for substantial expert curation. We have further extended the functionality of this tool so that validated and ingested SC files within the HCA Data Portal are transferred to Terra for further analysis. In addition, we verified data ingestion into Terra, hosted on Azure, and demonstrated the use of a workflow to analyze the first ingested porcine scRNA-seq dataset. Additionally, we have also developed prototype tools to visualize the output of scRNA-seq analyses on genome browsers to compare gene expression patterns across tissues and cell populations. This JBrowse tool now features distinct tracks, showcasing PBMC scRNA-seq alongside two bulk RNA-seq experiments. Discussion We intend to further build upon these existing tools to construct a scientist-friendly data resource and analytical ecosystem based on Findable, Accessible, Interoperable, and Reusable (FAIR) SC principles to facilitate SC-level genomic analysis through data ingestion, storage, retrieval, re-use, visualization, and comparative annotation across agricultural species.

Genetics & Heredity↗

First Plant Cell Atlas symposium report

The Plant Cell Atlas (PCA) community hosted a virtual symposium on December 9 and 10, 2021 on single cell and spatial omics technologies. The conference gathered almost 500 academic, industry, and government leaders to identify the needs and directions of the PCA community and to explore how establishing a data synthesis center would address these needs and accelerate progress. This report details the presentations and discussions focused on the possibility of a data synthesis center for a PCA and the expected impacts of such a center on advancing science and technology globally. Community discussions focused on topics such as data analysis tools and annotation standards; computational expertise and cyber-infrastructure; modes of community organization and engagement; methods for ensuring a broad reach in the PCA community; recruitment, training, and nurturing of new talent; and the overall impact of the PCA initiative. These targeted discussions facilitated dialogue among the participants to gauge whether PCA might be a vehicle for formulating a data synthesis center. The conversations also explored how online tools can be leveraged to help broaden the reach of the PCA (i.e., online contests, virtual networking, and social media stakeholder engagement) and decrease costs of conducting research (e.g., virtual REU opportunities). Major recommendations for the future of the PCA included establishing standards, creating dashboards for easy and intuitive access to data, and engaging with a broad community of stakeholders. The discussions also identified the following as being essential to the PCA's success: identifying homologous cell-type markers and their biocuration, publishing datasets and computational pipelines, utilizing online tools for communication (such as Slack), and user-friendly data visualization and data sharing. In conclusion, the development of a data synthesis center will help the PCA community achieve these goals by providing a centralized repository for existing and new data, a platform for sharing tools, and new analytical approaches through collaborative, multidisciplinary efforts. A data synthesis center will help the PCA reach milestones, such as community-supported data evaluation metrics, accelerating plant research necessary for human and environmental health.

59 BASIC BIOLOGICAL SCIENCES↗

Hosting downscaled decision-relevant community data products in ESGF2-US

As regionally-relevant high-resolution Earth system data is increasingly relied upon across scientific, policy, and practitioner communities, there is an urgent need for coordinated and federated infrastructure to store, manage, standardize, and distribute decision-relevant community data products. Substantial effort is required to ensure that these products, which are often critical for regional impact assessments and decision-making, are findable, accessible, interoperable, and reusable. The Earth System Grid Federation US project (ESGF2-US) is addressing this challenge by expanding its open-source, distributed platform to support the hosting and dissemination of downscaled Earth system datasets. This expansion includes aligning new downscaled datasets with developing community standards for metadata and file structure, consistent with existing ESGF archives. This includes ensuring CF-compliance, applying CMORization where appropriate, and developing tools to streamline user access. In this paper, we highlight the technical and coordination work required to bring downscaled data into ESGF2-US and aim to inform the broader Earth system data user community about the growing availability and utility of these curated resources.

ESGF↗

Evolution of standardization and dissemination of cryo-EM structures and data jointly by the community, PDB, and EMDB

Cryogenic electron microscopy (cryo-EM) methods began to be used in the mid-1970s to study thin and periodic arrays of proteins. Following a half-century of development in cryo-specimen preparation, instrumentation, data collection, data processing and modeling software, cryo-EM has become a routine method for solving structures from large biological assemblies to small biomolecules at near to true atomic resolution. This review explores the critical roles played by the Protein Data Bank (PDB) and Electron Microscopy Data Bank (EMDB) in partnership with the community to develop the necessary infrastructure to archive cryo-EM maps and associated models. Public access to cryo-EM structure data has in turn facilitated better understanding of structure-function relationships and advancement of image processing and modeling tool development. The partnership between the global cryo-EM community and PDB and EMDB leadership has synergistically shaped the standards for metadata, one-stop deposition of maps and models, and validation metrics to assess the quality of cryo-EM structures. The advent of cryo-electron tomography (cryo-ET) for in situ molecular cell structures at a broad resolution range and their correlations with other imaging data introduces new data archival challenges in terms of data size and complexity in the years to come.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Perspectives for artificial intelligence in bioprocess automation

Recent advances in artificial intelligence (AI) have rapidly changed the lab automation landscape, promoting self-driving laboratories (SDLs) that enable autonomous scientific discovery. These trends are increasingly applied in bioprocess development, yet bioprocessing faces unique challenges - biological complexity, regulatory and safety requirements, and multiscale experimentation - that distinguish it from other automation domains. Rather than pursuing full autonomy, we foresee that hybrid SDLs, combining AI-driven decision-making with sustained human oversight, represent the most practical near-term trajectory. This review examines three interconnected perspectives: (i) hybrid human-machine decision-making for bioprocessing; (ii) laboratory design considerations in the era of AI; and (iii) scale-up challenges when transitioning from screening to manufacturing. We highlight critical gaps in data standardization and the required community efforts necessary to realize autonomous bioprocess innovation.

Helleckes, Laura Marie↗

Principles of the Battery Data Genome

Batteries are central to modern society. They are no longer just a convenience but a critical enabler of the transition to a resilient, low-carbon economy. Battery development capabilities are provided by communities spanning materials discovery, battery chemistry and electrochemistry, cell and pack design, scale-up, manufacturing, and deployments. Despite their relative maturity, data-science practices among these diverse groups are far behind the state of the art in other fields, which have demonstrated an ability to significantly improve innovation and economic impact. The negative consequences of the present paradigm include incremental improvements but few breakthroughs, significant manufacturing uncertainties, and cascading investment risks that collectively slow deployments. The primary roadblock to a battery-data-science renaissance is the requirement for large amounts of high-quality data, which are not available in the current fragmented ecosystem. Here, in this study, we identify gaps and propose principles that enable the solution by building a robust community of data hubs with standardized practices and flexible sharing options that will seed advanced tools spanning innovation to deployment. Precedents are offered that demonstrate that both public good and immense economic gains will arise from sharing valuable battery data. The proposed Battery Data Genome looks to broadly transform innovations and revolutionize their translation from research to societal impact.

25 ENERGY STORAGE↗

Community standards and future opportunities for synthetic communities in plant–microbiota research

Harnessing beneficial microorganisms is seen as a promising approach to enhance sustainable agriculture production. Synthetic communities (SynComs) are increasingly being used to study relevant microbial activities and interactions with the plant host. Yet, the lack of community standards limits the efficiency and progress in this important area of research. Here, to address this gap, we recommend three actions: (1) defining reference SynComs; (2) establishing community standards, protocols and benchmark data for constructing and using SynComs; and (3) creating an infrastructure for sharing strains and data. We also outline opportunities to develop SynCom research through technical advances, linking to field studies, and filling taxonomic blind spots to move towards fully representative SynComs.

59 BASIC BIOLOGICAL SCIENCES↗

A customizable data management framework for high-repetition-rate high-energy-density science

The high-energy-density (HED) physics community is moving toward a new paradigm of high-repetition-rate (HRR) operation. To fully leverage the scientific power of HRR HED facilities, all of the components of each subsystem (laser, targetry, and performance diagnostics) must be connected and synchronized in a reliable and robust manner while the data acquired are tagged and archived in real time. To this end, GA has begun developing a generalized NoSQL-database framework, the MongoDB repository for information and archiving. An organizational strategy has been developed that shifts HED data organization from a shot-based to a diagnostic-based approach in order to increase archival and retrieval efficiency that lends itself to optimization applications. This work is a first step in pushing HRR HED science toward data management solutions that emphasize machine actionability and aim to stimulate community engagement to define data standards in HED science.

Instruments & Instrumentation↗

Improving the Quality of Geothermal Data Through Data Standards and Pipelines Within the Geothermal Data Repository: Preprint

For machine learning outputs to be applicable to real world problems, high quality data are needed to ensure high quality results. With the more recent emphasis on machine learning in geothermal, there is an increasing need for greater focus on the quality of the data available for use in these projects. For example, Geothermal Operational Optimization Using Machine Learning (GOOML) utilized large quantities of geothermal power plant operational data to inform power plant operational configurations to maximize power generation. High quality datasets result from dependable sensors or devices collecting data, high frequency of measurements, sufficient data points, adequate metadata, reliable storage of data, and sufficient data curation. Another component that contributes to high quality data is reusability, which can be enhanced through data standardization. Data Standardization creates consistency in formatting and contents of like datasets, lessening preprocessing requirements and ensuring adequate information provided by a given dataset. The Geothermal Data Repository (GDR) aims to help improve data quality through automated data standardization for high-value datasets through the implementation of data pipelines alongside reliable and accessible long-term storage for datasets. As such, the GDR has decided to shift away from recommending the use of Excel-based content models and towards the implementation of automated data pipelines. This takes the burden of data standardization off the user and project team and will increase the availability of standardized geothermal data available through the GDR. A set of recommendations, or a data standard for each data type will exist with each data pipeline in order to advise data collection for maximum usability for future research. This paper serves to describe the GDR's proposed transition towards data standardization through automated data pipelines, to discuss the need for and value of such a shift, and to call for suggestions from the community regarding the most useful data standards and pipelines.

data↗

HPC ODA Commons [SWR-26-003]

HPC ODA Commons is a community-driven platform for standardizing HPC operational data analytics. HPC sites generate enormous volumes of operational data - scheduler logs, accounting records, monitoring streams - but turning that data into actionable insight is needlessly hard. Each site builds bespoke parsers, schemas, and evaluation pipelines. Results can't be compared across institutions. Promising analytics ideas stay siloed because there's no shared language for describing the data, the experiments, or the outcomes. HPC ODA Commons fixes this by establishing community-governed contracts - versioned schemas, canonical artifacts, and benchmark recipes - that make ODA workflows discoverable, reproducible, and comparable. It pairs these standards with a practical, CLI-first toolkit that lets operators and researchers go from raw logs to standardized results without sending data off-cluster.

Menear, Kevin [National Laboratory of the Rockies ↗

Templates for developing and versioning data standards and reporting formats using GitHub

This data package contains three templates that can be used for creating README files and Issue Templates, written in the markdown language, that support community-led data reporting formats. We created these templates based on the results of a systematic review (see related references) that explored how groups developing data standard documentation use the Version Control platform GitHub, to collaborate on supporting documents. Based on our review of 32 GitHub repositories, we make recommendations for the content of README Files (e.g., provide a user license, indicate how users can contribute) and so 'README_template.md' includes headings for each section. The two issue templates we include ('issue_template_for_all_other_changes.md' and 'issue_template_for_documentation_change.md') can be used in a GitHub repository to help structure user-submitted issues, or can be modified to suit the needs of data standard developers. We used these templates when establishing ESS-DIVE's community space on GitHub (https://github.com/ess-dive-community) that includes documentation for community-led data reporting formats. We also include file-level metadata 'flmd.csv' that describes the contents of each file within this data package. Lastly, the temporal range that we indicate in our metadata is the time range during which we searched for data standards documented on GitHub.

54 ENVIRONMENTAL SCIENCES↗

The need for standardization and improved open (meta)data practices in metaproteomics

Metaproteomics enables functional insight into microbial communities by identifying and quantifying proteins in complex samples. Yet, heterogeneous analytical workflows and the lack of standardization across experimental and bioinformatics stages hinder reproducibility and comparability, limiting integration with other omics data. We here present a community-developed reporting checklist tailored to the specific needs of metaproteomics. We also outline current efforts to enable structured and interoperable metadata capture, drawing on standards from proteomics and microbiome research wherever possible. By promoting transparent reporting and advancing metadata practices, our recommendations aim to align metaproteomics more closely with FAIR principles and support reproducible and interoperable research practices.

Armengaud, Jean [Universite Paris-Saclay, France]↗

mzPeak: Designing a Scalable, Interoperable, and Future-Ready Mass Spectrometry Data Format

Advances in mass spectrometry (MS) instrumentation, such as higher resolution, faster scan speeds, and improved sensitivity, have significantly increased the volume and complexity of data. The growing adoption of imaging and ion mobility further amplifies these challenges across MS-based omics fields, including proteomics, metabolomics, and lipidomics. While these technologies unlock new possibilities, they also present significant challenges in data management, storage, and accessibility. Existing open formats, such as the XML-based community standards mzML and imzML, struggle to meet the demands of modern MS workflows due to their large file sizes, slow data access, and limited metadata support. Vendor-specific formats, while optimized for proprietary instruments, lack interoperability, comprehensive metadata support and long-term archival reliability. This white paper lays the groundwork for mzPeak, a next-generation community data format designed to address these challenges and support high-throughput, multi-dimensional MS workflows. By adopting a hybrid model that combines efficient binary storage for numerical data and both human and machine-readable metadata storage, mzPeak will reduce file sizes, accelerate data access, and offer a scalable, adaptable solution for evolving MS technologies. For researchers, mzPeak will enable enhanced interoperability across platforms, seamless support for complex workflows including ion mobility and MS imaging, and faster data access compared to existing community formats such as mzML. Its design will ensure data is managed in compliance with regulatory standards, essential for applications such as precision medicine and chemical safety, where long-term data integrity and accessibility are critical. For vendors, mzPeak provides a streamlined, open alternative to proprietary formats, reducing the burden of regulatory compliance while aligning with the industry's push for transparency and standardization. By offering a high-performance, interoperable solution, mzPeak positions vendors to meet customer demands for sustainable data management tools which will be able to handle emerging and future data types and workflows. mzPeak aspires to become the cornerstone of MS data management, empowering researchers, vendors, and developers to innovate and collaborate more effectively.

data formats↗

Review of Experimental Data for Validating Computer Codes Used in Shielding Calculations for Spent Fuel Storage and Transportation Systems

This report presents a review of available radiochemical assay data and shielding benchmarks applicable to spent nuclear fuel (SNF) shielding calculations. The relevant information reviewed herein includes the Spent Fuel Composition (SFCOMPO) database, the Shielding Integral Benchmark Archive and Database (SINBAD), the International Handbook of Evaluated Criticality Safety Benchmark Experiments, and published measurements of external dose rates of casks loaded with SNF. The relevant experimental data identified in this report may be used to support verification and validation of computer codes used in SNF cask/transport shielding applications, as well as development of calculation uncertainties. It should be noted that a relatively small subset of the identified experimental data (e.g., criticality alarm experiments) is available in a standard format established by the international community participating in experimental isotopic and shielding data evaluations. An effort of the SFCOMPO Technical Review Group (TRG) is underway to publish first isotopic evaluations of individual assay data using a standard data evaluation format. The SINBAD TRG has recently initiated benchmark evaluations and modernization of the database. Therefore, more relevant information is expected in the future that will enable users to select quality experimental data in depletion code and shielding code validations for SNF applications.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

A practical approach to using the Genomic Standards Consortium MIxS reporting standard for comparative genomics and metagenomics

Comparative analysis of (meta)genomes necessitates aggregation, integration, and synthesis of well-annotated data using standards. The Genomic Standards Consortium (GSC) collaborates with the research community to develop and maintain the Minimal Information about any (x) Sequence (MIxS) reporting standard for genomic data. To facilitate use of the GSC’s MIxS reporting standard, we provide a description of the structure and terminology, how to navigate ontologies for required terms in MIxS, and demonstrate practical usage through a soil metagenome example.

standards, metadata, genome, metagenome, schema, v↗

Path Forward: Materials Data Modernization for ASME Codes and Standards in the Artificial Intelligence Era

Development of the ASME Materials Properties Database was initiated in the early 2010s to support the ASME Codes and Standards. As information technologies advance at an accelerated pace with the artificial intelligence era on the horizon, the ASME Materials Properties Database must be further modernized from a database to a knowledgebase to ride the wave of digital information revolution and effectively support the ASME Codes and Standards in the new era.This paper is intended to provide an overview of the ASME Materials Properties Database and discuss a roadmap for its future development to facilitate understanding of and participation from different sectors of the Codes and Standards community. It first reviews the basic concepts of data, information, knowledge, database, and database system; as well as the pros and cons in different types of data management, and then discusses the path forward for a desired evolution of the database into a self-explanatory and machine-readable knowledgebase that is consistent with human cognitive processes for the Codes and Standards development and furthermore provides resources for data processing and analysis to reach an eventual goal of streamlining the Codes and Standards development from the initial inquiry, throughout data submission, analysis, …, to Codes and Standards rule establishment for final publication.

Ren, Weiju↗