Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “community data standard”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Hosting downscaled decision-relevant community data products in ESGF2-US

As regionally-relevant high-resolution Earth system data is increasingly relied upon across scientific, policy, and practitioner communities, there is an urgent need for coordinated and federated infrastructure to store, manage, standardize, and distribute decision-relevant community data products. Substantial effort is required to ensure that these products, which are often critical for regional impact assessments and decision-making, are findable, accessible, interoperable, and reusable. The Earth System Grid Federation US project (ESGF2-US) is addressing this challenge by expanding its open-source, distributed platform to support the hosting and dissemination of downscaled Earth system datasets. This expansion includes aligning new downscaled datasets with developing community standards for metadata and file structure, consistent with existing ESGF archives. This includes ensuring CF-compliance, applying CMORization where appropriate, and developing tools to streamline user access. In this paper, we highlight the technical and coordination work required to bring downscaled data into ESGF2-US and aim to inform the broader Earth system data user community about the growing availability and utility of these curated resources.

ESGF↗

Evolution of standardization and dissemination of cryo-EM structures and data jointly by the community, PDB, and EMDB

Cryogenic electron microscopy (cryo-EM) methods began to be used in the mid-1970s to study thin and periodic arrays of proteins. Following a half-century of development in cryo-specimen preparation, instrumentation, data collection, data processing and modeling software, cryo-EM has become a routine method for solving structures from large biological assemblies to small biomolecules at near to true atomic resolution. This review explores the critical roles played by the Protein Data Bank (PDB) and Electron Microscopy Data Bank (EMDB) in partnership with the community to develop the necessary infrastructure to archive cryo-EM maps and associated models. Public access to cryo-EM structure data has in turn facilitated better understanding of structure-function relationships and advancement of image processing and modeling tool development. The partnership between the global cryo-EM community and PDB and EMDB leadership has synergistically shaped the standards for metadata, one-stop deposition of maps and models, and validation metrics to assess the quality of cryo-EM structures. The advent of cryo-electron tomography (cryo-ET) for in situ molecular cell structures at a broad resolution range and their correlations with other imaging data introduces new data archival challenges in terms of data size and complexity in the years to come.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Perspectives for artificial intelligence in bioprocess automation

Recent advances in artificial intelligence (AI) have rapidly changed the lab automation landscape, promoting self-driving laboratories (SDLs) that enable autonomous scientific discovery. These trends are increasingly applied in bioprocess development, yet bioprocessing faces unique challenges - biological complexity, regulatory and safety requirements, and multiscale experimentation - that distinguish it from other automation domains. Rather than pursuing full autonomy, we foresee that hybrid SDLs, combining AI-driven decision-making with sustained human oversight, represent the most practical near-term trajectory. This review examines three interconnected perspectives: (i) hybrid human-machine decision-making for bioprocessing; (ii) laboratory design considerations in the era of AI; and (iii) scale-up challenges when transitioning from screening to manufacturing. We highlight critical gaps in data standardization and the required community efforts necessary to realize autonomous bioprocess innovation.

Helleckes, Laura Marie↗

Principles of the Battery Data Genome

Batteries are central to modern society. They are no longer just a convenience but a critical enabler of the transition to a resilient, low-carbon economy. Battery development capabilities are provided by communities spanning materials discovery, battery chemistry and electrochemistry, cell and pack design, scale-up, manufacturing, and deployments. Despite their relative maturity, data-science practices among these diverse groups are far behind the state of the art in other fields, which have demonstrated an ability to significantly improve innovation and economic impact. The negative consequences of the present paradigm include incremental improvements but few breakthroughs, significant manufacturing uncertainties, and cascading investment risks that collectively slow deployments. The primary roadblock to a battery-data-science renaissance is the requirement for large amounts of high-quality data, which are not available in the current fragmented ecosystem. Here, in this study, we identify gaps and propose principles that enable the solution by building a robust community of data hubs with standardized practices and flexible sharing options that will seed advanced tools spanning innovation to deployment. Precedents are offered that demonstrate that both public good and immense economic gains will arise from sharing valuable battery data. The proposed Battery Data Genome looks to broadly transform innovations and revolutionize their translation from research to societal impact.

25 ENERGY STORAGE↗

Community standards and future opportunities for synthetic communities in plant–microbiota research

Harnessing beneficial microorganisms is seen as a promising approach to enhance sustainable agriculture production. Synthetic communities (SynComs) are increasingly being used to study relevant microbial activities and interactions with the plant host. Yet, the lack of community standards limits the efficiency and progress in this important area of research. Here, to address this gap, we recommend three actions: (1) defining reference SynComs; (2) establishing community standards, protocols and benchmark data for constructing and using SynComs; and (3) creating an infrastructure for sharing strains and data. We also outline opportunities to develop SynCom research through technical advances, linking to field studies, and filling taxonomic blind spots to move towards fully representative SynComs.

59 BASIC BIOLOGICAL SCIENCES↗

A customizable data management framework for high-repetition-rate high-energy-density science

The high-energy-density (HED) physics community is moving toward a new paradigm of high-repetition-rate (HRR) operation. To fully leverage the scientific power of HRR HED facilities, all of the components of each subsystem (laser, targetry, and performance diagnostics) must be connected and synchronized in a reliable and robust manner while the data acquired are tagged and archived in real time. To this end, GA has begun developing a generalized NoSQL-database framework, the MongoDB repository for information and archiving. An organizational strategy has been developed that shifts HED data organization from a shot-based to a diagnostic-based approach in order to increase archival and retrieval efficiency that lends itself to optimization applications. This work is a first step in pushing HRR HED science toward data management solutions that emphasize machine actionability and aim to stimulate community engagement to define data standards in HED science.

Instruments & Instrumentation↗

TRMM .25 deg x .25 deg Gridded Precipitation Text Product

Since the launch of the Tropical Rainfall Measuring Mission (TRMM), the Precipitation Measurement Missions science team has endeavored to provide TRMM precipitation retrievals in a variety of formats that are more easily usable by the broad science community than the standard Hierarchical Data Format (HDF) in which TRMM data is produced and archived. At the request of users, the Precipitation Processing System (PPS) has developed a .25 x .25 gridded product in an easily used ASCII text format. The entire TRMM mission data has been made available in this format. The paper provides the details of this new precipitation product that is designated with the TRMM designator 3G68.25. The format is packaged into daily files. It provides hourly precipitation information from the TRMM microwave imager (TMI), precipitation radar (PR), and TMI/PR combined rain retrievals. A major advantage of this approach is the inclusion only of rain data, compression when a particular grid has no rain from the PR or combined, and its direct ASCII text format. For those interested only in rain retrievals and whether rain is convection or stratiform, these products provide a huge reduction in the data volume inherent in the standard TRMM products. This paper provides examples of the 3G68 data products and their uses. It also provides information about C tools that can be used to aggregate daily files into larger time samples. In addition, it describes the possibilities inherent in the spatial sampling which allows resampling into coarser spatial sampling. The paper concludes with information about downloading the gridded text data products.

Stocker, Erich↗

Improving the Quality of Geothermal Data Through Data Standards and Pipelines Within the Geothermal Data Repository: Preprint

For machine learning outputs to be applicable to real world problems, high quality data are needed to ensure high quality results. With the more recent emphasis on machine learning in geothermal, there is an increasing need for greater focus on the quality of the data available for use in these projects. For example, Geothermal Operational Optimization Using Machine Learning (GOOML) utilized large quantities of geothermal power plant operational data to inform power plant operational configurations to maximize power generation. High quality datasets result from dependable sensors or devices collecting data, high frequency of measurements, sufficient data points, adequate metadata, reliable storage of data, and sufficient data curation. Another component that contributes to high quality data is reusability, which can be enhanced through data standardization. Data Standardization creates consistency in formatting and contents of like datasets, lessening preprocessing requirements and ensuring adequate information provided by a given dataset. The Geothermal Data Repository (GDR) aims to help improve data quality through automated data standardization for high-value datasets through the implementation of data pipelines alongside reliable and accessible long-term storage for datasets. As such, the GDR has decided to shift away from recommending the use of Excel-based content models and towards the implementation of automated data pipelines. This takes the burden of data standardization off the user and project team and will increase the availability of standardized geothermal data available through the GDR. A set of recommendations, or a data standard for each data type will exist with each data pipeline in order to advise data collection for maximum usability for future research. This paper serves to describe the GDR's proposed transition towards data standardization through automated data pipelines, to discuss the need for and value of such a shift, and to call for suggestions from the community regarding the most useful data standards and pipelines.

data↗

HPC ODA Commons [SWR-26-003]

HPC ODA Commons is a community-driven platform for standardizing HPC operational data analytics. HPC sites generate enormous volumes of operational data - scheduler logs, accounting records, monitoring streams - but turning that data into actionable insight is needlessly hard. Each site builds bespoke parsers, schemas, and evaluation pipelines. Results can't be compared across institutions. Promising analytics ideas stay siloed because there's no shared language for describing the data, the experiments, or the outcomes. HPC ODA Commons fixes this by establishing community-governed contracts - versioned schemas, canonical artifacts, and benchmark recipes - that make ODA workflows discoverable, reproducible, and comparable. It pairs these standards with a practical, CLI-first toolkit that lets operators and researchers go from raw logs to standardized results without sending data off-cluster.

Menear, Kevin [National Laboratory of the Rockies ↗

Templates for developing and versioning data standards and reporting formats using GitHub

This data package contains three templates that can be used for creating README files and Issue Templates, written in the markdown language, that support community-led data reporting formats. We created these templates based on the results of a systematic review (see related references) that explored how groups developing data standard documentation use the Version Control platform GitHub, to collaborate on supporting documents. Based on our review of 32 GitHub repositories, we make recommendations for the content of README Files (e.g., provide a user license, indicate how users can contribute) and so 'README_template.md' includes headings for each section. The two issue templates we include ('issue_template_for_all_other_changes.md' and 'issue_template_for_documentation_change.md') can be used in a GitHub repository to help structure user-submitted issues, or can be modified to suit the needs of data standard developers. We used these templates when establishing ESS-DIVE's community space on GitHub (https://github.com/ess-dive-community) that includes documentation for community-led data reporting formats. We also include file-level metadata 'flmd.csv' that describes the contents of each file within this data package. Lastly, the temporal range that we indicate in our metadata is the time range during which we searched for data standards documented on GitHub.

54 ENVIRONMENTAL SCIENCES↗

Technologies and Methods Used at the Laboratory for Atmospheric and Space Physics (LASP) to Serve Solar Irradiance Data

The Laboratory for Atmospheric and Space Physics (LASP) at the University of Colorado in Boulder, USA operates the Solar Radiation and Climate Experiment (SORCE) NASA mission, as well as several other NASA spacecraft and instruments. Dozens of Solar Irradiance data sets are produced, managed, and disseminated to the science community. Data are made freely available to the scientific immediately after they are produced using a variety of data access interfaces, including the LASP Interactive Solar Irradiance Datacenter (LISIRD), which provides centralized access to a variety of solar irradiance data sets using both interactive and scriptable/programmatic methods. This poster highlights the key technological elements used for the NASA SORCE mission ground system to produce, manage, and disseminate data to the scientific community and facilitate long-term data stewardship. The poster presentation will convey designs, technological elements, practices and procedures, and software management processes used for SORCE and their relationship to data quality and data management standards, interoperability, NASA data policy, and community expectations.

Pankratz, Chris↗

The need for standardization and improved open (meta)data practices in metaproteomics

Metaproteomics enables functional insight into microbial communities by identifying and quantifying proteins in complex samples. Yet, heterogeneous analytical workflows and the lack of standardization across experimental and bioinformatics stages hinder reproducibility and comparability, limiting integration with other omics data. We here present a community-developed reporting checklist tailored to the specific needs of metaproteomics. We also outline current efforts to enable structured and interoperable metadata capture, drawing on standards from proteomics and microbiome research wherever possible. By promoting transparent reporting and advancing metadata practices, our recommendations aim to align metaproteomics more closely with FAIR principles and support reproducible and interoperable research practices.

Armengaud, Jean [Universite Paris-Saclay, France]↗

Application of ESE Data and Tools to Air Quality Management: Services for Helping the Air Quality Community use ESE Data (SHAirED)

The goal of this REASoN applications and technology project is to deliver and use Earth Science Enterprise (ESE) data and tools in support of air quality management. Its scope falls within the domain of air quality management and aims to develop a federated air quality information sharing network that includes data from NASA, EPA, US States and others. Project goals were achieved through a access of satellite and ground observation data, web services information technology, interoperability standards, and air quality community collaboration. In contributing to a network of NASA ESE data in support of particulate air quality management, the project will develop access to distributed data, build Web infrastructure, and create tools for data processing and analysis. The key technologies used in the project include emerging web services for developing self describing and modular data access and processing tools, and service oriented architecture for chaining web services together to assemble customized air quality management applications. The technology and tools required for this project were developed within DataFed.net, a shared infrastructure that supports collaborative atmospheric data sharing and processing web services. Much of the collaboration was facilitated through community interactions through the Federation of Earth Science Information Partners (ESIP) Air Quality Workgroup. The main activities during the project that successfully advanced DataFed, enabled air quality applications and established community-oriented infrastructures were: develop access to distributed data (surface and satellite), build Web infrastructure to support data access, processing and analysis create tools for data processing and analysis foster air quality community collaboration and interoperability.

Falke, Stefan↗

Achieving Global Consensus on Acceptable Sound Levels for Overland Supersonic Flight

The National Aeronautics and Space Administration has made a commitment to deliver to the International Civil Aviation Organization’s Committee on Aviation Environmental Protection (ICAO CAEP) data defining community response to sounds from supersonic aircraft designed such that their sonic boom is replaced with a soft “thump” sound. The dataset will be a correlation of public perceptions of these sounds to the corresponding acoustic levels. The data will support efforts to develop international standards for permissible noise from supersonic overflight. NASA is planning and preparing for a series of community overflight tests with the X-59, a unique research aircraft capable of generating the “sonic thump”. NASA will begin these tests in 2024. With an eye toward achieving global consensus for noise standards, NASA’s goal is that the community response data be as broadly representative of the response of the international population as possible. As such, NASA is engaging the international regulatory and research communities in both the planning and execution of these tests, through status briefings at ICAO CAEP-sponsored meetings and through workshops with international participation. As part of this outreach NASA held a virtual workshop in December 2021 focused on strategies and considerations for estimating noise exposure levels and conducting surveys to characterize community annoyance levels relative to the “thump” sounds. This paper will present an overview of NASA’s effort, with a focus on the plans and technical goals for the community response tests. In addition, results of the recent workshop will be briefed, including considerations and approaches for ensuring broad representativeness of results and approaches for estimation of the sound levels across the test community. Participant feedback from both the workshop and previous engagements will be discussed, along with how it is being addressed in NASA’s ongoing planning efforts.

Supersonic Flight↗

Review of Experimental Data for Validating Computer Codes Used in Shielding Calculations for Spent Fuel Storage and Transportation Systems

This report presents a review of available radiochemical assay data and shielding benchmarks applicable to spent nuclear fuel (SNF) shielding calculations. The relevant information reviewed herein includes the Spent Fuel Composition (SFCOMPO) database, the Shielding Integral Benchmark Archive and Database (SINBAD), the International Handbook of Evaluated Criticality Safety Benchmark Experiments, and published measurements of external dose rates of casks loaded with SNF. The relevant experimental data identified in this report may be used to support verification and validation of computer codes used in SNF cask/transport shielding applications, as well as development of calculation uncertainties. It should be noted that a relatively small subset of the identified experimental data (e.g., criticality alarm experiments) is available in a standard format established by the international community participating in experimental isotopic and shielding data evaluations. An effort of the SFCOMPO Technical Review Group (TRG) is underway to publish first isotopic evaluations of individual assay data using a standard data evaluation format. The SINBAD TRG has recently initiated benchmark evaluations and modernization of the database. Therefore, more relevant information is expected in the future that will enable users to select quality experimental data in depletion code and shielding code validations for SNF applications.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Collaborative Data Publication Utilizing the Open Data Repository's (ODR) Data Publisher

Introduction: For small communities in diverse fields such as astrobiology, publishing and sharing data can be a difficult challenge. While large, homogenous fields often have repositories and existing data standards, small groups of independent researchers have few options for publishing standards and data that can be utilized within their community. In conjunction with teams at NASA Ames and the University of Arizona, the Open Data Repository's (ODR) Data Publisher has been conducting ongoing pilots to assess the needs of diverse research groups and to develop software to allow them to publish and share their data collaboratively. Objectives: The ODR's Data Publisher aims to provide an easy-to-use and implement software tool that will allow researchers to create and publish database templates and related data. The end product will facilitate both human-readable interfaces (web-based with embedded images, files, and charts) and machine-readable interfaces utilizing semantic standards. Characteristics: The Data Publisher software runs on the standard LAMP (Linux, Apache, MySQL, PHP) stack to provide the widest server base available. The software is based on Symfony (www.symfony.com) which provides a robust framework for creating extensible, object-oriented software in PHP. The software interface consists of a template designer where individual or master database templates can be created. A master database template can be shared by many researchers to provide a common metadata standard that will set a compatibility standard for all derivative databases. Individual researchers can then extend their instance of the template with custom fields, file storage, or visualizations that may be unique to their studies. This allows groups to create compatible databases for data discovery and sharing purposes while still providing the flexibility needed to meet the needs of scientists in rapidly evolving areas of research. Research: As part of this effort, a number of ongoing pilot and test projects are currently in progress. The Astrobiology Habitable Environments Database Working Group is developing a shared database standard using the ODR's Data Publisher and has a number of example databases where astrobiology data are shared. Soon these databases will be integrated via the template-based standard. Work with this group helps determine what data researchers in these diverse fields need to share and archive. Additionally, this pilot helps determine what standards are viable for sharing these types of data from internally developed standards to existing open standards such as the Dublin Core (http://dublincore.org) and Darwin Core (http://rs.twdg.org) metadata standards. Further studies are ongoing with the University of Arizona Department of Geosciences where a number of mineralogy databases are being constructed within the ODR Data Publisher system. Conclusions: Through the ongoing pilots and discussions with individual researchers and small research teams, a definition of the tools desired by these groups is coming into focus. As the software development moves forward, the goal is to meet the publication and collaboration needs of these scientists in an unobtrusive and functional way.

easy to use and implement software tool↗

Examining Mars with SPICE

The International Mars Conference highlights the wealth of scientific data now and soon to be acquired from an international armada of Mars-bound robotic spacecraft. Underlying the planning and interpretation of these scientific observations around and upon Mars are ancillary data and associated software needed to deal with trajectories or locations, instrument pointing, timing and Mars cartographic models. The NASA planetary community has adopted the SPICE system of ancillary data standards and allied tools to fill the need for consistent, reliable access to these basic data and a near limitless range of derived parameters. After substantial rapid growth in its formative years, the SPICE system continues to evolve today to meet new needs and improve ease of use. Adaptations to handle landers and rovers were prototyped on the Mars pathfinder mission and will next be used on Mars '01-'05. Incorporation of new methods to readily handle non-inertial reference frames has vastly extended the capability and simplified many computations. A translation of the SPICE Toolkit software suite to the C language has just been announced. To further support cartographic calculations associated with Mars exploration the SPICE developers at JPL have recently been asked by NASA to work with cartographers to develop standards and allied software for storing and accessing control net and shape model data sets; these will be highly integrated with existing SPICE components. NASA specifically supports the widest possible utilization of SPICE capabilities throughout the international space science community. With NASA backing the Russian Space Agency and Russian Academy of Science adopted the SPICE standards for the Mars 96 mission. The SPICE ephemeris component will shortly become the international standard for agencies using the Deep Space Network. U.S. and European scientists hope that ESA will employ SPICE standards on the Mars Express mission. SPICE is an open set of standards, and all related specifications and software are freely distributed around the world. This poster describes the current state of SPICE system development, with special emphasis on current and planned support for Mars exploration missions.

Acton, Charles H.↗