Engineering PapersSearch

SEARCH · Engineering Papers

Results for “data stewardship”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Aligning NASA Earth Science Data Stewardship with FAIR Principles: Outcomes, Recommendations, and Future Directions

The FAIR Principles—Findable, Accessible, Interoperable, and Reusable—offer a widely accepted framework for improving the sharing and reuse of digital scientific data by both human and machine users. Following these principles is critical for effective scientific data stewardship, broader scientific collaboration, and compliance with federal and agency data policies. This paper, based on the work of NASA’s Open, Free, and FAIR Working Group (O’FAIR WG) under the Earth Science Data Systems Program, presents an overview of how FAIR is being applied within NASA’s Earth science data landscape. It highlights ongoing progress and challenges, identifies FAIR-enabling resources, and offers recommendations and strategic actions to enhance the FAIRness of NASA-funded open and free Earth science data products. The FAIR-enabling resources identified underscore the vital role of NASA's existing enterprise processes, standards, tools, and infrastructures in supporting FAIR implementation. Our findings show strong performance in making NASA Earth science data more findable and accessible. However, further work is needed—especially in enhancing interoperability, so that different systems and tools can better understand and exchange data. This is especially important for enabling machine-driven discovery and analysis. We emphasize the importance of a balanced strategy that combines a centralized, top-down approach—focused on building enterprise-level capabilities and processes—with a decentralized, bottom-up approach driven by discipline-specific needs and community practices. We advocate for coordinated efforts to enhance (meta)data interoperability to facilitate seamless data and information sharing and exchange of Earth science data both within NASA and across other agencies managing Earth science data.

Data Product

FAIR Data Meets FAIR Software

Modern scientific research is increasingly defined by the interplay between data, software, and the workflows that connect them. Yet while the FAIR (Findable, Accessible, Interoperable, Reusable) principles have become foundational for scientific data stewardship, the same level of structure and expectation has only recently begun to extend to research software. This talk covers why and how FAIR principles are being applied to data and software to support data reuse. It outlines the gaps in current sharing norms, the growing federal emphasis on persistent identifiers and public access, and the opportunities created when datasets, computational workflows, code, and models are linked through rich, standardized metadata. Practical implementation pathways for the EIC and JLab communities are described, including datacards for structured dataset documentation and provenance-aware workflows. By aligning data lifecycle management with FAIR-aligned software practices, the scientific community can advance toward autonomous knowledge graphs, generative workflows, and high-quality, AI-ready scientific datasets.

McSpadden, Diana [Thomas Jefferson National Accele

PDB-IHM: A System for Deposition, Curation, Validation, and Dissemination of Integrative Structures

Structures of many large biomolecular assemblies are now being determined using integrative approaches. In these approaches, information derived from multiple experimental and computational methods is combined to compute three-dimensional structures of multi-protein complexes and other macromolecular machines. A standalone prototype data resource for integrative structures called PDB-Dev was built, based on recommendations of the Integrative and Hybrid Methods (IHM) Task Force of the Worldwide Protein Data Bank (wwPDB). This effort included developing data standards and software tools for collecting, curating, validating, visualizing, archiving, and disseminating integrative structures that span diverse spatiotemporal scales and conformational states. Mechanisms have been created to validate integrative structures based on the experimental data underpinning them. Building upon this foundational framework, PDB-Dev has been further expanded to handle large dynamic macromolecular systems and integrative structures that combine, for example, experimental restraints with atomic coordinates computed by machine learning algorithms. Data standards and supporting tools have also been extended to capture information about biomolecular dynamics, such as conformational transitions and related kinetic data derived from biophysical methods. Recently, PDB-Dev was unified with the PDB archive and rebranded as PDB-IHM (pdb-ihm.org), further promoting FAIR (Findable, Accessible, Interoperable, and Reusable) principles of data stewardship for integrative structural biology.

IHMCIF

ESnet Data and AI Workshop Report

In February 2025, the DOE user facility Energy Sciences Network (ESnet) held a three-day Data and AI Workshop in Berkeley, California. The objective of the workshop was to identify challenges within ESnet that could be addressed through data-driven methods, to help define ESnet’s data-analysis requirements, and to shape its AI strategy, guiding data-stewardship efforts and the direction of AI research and AIOps exploration for ESnet7, the next iteration of ESnet’s network. This report summarizes the multi-faceted discussions and findings and presents a set of recommendations for next steps.

97 MATHEMATICS AND COMPUTING

Microbiome data management in action workshop: Atlanta, GA, USA, June 12–13, 2024

Microbiome research is revolutionizing human and environmental health, but the value and reuse of microbiome data are significantly hampered by the limited development and adoption of data standards. While several ongoing efforts are aimed at improving microbiome data management, significant gaps still remain in terms of defining and promoting adoption of consensus standards for these datasets. The Strengthening the Organization and Reporting of Microbiome Studies (STORMS) guidelines for human microbiome research have been endorsed and successfully utilized by many research organizations, publishers, and funding agencies, and have been recognized as a consensus community standard. No equivalent effort has occurred for environmental, synthetic, and non-human host-associated microbiomes. To address this growing need within the microbiome research community, we convened the Microbiome Data Management in Action Workshop (June 12–13, 2024, in Atlanta, GA, USA), to bring together key decision makers in microbiome science including researchers, publishers, funders, and data repositories. The 50 attendees, representing the diverse and interdisciplinary nature of microbiome research, discussed recent progress and challenges, and brainstormed actionable recommendations and paths forward for coordinated environmental microbiome data management and the modifications necessary for the STORMS guidelines to be applied to environmental, non-human host, and synthetic microbiomes. The outcomes of this workshop will form the basis of a formalized data management roadmap to be implemented across the field. These best practices will drive scientific innovation now and in years to come as these data continue to be used not only in targeted reanalyses but in large-scale models and machine learning efforts.

54 ENVIRONMENTAL SCIENCES

Hanford Site Mule Deer Monitoring Report for Fiscal Years 2024 and 2026

The U.S. Department of Energy, Hanford Field Office (HFO) conducts ecological monitoring on the Hanford Site to collect and track data needed to ensure compliance with environmental laws, regulations, and policies governing Department of Energy activities. The vision for the HFOmanaged portion of the Hanford Site, hereby referred to as Central Hanford, focuses not only on the cleanup of nuclear facilities and waste sites but on the protection and restoration of the Hanford Site lands. As the HFO moves toward accomplishing this vision, understanding of the ecological resources present and the need for conservation and/or protection of those resources will be critical for making informed decisions for responsible site stewardship. Ecological monitoring data provides baseline information about the plants, animals, and habitats under HFO stewardship at Central Hanford required for decision-making under the National Environmental Policy Act of 1969 (NEPA) and Comprehensive Environmental Response, Compensation, and Liability Act of 1980.

54 ENVIRONMENTAL SCIENCES

Genesis Data Card Schema, Template and Supporting Tools

Genesis Data Cards provide a standardized template and schema for documenting scientific datasets in support of discovery, access, interoperability, reusability, governed use, and AI usability. This release of the Genesis Data Card repository includes a versioned Markdown template, a LinkML schema with generated Pydantic and JSON artifacts, schema documentation, and example completed data cards. Validation tooling is provided to ensure that completed data cards conform to the schema prior to submission. Accompanying documentation for the structured metadata is provided as a Field Reference Guide. The schema and accompanying template provided in this repository address the call for actionable context that enables humans and AI systems to find, access, interpret, cite, and reuse data, and, when appropriate, integrate it into AI and machine learning workflows. The data card is intended to serve as a common metadata artifact intended to support standardized, cross-program dataset documentation across Department of Energy (DOE)-aligned efforts, including but not limited to Genesis Mission-related implementations, the Office of Science, National Nuclear Security Administration (NNSA), and Advanced Simulation and Computing (ASC) data governance and stewardship initiatives.

data card

DOE Repository Metadata Profile (DRMP): A Metadata Framework for Advancing Interoperability and AI Readiness Across Scientific Repositories

The Department of Energy (DOE) funds a diverse and distributed ecosystem of repositories that steward scientific data, publications, and software across its research programs, user facilities, and national laboratories. While significant progress has been made in standardizing dataset-level metadata, the metadata describing repositories themselves (their identity, governance, access interfaces, policies, and technical capabilities) remains inconsistent and fragmented across DOE-funded systems. This variability limits discoverability, interoperability, automated validation, and AI-driven analysis, all of which are increasingly essential for modern scientific workflows. To address this gap, the DOE Data Curation Working Group (DCWG) developed the DOE Repository Metadata Profile (DRMP). The DRMP is a practical, community-driven framework that defines how repositories can describe themselves in a consistent, machine-actionable, and scalable manner. The DRMP is not a new metadata schema. Instead, it is a mapping profile and structured element set capturing the essential characteristics of DOE repositories. It harmonizes repository-level metadata across six widely adopted community schemas: RE3Data; DCAT-US v3; Schema.org; Dublin Core; DataCite 4.6; and PREMIS 3.0. This harmonization eliminates reinvention and enables interoperability within DOE and across the broader scientific ecosystem. A core objective of the DRMP is to reduce burden on repositories by allowing them to reuse their existing metadata through a Rosetta-style crosswalk rather than redesigning local implementations. The profile introduces a three-level conformance model that supports incremental adoption: • Level 1 – Minimum Viable Record (MVR): foundational identification elements required for workflows, project registration, and basic repository presence. • Level 2 – Interoperable: structured metadata enabling alignment with national and international discovery systems. • Level 3 – AI-Ready: enhanced provenance, policy transparency, fixity, semantic context, and capabilities that support automated reasoning, model training governance, and machine-assisted curation. To support implementation, the DRMP includes JSON Schema definitions, OpenAPI patterns, and MCP templates that allow repositories to publish machine-readable metadata directly within existing platforms. These resources are modular and lightweight, enabling adoption without major architectural change. Adopting the DRMP enables repositories to: • Enhance discoverability and interoperability by aligning identifiers, classifications, and descriptive elements across widely used schema standards. • Support federated discovery and cross-registration across DOE systems, Data.gov, and international catalogs. • Enable AI agents and workflow orchestration systems to interpret repository-level metadata within the American Science Cloud (AmSC) through Model Context Protocol (MCP)-based context publication. • Demonstrate alignment with DOE’s open science, stewardship, and FAIR data priorities. This guidance represents a community-driven step forward. Through voluntary adoption and continued feedback, the DRMP advances a cohesive, machine-actionable description of DOE repositories that supports FAIR data practices, preparing the infrastructure for AI-enabled research, and strengthening the discoverability and reuse of DOE’s scientific outputs.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION

Materials Data Science Ontology(MDS-Onto): Unifying Domain Knowledge in Materials and Applied Data Science

Ontologies have gained popularity in the scientific community as a way to standardize terminologies in organizations’ data. Although certain cohorts have created frameworks with rules and guidelines on creating ontologies, there exist significant variations in how Materials Science ontologies are currently developed. We seek to provide guidance in the form of a unified automated framework for developing interoperable and modular ontologies for Materials Data Science that simplifies the ontology terms matching by establishing a semantic bridge up to the Basic Formal Ontology(BFO). This framework provides key recommendations on how ontologies should be positioned within the semantic web, what knowledge representation language is recommended, and where ontologies should be published online to boost their findability and interoperability. Two fundamental components of the MDS-Onto framework are the bilingual package called FAIRmaterials for ontology creation and FAIRLinked, for FAIR data creation. To showcase the practical capabilities of FAIRmaterials, we present two exemplar domain ontologies of MDS-Onto: Synchrotron X-Ray Diffraction and Photovoltaics.

29 ENERGY PLANNING, POLICY, AND ECONOMY

FAIRmaterials: Ontology Tools with Data FAIRification in Development

The bilingual FAIRmaterials package simplifies the creation and visualization of materials and data science ontologies. FAIRmaterials, available in the Python and R languages, addresses the complexities associated with traditional ontology editors based on manual user input such as Protege with an intuitive workflow and easy-to-use templates, making it accessible to users both experienced and inexperienced with ontologies. The FAIRmaterials package is its ability to programatically convert simple and structured CSV inputs into rich, well-defined ontologies. This capability is designed to support the findability, accessibility, interoperability, and reusability (FAIR) of research data and serve as a tool in the process of data FAIRification. Its additional features, such as automated ontology merging, static visualizations, and comprehensive documentation for outputs extend its utility, making it a valuable tool for any researcher engaged in knowledge management.

Bradley, Alexander Harding [Case Western Reserve U

FAIRLinked: Data FAIRification Tools for Materials Data Science

FAIRLinked is a software package created to support the FAIRification of materials science data, ensuring proper alignment with FAIR principles: Findable, Accessible, Interoperable, and Reusable. It is built to be compatible with MDS-Onto, an ontology designed to capture the semantics of various types of materials data, enabling integration and sharing across different research workflows. The package is subdivided into three subpackages: InterfaceMDS, RDFTableConversion, and QBWorkflow. The first subpackage, InterfaceMDS allows users to search for terms using either string search or various filters, explore different domains and subdomains, and add terms to MDS-Onto. RDFTableConversion is used for serialization and deserialization of data from CSV into JSONLDs and vice versa in a way that captures the semantics of the data using MDS-Onto. Lastly, QBWorkflow is a serialization and deserialization workflow that incorporates RDF Data Cube vocabulary, useful for working with multidimensional datasets. By offering these packages, FAIRLinked lowers the barrier of creating FAIR, machine-actionable data for researchers in the materials science community.

FAIR

AI-Ready Data Pilot Project Report

The proliferation of artificial intelligence in scientific research has created an urgent need to define "AI-ready data" for researchers and, more importantly, provide resources to help them produce AI-ready data. At Pacific Northwest National Laboratory, we conducted a pilot study with three data scientists evaluating three CSV datasets from different scientific domains, followed by semi-structured interviews capturing assessment practices. Our findings reveal that AI-readiness evaluation is intuition-based, with practitioners asking "How fast can I go from raw data to my machine learning pipeline?" Data scientists consistently prioritized workflow efficiency, human interpretability, and quality stewardship signals. From these insights, we developed a practical evaluation framework comprising data requirements, metadata standards, and validation tests that provides actionable criteria for producing and curating AI-ready datasets, addressing the gap between theoretical understanding and practical implementation.

97 MATHEMATICS AND COMPUTING

C 12 ( n , n 1 ′ γ ) partial γ -ray cross section measured using the GENESIS array

Improved neutron inelastic scattering cross sections have repeatedly been identified as a top priority nuclear data need, important for basic science and a range of applications in nuclear energy, stockpile stewardship, and proliferation detection. For the C 12 ( n , n ′ γ ) reaction in particular, recent measurements have unveiled some structural discrepancies, demonstrating incongruities among themselves and in relation to the ENDF/B-VIII.0 nuclear data evaluation. To help resolve these disagreements, a measurement was performed at the 88-Inch Cyclotron at Lawrence Berkeley National Laboratory using a broad-spectrum neutron beam and a 99.8% pure natural carbon target. The Gamma Energy Neutron Energy Spectrometer for Inelastic Scattering (GENESIS) was employed to measure energy-differential γ -ray emission spectra as a function of incident neutron energy in the energy range of 5.5 to 16.7 MeV. The C 12 partial γ -ray cross sections were extracted at 63 ∘ , 122 . 5 ∘ , and 150 ∘ with respect to the incoming neutron beam and integrated using angular distribution data available in the literature. The data show agreement with a recent literature measurement and evaluation from 11 to 15 MeV, but indicate a larger cross section for incident neutron energies between 5.5 and 8.5 MeV. The measured relative angular distributions are also reported and were found to agree with evaluation. Published by the American Physical Society 2025

Gordon, J. M. (ORCID:0009000789886897)

Commissioning of the neutron irradiation station at the University of Notre Dame

Cross section data for neutron-induced reactions are needed for applications in nuclear astrophysics, stockpile stewardship, and nuclear reactor design but often contain discrepancies and gaps. A neutron source with a high neutron flux, a well-characterized neutron energy distribution, and a wide range of available neutron energies is ideal for studying such reactions in order to improve on existing data and provide new experimental results. For this purpose, the Neutron Irradiation Station (NIS) was developed using the 7 Li(p,n) 7 Be reaction. Neutron energy profiles were obtained using 3 He spectrometers and neutron standards were used to determine the neutron flux using the activation method. The neutron energy profile was found to be quasi-monoenergetic and neutrons were produced at energies from 1 keV to 1 MeV. Finally, the flux at the production target was found to be (1.42 ± 0.20) x 10 8 s −1 cm −2 .

73 NUCLEAR PHYSICS AND RADIATION PHYSICS

Hanford Site Roadside Bird Surveys Report for Calendar Years 2021-2023

The U.S. Department of Energy, Hanford Field Office (HFO) conducts ecological monitoring on the Hanford Site to collect and track data needed to ensure compliance with an array of environmental laws, regulations, and policies governing HFO activities. Ecological monitoring data provide baseline information about the plants, animals, and habitats under HFO stewardship at Hanford which is required for decision-making under the National Environmental Policy Act (NEPA) and Comprehensive Environmental Response, Compensation, and Liability Act (CERCLA). The Hanford Site Comprehensive Land Use Plan (CLUP, DOE/EIS-0222-F), which is the Environmental Impact Statement for Hanford Site activities, helps ensure that HFO, its contractors, and other entities conducting activities on the Hanford Site are in compliance with NEPA.

54 ENVIRONMENTAL SCIENCES

Hanford Reach Fall Chinook Salmon Redd Monitoring Report for Calendar Year 2024

The U.S. Department of Energy, Hanford Field Office (HFO) conducts ecological monitoring on the Hanford Site to collect and track data needed to ensure compliance with an array of environmental laws, regulations, and policies governing HFO activities. Ecological monitoring data provide baseline information about the plants, animals, and habitats under HFO stewardship at the Hanford Site required for decision making under the National Environmental Policy Act (NEPA) and Comprehensive Environmental Response, Compensation, and Liability Act. DOE/EIS-0222, Final Hanford Comprehensive Land-Use Plan Environmental Impact Statement, (CLUP) evaluates the potential environmental impacts associated with implementing a comprehensive land-use plan for the Hanford Site for at least the next 50 years, and ensures that HFO, its contractors, and other entities conduct activities on the Hanford Site in compliance with NEPA.

54 ENVIRONMENTAL SCIENCES

Hanford Reach Fall Chinook Salmon Redd Monitoring Report for Calendar Year 2023

The U.S. Department of Energy, Hanford Field Office (HFO) conducts ecological monitoring on the Hanford Site to collect and track data needed to ensure compliance with an array of environmental laws, regulations, and policies governing HFO activities. Ecological monitoring data provide baseline information about the plants, animals, and habitats under HFO stewardship at the Hanford Site required for decision making under the National Environmental Policy Act (NEPA) and Comprehensive Environmental Response, Compensation, and Liability Act. DOE/EIS-0222, Final Hanford Comprehensive Land-Use Plan Environmental Impact Statement, (CLUP) evaluates the potential environmental impacts associated with implementing a comprehensive land-use plan for the Hanford Site for at least the next 50 years, and ensures that HFO, its contractors, and other entities conduct activities on the Hanford Site in compliance with NEPA.

54 ENVIRONMENTAL SCIENCES