Engineering PapersSearch

SEARCH · Engineering Papers

Results for “data standards”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Lidar / Processed Data

This dataset contains processed, standardized data from the UND scanning Doppler lidar at WFIP3's BARG site, consisting of range- and time-resolved measurements of radial velocity, attenuated backscatter, intensity, and spectral width. We note that the beam azimuth angles are NOT referenced to true north.

17 WIND ENERGY

CACO Site - ANL Scanning Doppler Lidar / Processed Data

This dataset contains processed, standardized data from the ANL scanning Doppler lidar, consisting of range- and time-resolved measurements of radial velocity, attenuated backscatter, intensity, and spectral width. We note that the beam azimuth angles are NOT referenced to true north.

17 WIND ENERGY

Lidar / Processed Data

This dataset contains processed, standardized data from the UND scanning Doppler lidar at WFIP3's BARG site, consisting of range- and time-resolved measurements of radial velocity, attenuated backscatter, intensity, and spectral width. We note that the beam azimuth angles are NOT referenced to true north.

17 WIND ENERGY

NANT Site - Lidar / Processed Data Reformatted

This dataset contains standardized data from the PNNL scanning Doppler lidar (S/N 184), consisting of range- and time-resolved measurements of radial velocity, attenuated backscatter, intensity, and spectral width. We note that the beam azimuth angles are NOT referenced to true north.

17 WIND ENERGY

LinkML: an open data modeling framework

Background Scientific research relies on well-structured, standardized data; however, much of it is stored in formats such as free-text lab notebooks, nonstandardized spreadsheets, or data repositories. This lack of structure challenges interoperability, making data integration, validation, and reuse difficult. Findings LinkML (Linked Data Modeling Language) is an open framework that simplifies the process of authoring, validating, and sharing data. LinkML can describe a range of data structures, from flat, list-based models to complex, interrelated, and normalized models that utilize polymorphism and compound inheritance. It offers an approachable syntax that is not tied to any one technical architecture and can be integrated seamlessly with many existing frameworks. The LinkML syntax provides a standard way to describe schemas, classes, and relationships, allowing modelers to build well-defined, stable, and optionally ontology-aligned data structures. Once defined, LinkML schemas may be imported into other LinkML schemas. These key features make LinkML an accessible platform for interdisciplinary collaboration and a reliable way to define and share data semantics. Conclusions LinkML helps reduce heterogeneity, complexity, and the proliferation of single-use data models while simultaneously enabling compliance with FAIR (Findable, Accessible, Interoperable, and Reusable) data standards. LinkML has seen increasing adoption in various fields, including biology, chemistry, biomedicine, microbiome research, finance, electrical engineering, transportation, and commercial software development. In short, LinkML makes implicit models explicitly computable and allows data to be standardized at their origin. LinkML documentation and code are available at https://linkml.io/.

AI-ready data

HPC and Cloud Convergence Beyond Technical Boundaries: Strategies for Economic Sustainability, Standardization, and Data Accessibility

At the IEEE/ACM International Conference for High-Performance Computing, Networking, Storage, and Analysis (SC23), held in Denver, experts discussed the convergence of high-performance computing and cloud computing. Experts explored how this integration could address current scientific computing limitations, enhance computational capabilities, and foster global collaboration while focusing on economic, security, technical, and community challenges and opportunities.

97 MATHEMATICS AND COMPUTING

Data Cards for Standardized Metadata Across DOE-Aligned Data Initiatives: Toward Transparent, Interoperable, and Governed Dataset Documentation

As data-intensive research, advanced computing, and artificial intelligence become increasingly central to scientific and operational workflows, the need for consistent, transparent, and machine-actionable documentation has grown correspondingly. Multiple DOE-aligned communities—including Office of Science, Genesis Mission, American Science Cloud (AmSC), National Nuclear Security Administration (NNSA) stewardship and governance, and related cross-laboratory collaborations—have independently developed metadata practices to support discovery, access, reuse, repository deposit, and compliance.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION

2025 TEM Workshop

The TEM Data Management Workshop will take place on August 26 from 9 a.m. to 12 p.m. MT, and will be held virtually on TEAMS. The primary goal of this workshop is to engage NSUF users and stakeholders in discussions about the data needs for the utilization of AI and ML in the analysis of TEM data. Key topics to be covered include data storage, data sharing, data tagging, metadata inclusion, standardized data formats, data augmentation, and annotated training datasets. Additionally, the workshop will provide valuable insights into resources such as the Nuclear Research Data System (NRDS) for data storage and sharing, as well as open-source codes for data analysis.

Bachhav, Mukesh

Evaluating the factors influencing accuracy, interpretability, and reproducibility in the use of machine learning classifiers in biology to enable standardization

The complexity and variability of biological data has promoted the increased use of machine learning methods to understand processes and predict outcomes. These same features complicate reliable, reproducible, interpretable, and responsible use of such methods, resulting in questionable relevance of the derived. outcomes. Here we systematically explore challenges associated with applying machine learning to predict and understand biological processes using a well- characterized in vitro experimental system. We evaluated factors that vary while applying machine learning classifers: (1) type of biochemical signature (transcripts vs. proteins), (2) data curation methods (pre- and post-processing), and (3) choice of machine learning classifier. Using accuracy, generalizability, interpretability, and reproducibility as metrics, we found that the above factors significantly mod- ulate outcomes even within a simple model system. Our results caution against the unregulated use of machine learning methods in the biological sciences, and strongly advocate the need for data standards and validation tool-kits for such studies.

59 BASIC BIOLOGICAL SCIENCES

Comparability of Liquid Chromatography Tandem Mass Spectrometry Analysis of Dissolved Organic Matter across Laboratories

Non-targeted liquid chromatography tandem highresolution mass spectrometry (LC−MS/MS) is increasingly applied for the structure-resolved chemical analysis of dissolved organic matter (DOM). With new developments in MS instrumentation and analysis software, the approach has gained substantial momentum over the past decade. However, achieving high-quality analytical data that is reproducible and comparable across laboratories can be a bottleneck in non-targeted metabolomics and organic matter chemical analysis, especially for data reuse in repository-scale analyses. Understanding the capabilities as well as challenges of comparing LC−MS/MS data from different laboratories is necessary for inferring global trends from public data sets. To illuminate instrumentation factors that drive differences and variability, we used a standardized data analysis pipeline, including classical (CMN) and featurebased molecular networking (FBMN), to analyze data from a ring trial by 24 laboratories on identical sample sets of algal and DOM extracts that were mixed in predefined concentrations and spiked with standards. Our results showed that data sets from similar mass spectrometer types with unified instrument parameters were qualitatively comparable, resolving the same general trends and shared mass spectral features. Interlaboratory comparability was best for high-intensity features, while low-intensity features showed greater detection variability. Our analysis also highlights challenges when comparing data from instruments with different acquisition rates or operating with less standardized methods. Lastly, we provide recommendations for data integration, public data sharing, standardization, and best practices for standardized LC−MS/MS data acquisition, which will be critical for long-term time series and intercomparability of DOM chemical analyses.

DOM

Best Practices for Nuclear Experiment Data Preservation at Idaho National Laboratory: A Guide for Researchers and Reactor Operators

Preserving experimental data is essential for supporting advancements in nuclear science and ensuring the longevity of Idaho National Laboratory's contributions to reactor technology and safety. This report provides a comprehensive guide to best practices for experimental data management and preservation, focusing on standardized data formats, redundancy in storage, metadata documentation, and alignment with international standards. By following these recommendations, experimentalists and reactor operators can enhance the accessibility, reproducibility, and utility of critical datasets for regulatory review, validation computational methods, and future research.

22 GENERAL STUDIES OF NUCLEAR REACTORS

RC-SFA Data Management Templates and Guidance for Standardized, Reusable AI-Ready Data Packages

This data package provides templates and supporting documentation developed by the River Corridor Science Focus Area (RC-SFA; https://www.pnnl.gov/projects/river-corridor) to communicate its approach to managing and publishing AI-ready data. The package is intended to help data users and data producers understand the structures, metadata practices, and quality-control approaches that support consistent, reusable, and machine-actionable data products across RC-SFA studies. Rather than focusing on a single experimental dataset, this package documents the data management framework used to make RC-SFA data easier to find, ingest, navigate, and interpret. The materials in this package reflect RC-SFA practices for standardized data package organization, including the use of a human- and machine-readable README, file-level metadata, data dictionaries, descriptive file naming, method identifiers, and automated and review-based quality assurance procedures. Together, these components illustrate how RC-SFA extends FAIR data principles toward AI-readiness by prioritizing deep metadata, consistency across data packages, and support for informed downstream reuse by both humans and computational tools. This dataset is comprised of (1) readme; (2) presentation slides with an overview of RC-SFA approach and guidance; (3) document of RC-SFA best practices; (4) data dictionary (dd); (5) file level metadata (flmd); and a subfolder containing templates for dd and flmd. All files are .csv and .pdf. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About.

AI-readiness

Steam Condensation Scaled Experiment in the Presence of Non-condensable Gas for Reactor Containment Passive Safety Analysis

This study presents scaled experiments using steam condensation with non-condensable gas (NCG)—helium, simulating hydrogen—as these experiments are pivotal for water-cooled reactor passive containment cooling system (PCCS) design and analysis. Research into PCCSs for small modular reactors (SMRs) is especially important in light of SMR system design; however, studies in the literature reflect limitations due to test geometry and operational condition variations, without considering SMR prototypic design. To address these challenges, a scaled test facility was developed to accurately replicate SMR PCCSs. This facility includes vertical down-flow condensing test sections with 1-, 2-, and 4-in.-diameter condensing tubes, accompanied by annular water cooling. Experiments were conducted using both superheated and saturated steam, with steam mass flow rates varying from 55 to 66 kg/hr., in the presence of helium as the NCG mass flow rate ranges from 1.8 to 22 kg/hr. Test data were collected on (a) the axial temperatures of the annular cooling water; (b) the outer wall temperature of the condensers; and (c) the mass flow rate, temperature, and pressure at the test section inlets and outlets. These primary test data were used in conjunction with a standard data reduction methodology to estimate essential thermal parameters such as heat fluxes, heat transfer coefficients, and condensation rates. The effects of NCGs on steam condensation within the geometry of the scaled test sections were then presented in regard to various testing conditions.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS

Presentation: Steam Condensation Scaled Experiment in the Presence of Non-condensable Gas for Reactor Containment Passive Safety Analysis

This study presents scaled experiments using steam condensation with non-condensable gas (NCG)—helium, simulating hydrogen—as these experiments are pivotal for water-cooled reactor passive containment cooling system (PCCS) design and analysis. Research into PCCSs for small modular reactors (SMRs) is especially important in light of SMR system design; however, studies in the literature reflect limitations due to test geometry and operational condition variations, without considering SMR prototypic design. To address these challenges, a scaled test facility was developed to accurately replicate SMR PCCSs. This facility includes vertical down-flow condensing test sections with 1-, 2-, and 4-in.-diameter condensing tubes, accompanied by annular water cooling. Experiments were conducted using both superheated and saturated steam, with steam mass flow rates varying from 55 to 66 kg/hr., in the presence of helium as the NCG mass flow rate ranges from 1.8 to 22 kg/hr. Test data were collected on (a) the axial temperatures of the annular cooling water; (b) the outer wall temperature of the condensers; and (c) the mass flow rate, temperature, and pressure at the test section inlets and outlets. These primary test data were used in conjunction with a standard data reduction methodology to estimate essential thermal parameters such as heat fluxes, heat transfer coefficients, and condensation rates. The effects of NCGs on steam condensation within the geometry of the scaled test sections were then presented in regard to various testing conditions.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS

Microbiome data management in action workshop: Atlanta, GA, USA, June 12–13, 2024

Microbiome research is revolutionizing human and environmental health, but the value and reuse of microbiome data are significantly hampered by the limited development and adoption of data standards. While several ongoing efforts are aimed at improving microbiome data management, significant gaps still remain in terms of defining and promoting adoption of consensus standards for these datasets. The Strengthening the Organization and Reporting of Microbiome Studies (STORMS) guidelines for human microbiome research have been endorsed and successfully utilized by many research organizations, publishers, and funding agencies, and have been recognized as a consensus community standard. No equivalent effort has occurred for environmental, synthetic, and non-human host-associated microbiomes. To address this growing need within the microbiome research community, we convened the Microbiome Data Management in Action Workshop (June 12–13, 2024, in Atlanta, GA, USA), to bring together key decision makers in microbiome science including researchers, publishers, funders, and data repositories. The 50 attendees, representing the diverse and interdisciplinary nature of microbiome research, discussed recent progress and challenges, and brainstormed actionable recommendations and paths forward for coordinated environmental microbiome data management and the modifications necessary for the STORMS guidelines to be applied to environmental, non-human host, and synthetic microbiomes. The outcomes of this workshop will form the basis of a formalized data management roadmap to be implemented across the field. These best practices will drive scientific innovation now and in years to come as these data continue to be used not only in targeted reanalyses but in large-scale models and machine learning efforts.

54 ENVIRONMENTAL SCIENCES

From Silos to Synergy: Identifying a Roadmap for Cross-Sector Research to Accelerate the Clean Energy Transition

The U.S. Department of Energy's blueprints for the transportation, buildings, and electricity sectors call for substantial reductions in greenhouse gas (GHG) emissions by 2050. These plans focus on zero-emission vehicles, investments in transit, energy-efficient buildings, and the widespread adoption and deployment of renewable energy technologies like solar photovoltaics (PV), energy storage and energy-efficient appliances. However, these sectors are often studied and modeled in isolation, overlooking how household decisions to adopt clean technologies in one sector influence others. This study, led by an interdisciplinary team at the National Renewable Energy Laboratory (NREL), explores opportunities for cross-sector collaboration to drive more effective and equitable decarbonization. Through discussions with 22 NREL researchers across transportation, building, solar, and grid sectors, the study highlights the need for integrated tools and models that capture interactions between these sectors. Key insights include the need for data standardization and interoperability to enable cross-sector analysis and decision-making. Strengthening utility partnerships is also critical to align energy policies with decarbonization goals and manage the increased demand for renewable energy. The study also emphasizes the importance of equity in the clean energy transition, calling for targeted incentives and support to ensure that low-income and underserved communities benefit from clean technologies like electric vehicles and energy-efficient appliances. To support these efforts, innovative funding mechanisms must be expanded to facilitate interdisciplinary research, such as city-specific decarbonization plans and federal projects like DOE"s Standard Scenarios. By encouraging collaboration and integrating cross-sector insights, this study aims to provide a roadmap to accelerate the clean energy transition and ensure it is both sustainable and inclusive.

14 SOLAR ENERGY

Livewire: Automatic Annotations

Diogenes processes datasets to provide data quality metrics for the Livewire platform and creates standardized data dictionaries from data annotations. Diogenes needs data annotations that clearly outline thenformat and organization of the data. It also relies on the type, class, and unit of each data piece for comprehensive analysis, which it cannot determine independently. The Annotation Tool significantly reduces the time needed to create annotations for Diogenes by generating data annotations with the correct formatting and content. It also employs machine learning and hard-coded models to automatically annotate data class, quality type, and data units.

33 - ADVANCED PROPULSION SYSTEMS