Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data standard”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Global root traits (GRooT) database

Motivation: Trait data are fundamental to the quantitative description of plant form and function. Although root traits capture key dimensions related to plant responses to changing environmental conditions and effects on ecosystem processes, they have rarely been included in large-scale comparative studies and global models. For instance, root traits remain absent from nearly all studies that define the global spectrum of plant form and function. Thus, to overcome conceptual and methodological roadblocks preventing a widespread integration of root trait data into large-scale analyses we created the Global Root Trait (GRooT) Database. GRooT provides readyto- use data by combining the expertise of root ecologists with data mobilization and curation. Specifically, we (a) determined a set of core root traits relevant to the description of plant form and function based on an assessment by experts, (b) maximized species coverage through data standardization within and among traits, and (c) implemented data quality checks. Main types of variables contained: GRooT contains 114,222 trait records on 38 continuous root traits. Spatial location and grain: Global coverage with data from arid, continental, polar, temperate and tropical biomes. Data on root traits were derived from experimental studies and field studies. Time period and grain: Data were recorded between 1911 and 2019. Major taxa and level of measurement: GRooT includes root trait data for which taxonomic information is available. Trait records vary in their taxonomic resolution, with subspecies or varieties being the highest and genera the lowest taxonomic resolution available. It contains information for 184 subspecies or varieties, 6,214 species, 1,967 genera and 254 families. Owing to variation in data sources, trait records in the database include both individual observations and mean values. Software format: GRooT includes two csv files. A GitHub repository contains the csv files and a script in R to query the database.

59 BASIC BIOLOGICAL SCIENCES↗

Closing the Gap between FAIR Data Repositories and Hierarchical Data Formats

Many in the scientific community, particularly in publicly funded research, are pushing to adhere to more accessible data standards to maximize the findability, accessibility, interoperability, and reusability (FAIR) of scientific data, especially with the growing prevalence of machine learning augmented research. Online FAIR data repositories, such as the Open Science Framework (OSF), help facilitate the adoption of these standards by providing frameworks for storage, access, search, APIs, and other features that create organized hubs of scientific data. However, the wider acceptance of such repositories is hindered by the lack of support of hierarchical data formats, such as Technical Data Management Streaming (TDMS) and Hierarchical Data Format 5 (HDF5), that many researchers rely on to organize their datasets. Various tools and strategies should be used to allow hierarchical data formats, FAIR data repositories, and scientific organizations to work more seamlessly together. A pilot project at Los Alamos National Laboratory (LANL) addresses the disconnect between them by integrating the OSF FAIR data repository with hierarchical data renderers, extending support for additional file types in their framework. The multifaceted interactive renderer displays a tree of metadata alongside a table and plot of the data channels in the file. This allows users to quickly and efficiently load large and complex data files directly in the OSF webapp. Users who are browsing files can quickly and intuitively see the files in the way they or their colleagues structured the hierarchical form and immediately grasp their contents. This solution helps bridge the gap between hierarchical data storage techniques and FAIR data repositories, making both of them more viable options for scientific institutions like LANL which have been put off by the lack of integration between them.

97 MATHEMATICS AND COMPUTING↗

Livewire: Automatic Annotations

Diogenes processes datasets to provide data quality metrics for the Livewire platform and creates standardized data dictionaries from data annotations. Diogenes needs data annotations that clearly outline thenformat and organization of the data. It also relies on the type, class, and unit of each data piece for comprehensive analysis, which it cannot determine independently. The Annotation Tool significantly reduces the time needed to create annotations for Diogenes by generating data annotations with the correct formatting and content. It also employs machine learning and hard-coded models to automatically annotate data class, quality type, and data units.

33 - ADVANCED PROPULSION SYSTEMS↗

VizBrick: A GUI-based Interactive Tool for Authoring Semantic Metadata for Building Datasets

Brick ontology is a unified semantic metadata schema to address the stand-ardization problem of buildings' physical, logical, and virtual assets and the relationships between them. Creating a Brick model for a building dataset means that the dataset's contents are semantically described using the standard terms defined in the Brick ontology. It will enable the benefits of data standardization, without having to recollect or reorganize the data and opens the possibility of automation leveraging the machine readability of the semantic metadata. The problem is that authoring Brick models for building datasets often requires knowledge of semantic technology (e.g., on-tology declarations and RDF syntax) and leads to repeated manual trial and error processes, which can be time-consuming and challenging to do with-out an interactive visual representation of the data. We developed VizBrick, a tool with a graphical user interface that can assist users in creating Brick models visually and interactively without having to understand the Re-source Description Framework (RDF) syntax. VizBrick provides handy ca-pabilities such as keyword search for easy find of relevant brick concepts and relations to their data columns and automatic suggestions of concept mapping. In this demonstration, we present a use-case of VizBrick to show-case how a Brick model can be created for a real-world building dataset.

Lee, Sangkeun (Matt)↗

Measurement of the 252 Cf ⁢(sf) prompt fission neutron spectrum utilizing 12 C ⁡(𝑛, 𝑛) and 9 Be ⁢(𝑛, 𝑛) neutron scattering reference measurements

The 252 Cf spontaneous fission (sf), prompt fission neutron spectrum (PFNS) is a fundamental quantity for nuclear physics measurements of neutron-emitting reactions. This energy distribution of neutrons emitted from fission has been considered a neutron data standard for decades and has been utilized as a reference for neutron detection efficiency, validation of Monte Carlo simulations, benchmarking of dosimetry standards, and more. A significant portion of the global collection of nuclear data on neutron-induced reactions is correlated with the 252 Cf ⁢(sf) PFNS. Despite the reliance on this quantity by the nuclear physics community, the historical collection of 252 Cf PFNS measurements display systematic disagreements that are not understood or easily explained. These experimental discrepancies could potentially bias the 252 Cf PFNS Standard evaluation. On top of this, these past experiments frequently employed correlated experimental measurement or analysis methods. The artificial intelligence (AI)/machine learning (ML)-informed californium chi-nuclear data experiment (AIACHNE) project was formed to (a) investigate these discrepancies utilizing AI/ML methods to identify outlying regions of literature data, assign these regions to features of the experiment itself, and perform an improved evaluation of the 252 Cf PFNS and (b) perform a new experimental measurement of this quantity designed to improve upon the existing literature database. Here, in this work, we report on the AIACHNE 252 Cf PFNS experiment utilizing a new analysis method uncorrelated with all previous measurements: neutron efficiency determinations based on elastic neutron scattering on 12 C and 9 Be . This new method provides an independent test of the existing literature data and evaluation of the 252 Cf ⁢(sf) PFNS. The method is described with detailed covariance quantification procedures, as well as a direct discussion of the sources of uncertainty described as requirements in the “Templates” series of papers. The 252 Cf ⁢(sf) PFNS reported in this work agrees well with the overall shape of the existing standard PFNS evaluation as well as many literature measurements, thus verifying the current evaluation utilizing new techniques. However, the results suggest that there are deficiencies in the angle-differential 12 C and 9 Be ⁢(𝑛, 𝑛) evaluated nuclear data, which produce unphysical structures in the reported result. While these structures are relatively minor, they become obvious because of the high statistical precision of the data and the expected smooth continuity of the 252 Cf ⁢(sf) PFNS.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Towards provision of regularly updated climate data from the Coupled Model Intercomparison Project

The Coupled Model Intercomparison Project (CMIP) is a flagship of the World Climate Research Programme (WCRP). CMIP has become a recognised ‘brand’ in climate circles evolving over the last thirty years from a targeted research activity by a small number of climate modelling centres intercomparing their Earth System Model (ESM) simulations to a broad international coordinated research effort (Durack et al, 2025). CMIP is organized as a research activity leveraging funded and in-kind contributions from experts within modelling centres and the broader scientific community supported more recently by a fully-funded International Project Office. Within CMIP, Model Intercomparison Projects (MIPs) are community-designed to understand past, present and future climate. CMIP data provides a valuable resource for climate research and is routinely used to assess model representation of climate processes and test scientific hypotheses in the context of model uncertainty and (forced and internal) variability as evident from its prolific use in scientific publications1 . The impact relies on enabling infrastructure (most prominently via the Earth System Grid Federation (ESGF)), which allows sharing of simulation output, provision of the boundary conditions used in each simulation, and definition of the data standards that are essential to facilitating wide use of the data. The impact is supplemented by the wide-ranging scrutiny to which model simulations are subjected. Beyond its use in research, CMIP data is a key resource for communities producing derived climate information from downscaling and impact studies, such as the Coordinated Regional Downscaling Experiment (CORDEX; Gutowski et al., 2016) and the Intersectoral Impacts MIP (ISIMIP; Frieler et al., 2024). Government, academic and commercial entities also increasingly rely on CMIP and its downstream data for climate risk assessments and climate services (for example, Copernicus Climate Change Service and World Bank portal). This means that, although CMIP is a research activity, it increasingly serves a secondary and very relevant role as a provider of climate data – a long-recognised dichotomy (Stevens, 2024). Research and applications have distinct needs, with the former requiring flexibility and generality and the latter consistency. Here we explain how the design of the research activity has been adapted to reduce the burdens imposed by applications and how the research infrastructure might evolve to further enable scientific inquiry. We propose one possible approach to consistently providing model information and projections for applications in the future.

Environmental sciences↗

Solar Resource Measurements in Eugene, OR: Cooperative Research and Development Final Report, CRADA Number CRD-07-00252

Site-specific, long-term, continuous, and high-resolution measurements of solar irradiance are important for developing renewable resource data. These data are used for several research and development activities consistent with the NLR mission: establish a national 3-year climatological database of measured solar irradiances; provide high quality ground-truth data for satellite remote sensing validation; support development of radiative transfer models for estimating solar irradiance from available meteorological observations; provide solar resource information needed for technology deployment and operations. Data acquired under this agreement will be available to the public through NLR's Measurement & Instrumentation Data Center – MIDC (http://www.nlr.gov/midc) Or the Renewable Resource Data Center - RReDC (http://rredc.nlr.gov). The MIDC offers a variety of standard data display, access, and analysis tools designed to address the needs of a wide user audience (e.g., industry, academia, and government interests).

14 SOLAR ENERGY↗

radkit v1.2

The radkit software suite (python) consists of three primary libraries: stark, trajan, and curie. The trajan library provides the tools to analyze and manipulate data from lidar and inertial measurement unit (IMU) devices as well as trajectories from algorithms such as simultaneous localization and mapping (SLAM). These components allow reading and writing standard data formats, performing rigid affine transformations, discretizing three-dimensional space, and visualizing data products. The curie library comprises a standard set of object-oriented tools for radiation data and analysis in the following modules: (1) listmode and binmode data classes with methods for manipulation, plotting, slicing and file IO; (2) radiological/nuclear source detection/identification analysis results; (3) source encounters of correlated analyses and (4) energy-dependent angular detector response functions. The stark package provides low-level tools that are leveraged by both curie and trajan. The tools are flexible for offline analysis as well as performant for real-time integrations. The radkit libraries have associated Robot Operating System packages for use in real-time and robotic systems.

Joshi, Tenzing↗

radkit base v1.6

The radkit (base) software suite (python) consists of three primary libraries: stark, trajan, and curie. The trajan library provides the tools to analyze and manipulate data from lidar and inertial measurement unit (IMU) devices, cameras, as well as trajectories from algorithms such as simultaneous localization and mapping (SLAM). These components allow reading and writing standard data formats, performing rigid affine transformations, discretizing three-dimensional space, and visualizing data products. The curie library comprises a standard set of object-oriented tools for radiation data and analysis in the following modules: (1) listmode and binmode data classes with methods for manipulation, plotting, slicing and file IO; (2) radiological/nuclear source detection/identification analysis results; (3) source encounters of correlated analyses and (4) energy-dependent angular detector response functions. The stark package provides low-level tools that are leveraged by both curie and trajan. The tools are flexible for offline analysis as well as performant for real-time integrations.

Salathe, Marco [Lawrence Berkeley National Laborat↗

Sandia National Laboratories Ecosystem for Open Science: Metadata Schema v0.2 Description.

The Ecosystem for Open Science (eOS) initiative was established in 2019. Its objective is improving openness and sharing of data and information across Defense Nuclear Nonproliferation (DNN) Research and Development (R&D) activities. To support this initiative, the eOS team at Sandia National Laboratories (SNL) developed metadata and data standards and proposed a machine-readable metadata schema. The nuclear explosion monitoring field was selected as a focus area due to its the wide range of pertinent phenomenologies.We developed the DCAT-eOS-AP metadata schema extending the Data Catalog Vocabulary version 2 (DCATv2) standard using an application profile (AP), to fit the needs of multi-disciplinary NA-22 projects. The DCAT-eOS-AP metadata schema describes data at different levels of granularity ranging from general descriptions to more domain-specific granular metadata. Its implementation and serialization is flexible with the ability to include new file or data types. Thus, it will scale with the ever-increasing data management needs of government research. Due to the multitude of phenomenologies represented in the DCAT-eOS-AP schema, we anticipate that it will be easily extensible to various projects across many DOE mission areas. This document describes data management challenges faced within the DNN R&D portfolio and provides insight on how metadata and data standards/guidelines combined with a comprehensive metadata schema can add value to programs throughout the Department of Energy (DOE). It reviews the importance of metadata standards, FAIR (Findability, Accessibility, Interoperability, and Reusability) data principles, and metadata schemas. Additionally, it summarizes input from subject matter experts (SME) at SNL and other National Laboratories that resulted in metadata and data standards/guidelines encompassing domains relevant to NA-22 projects. Finally, we discuss the DCAT-eOS-AP metadata development. Implementation recommendations and future development directions are included for those keen on adopting the DCAT-eOS-AP metadata schema.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Legacy Survey of Space and Time Data Preview 2: standard_passband dataset type

We present Rubin Data Preview 2 (DP2), the second data preview from the NDF-DOE Vera C. Rubin Observatory. Data Preview 2 (DP2) comprises coadds, detection catalogs, and ancillary data products; and when fully released will also include single-epoch images and difference images. DP2 is derived from observations acquired by the LSST Science Camera (LSSTCam) on the Simonyi Survey Telescope at the Summit Facility on Cerro Pachón, Chile, primarily during the on-sky commissioning campaign between 2025-04-16 and 2025-09-21, supplemented by observations taken between 2025-10-25 and 2026-01-06 that overlap the commissioning footprint. The DP2 footprint comprises the Science Validation wide-area survey, five Deep Drilling Fields, and a number of targeted small-field regions, including Trifid and Lagoon, Prawn, M49, and New Horizons, all observed as part of the Rubin First Look campaign. Each field was imaged in up to six broad photometric bands, ugrizy, and coadded to produce deep imaging covering an estimated 3,000 deg2. The addition of single-visit-only areas expands the total DP2 footprint to an estimated 15,000 deg2, with coverage in at least one filter. The median per-visit PSF FWHM across the wide-area survey ranges from 1.17 arcsec in the z band to 1.26 arcsec in g and r bands. The deepest field, reaches estimated coadded 5σ depths of u=26 mag, g=26.8 mag, r=26.3 mag, i=26.1 mag, z=25.3 mag, y=23.9 mag. Based on a roughly five-month primary observing baseline and covering only part of the eventual LSST footprint, DP2's area, depth, and multiband coverage nonetheless support a broad range of early science investigations ahead of LSST Data Release This dataset is a subset of the full data release consisting of the standard_passband dataset type. These are the LSSTCam filter bandpasses. This release contains 6 datasets of this type.

79 ASTRONOMY AND ASTROPHYSICS↗

From remotely-sensed solar-induced chlorophyll fluorescence to ecosystem structure, function, and service: Part II—Harnessing data

Although our observing capabilities of solar-induced chlorophyll fluorescence (SIF) have been growing rapidly, the quality and consistency of SIF datasets are still in an active stage of research and development. As a result, there are considerable inconsistencies among diverse SIF datasets at all scales and the widespread applications of them have led to contradictory findings. The present review is the second of the two companion reviews, and data oriented. It aims to (1) synthesize the variety, scale, and uncertainty of existing SIF datasets, (2) synthesize the diverse applications in the sector of ecology, agriculture, hydrology, climate, and socioeconomics, and (3) clarify how such data inconsistency superimposed with the theoretical complexities laid out in may impact process interpretation of various applications and contribute to inconsistent findings. We emphasize that accurate interpretation of the functional relationships between SIF and other ecological indicators is contingent upon complete understanding of SIF data quality and uncertainty. Additionally, biases and uncertainties in SIF observations can significantly confound interpretation of their relationships and how such relationships respond to environmental variations. Built upon our syntheses, we summarize existing gaps and uncertainties in current SIF observations. Further, we offer our perspectives on innovations needed to help improve informing ecosystem structure, function, and service under climate change, including enhancing in-situ SIF observing capability especially in “data desert” regions, improving cross-instrument data standardization and network coordination, and advancing applications by fully harnessing theory and data.

59 BASIC BIOLOGICAL SCIENCES↗

Descriptor: High Temporal Resolution Meteorological Data at Oak Ridge Reservation (ORR-HiResMet)

Access to continuous, quality assessed meteorological data is critical for understanding the climatology and atmospheric dynamics of a region. Research facilities like Oak Ridge National Laboratory (ORNL) rely on such data to assess site-specific climatology, model potential emissions, establish safety baselines, and prepare for emergency scenarios. To meet these needs, on-site towers at ORNL collect meteorological data at 15-minute and hourly intervals. However, data measurements from meteorological towers are affected by sensor sensitivity, degradation, lightning strikes, power fluctuations, glitching, and sensor failures, all of which can affect data quality. To address these challenges, we conducted a comprehensive quality assessment and processing of five years of meteorological data collected from ORNL at 15-minute intervals, including measurements of temperature, pressure, humidity, wind, and solar radiation. The time series of each variable was pre-processed and gap-filled using established meteorological data collection and cleaning techniques, i.e., the time series were subjected to structural standardization, data integrity testing, automated and manual outlier detection, and gap-filling. The data product and highly generalizable processing workflow developed in Python Jupyter notebooks are publicly accessible online. As a key contribution of this study, the evaluated 5-year data will be used to train atmospheric dispersion models that simulate dispersion dynamics across the complex ridge-and-valley topography of the Oak Ridge Reservation in East Tennessee.

Steckler, Morgan R. [Oak Ridge National Laboratory↗

A NWB-based dataset and processing pipeline of human single-neuron activity during a declarative memory task

A challenge for data sharing in systems neuroscience is the multitude of different data formats used. Neurodata Without Borders: Neurophysiology 2.0 (NWB:N) has emerged as a standardized data format for the storage of cellular-level data together with meta-data, stimulus information, and behavior. A key next step to facilitate NWB:N adoption is to provide easy to use processing pipelines to import/export data from/to NWB:N. Here, we present a NWB-formatted dataset of 1863 single neurons recorded from the medial temporal lobes of 59 human subjects undergoing intracranial monitoring while they performed a recognition memory task. We provide code to analyze and export/import stimuli, behavior, and electrophysiological recordings to/from NWB in both MATLAB and Python. The data files are NWB:N compliant, which affords interoperability between programming languages and operating systems. This combined data and code release is a case study for how to utilize NWB:N for human single-neuron recordings and enables easy re-use of this hard-to-obtain data for both teaching and research on the mechanisms of human memory.

Science & Technology - Other Topics↗

NuHepMC: A standardized event record format for neutrino event generators

Simulations of neutrino interactions are playing an increasingly important role in the pursuit of high-priority measurements for the field of particle physics. A significant technical barrier for efficient development of these simulations is the lack of a standard data format for representing individual neutrino scattering events. We propose and define such a universal format, named NuHepMC, as a common standard for the output of neutrino event generators. The NuHepMC format uses data structures and concepts from the HepMC3 event record library adopted by other subfields of high-energy physics. These are supplemented with an original set of conventions for generically representing neutrino interaction physics within the HepMC3 infrastructure.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Film condensation with high heat fluxes and scaled experiments using pure steam for reactor containment cooling

Condensation tests were performed using a newly developed test facility for scaling the passive containment cooling system (PCCS) to a small modular reactor (SMR). The PCCS of the SMR plays a pivotal role in ensuring greater safety, reliability, and compactness than what is afforded by traditional reactors. Therefore, a well-designed PCCS is essential to SMRs. However, previous studies and test data were unsuitable for scaling, due to high variation in the test geometry and operating conditions. This study intends to close this research gap by using a novel designed scaled test facility consisting of vertical condensing test sections featuring 1-, 2-, and 4-inch-diameter condensing tubes with annular water cooling, and by applying superheated and saturated steam with different steam mass flow ranges of 5–25 g/s. Further, the primary test data, including axial temperatures, mass flow rates, and pressures, were used in conjunction with a standard data reduction method to estimate critical parameters such as heat fluxes, heat transfer coefficients, and condensation rates. These scaled test data would support improving empirical correlations and validating condensation models to identify scaling distortion for SMR PCCSs.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

MusMorph, a database of standardized mouse morphology data for morphometric meta-analyses

Complex morphological traits are the product of many genes with transient or lasting developmental effects that interact in anatomical context. Mouse models are a key resource for disentangling such effects, because they offer myriad tools for manipulating the genome in a controlled environment. Unfortunately, phenotypic data are often obtained using laboratory-specific protocols, resulting in self-contained datasets that are difficult to relate to one another for larger scale analyses. To enable meta-analyses of morphological variation, particularly in the craniofacial complex and brain, we created MusMorph, a database of standardized mouse morphology data spanning numerous genotypes and developmental stages, including E10.5, E11.5, E14.5, E15.5, E18.5, and adulthood. To standardize data collection, we implemented an atlas-based phenotyping pipeline that combines techniques from image registration, deep learning, and morphometrics. Alongside stage-specific atlases, we provide aligned micro-computed tomography images, dense anatomical landmarks, and segmentations (if available) for each specimen (N = 10,056). Our workflow is open-source to encourage transparency and reproducible data collection.

59 BASIC BIOLOGICAL SCIENCES↗