Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “standardized data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Descriptor: High Temporal Resolution Meteorological Data at Oak Ridge Reservation (ORR-HiResMet)

Access to continuous, quality assessed meteorological data is critical for understanding the climatology and atmospheric dynamics of a region. Research facilities like Oak Ridge National Laboratory (ORNL) rely on such data to assess site-specific climatology, model potential emissions, establish safety baselines, and prepare for emergency scenarios. To meet these needs, on-site towers at ORNL collect meteorological data at 15-minute and hourly intervals. However, data measurements from meteorological towers are affected by sensor sensitivity, degradation, lightning strikes, power fluctuations, glitching, and sensor failures, all of which can affect data quality. To address these challenges, we conducted a comprehensive quality assessment and processing of five years of meteorological data collected from ORNL at 15-minute intervals, including measurements of temperature, pressure, humidity, wind, and solar radiation. The time series of each variable was pre-processed and gap-filled using established meteorological data collection and cleaning techniques, i.e., the time series were subjected to structural standardization, data integrity testing, automated and manual outlier detection, and gap-filling. The data product and highly generalizable processing workflow developed in Python Jupyter notebooks are publicly accessible online. As a key contribution of this study, the evaluated 5-year data will be used to train atmospheric dispersion models that simulate dispersion dynamics across the complex ridge-and-valley topography of the Oak Ridge Reservation in East Tennessee.

Steckler, Morgan R. [Oak Ridge National Laboratory↗

NuHepMC: A standardized event record format for neutrino event generators

Simulations of neutrino interactions are playing an increasingly important role in the pursuit of high-priority measurements for the field of particle physics. A significant technical barrier for efficient development of these simulations is the lack of a standard data format for representing individual neutrino scattering events. We propose and define such a universal format, named NuHepMC, as a common standard for the output of neutrino event generators. The NuHepMC format uses data structures and concepts from the HepMC3 event record library adopted by other subfields of high-energy physics. These are supplemented with an original set of conventions for generically representing neutrino interaction physics within the HepMC3 infrastructure.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Film condensation with high heat fluxes and scaled experiments using pure steam for reactor containment cooling

Condensation tests were performed using a newly developed test facility for scaling the passive containment cooling system (PCCS) to a small modular reactor (SMR). The PCCS of the SMR plays a pivotal role in ensuring greater safety, reliability, and compactness than what is afforded by traditional reactors. Therefore, a well-designed PCCS is essential to SMRs. However, previous studies and test data were unsuitable for scaling, due to high variation in the test geometry and operating conditions. This study intends to close this research gap by using a novel designed scaled test facility consisting of vertical condensing test sections featuring 1-, 2-, and 4-inch-diameter condensing tubes with annular water cooling, and by applying superheated and saturated steam with different steam mass flow ranges of 5–25 g/s. Further, the primary test data, including axial temperatures, mass flow rates, and pressures, were used in conjunction with a standard data reduction method to estimate critical parameters such as heat fluxes, heat transfer coefficients, and condensation rates. These scaled test data would support improving empirical correlations and validating condensation models to identify scaling distortion for SMR PCCSs.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

MusMorph, a database of standardized mouse morphology data for morphometric meta-analyses

Complex morphological traits are the product of many genes with transient or lasting developmental effects that interact in anatomical context. Mouse models are a key resource for disentangling such effects, because they offer myriad tools for manipulating the genome in a controlled environment. Unfortunately, phenotypic data are often obtained using laboratory-specific protocols, resulting in self-contained datasets that are difficult to relate to one another for larger scale analyses. To enable meta-analyses of morphological variation, particularly in the craniofacial complex and brain, we created MusMorph, a database of standardized mouse morphology data spanning numerous genotypes and developmental stages, including E10.5, E11.5, E14.5, E15.5, E18.5, and adulthood. To standardize data collection, we implemented an atlas-based phenotyping pipeline that combines techniques from image registration, deep learning, and morphometrics. Alongside stage-specific atlases, we provide aligned micro-computed tomography images, dense anatomical landmarks, and segmentations (if available) for each specimen (N = 10,056). Our workflow is open-source to encourage transparency and reproducible data collection.

59 BASIC BIOLOGICAL SCIENCES↗

Challenges and Vision for Standardization of Biopolymer Data Sets for Machine Learning

Machine learning (ML) is transforming materials research, yet potential for biopolymer discovery remains constrained by fragmented data and nonstandardized reporting. Biopolymers differ significantly from synthetic polymers, requiring specialized approaches to represent their biosynthetic origins, hierarchical structures, and application-specific metrics. In this Perspective, we identify three core challenges limiting biopolymer representation: information encoding, data quality, and data sharing. We describe the most pressing issues and propose commensurate approaches to address each key challenge. Recommendations include the design and adoption of biopolymer-specific fingerprinting and representation frameworks, development of hybrid human-large language model (LLM) data extraction strategies, and expanding Findable, Accessible, Interoperable, Reusable (FAIR)-compliant repositories. We propose a robust foundation to define interoperable, high-quality data sets that capture the full context of biopolymer materials. Standardized metadata, shared ontologies, and community-driven infrastructure would enable scalable, reproducible workflows and accelerate the ML-driven development of biopolymers.

36 MATERIALS SCIENCE↗

Classification of events from α -induced reactions in the MUSIC detector via statistical and ML methods

The Multi-Sampling Ionization Chamber (MUSIC) detector is typically used to measure nuclear reaction cross sections relevant for nuclear astrophysics, fusion studies, and other applications. From the MUSIC data produced in one experiment scientists carefully extract an order of 10 3 events of interest from about 10 9 total events, where each event can be represented by an 18-dimensional vector. However, the standard data classification process is based on expert driven, manually intensive data analysis techniques that require several months to identify patterns and classify the relevant events from the collected data. Here, to address this issue, we present a method for the classification of events originating from specific α-induced reactions by combining statistical and machine learning methods that require significantly less input from the domain scientist, relative to the standard technique. Here, we applied the new method to two experimental data sets and compared our results with those obtained using traditional methods. With few exceptions, the number of events classified by our method agrees within ±20% with the results obtained using traditional methods. With the present method, which is the first of its kind for the MUSIC data, we have established the foundation for the automated extraction of physical events of interest from experiments using the MUSIC detector.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Framework for custom event sample augmentations for ATLAS analysis data

For HEP event processing, data is typically stored in column-wise synchronized containers, such as most prominently ROOT’s TTree, which have been used for several decades to store by now over 1 exabyte. These containers can combine row-wise association capabilities needed by most HEP event processing frameworks (e.g. Athena for ATLAS) with column-wise storage, which typically results in better compression and more efficient support for many analysis use-cases. One disadvantage is that these containers, TTree in the HEP use-case, require to contain the same attributes for each entry/row (representing events), which can make extending the list of attributes very costly in storage, even if those are only required for a small subsample of events. Since the initial design, the ATLAS software framework features powerful navigational infrastructure to allow storing custom data extensions for subsamples of events in separate, but synchronized containers. This allows adding event augmentations to ATLAS standard data products (such as DAOD-PHYS or PHYSLITE) avoiding duplication of those core data products, while limiting their size increase. For this functionality, the framework does not rely on any associations made by the I/O technology (i.e. ROOT), however it supports TTree friends and builds the associated index to allow for analysis outside of the ATLAS framework. A prototype based on the Long-Lived Particle search is implemented and preliminary results with this prototype will be presented. At this point, augmented data are stored within the same file as the core data. Storing them in separate files will be investigated in future, as this could provide more flexibility, e.g. certain sites may only want a subset of several augmentations or augmentations can be archived to tape once their analysis is complete.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

What's New in the NSRDB

The National Solar Radiation Database (NSRDB) provides solar resource data across the globe at a high temporal and spatial resolution. This data is primarily used in solar energy modeling. The NSRDB is updated annually for the United States and North, Central and South America and the data is currently available from 1998-2021. In 2022 the NSRDB was updated using the latest version of the underlying Physical Solar Model (PSM). This update includes improved surface albedo and gap-filling of cloud properties. The inclusion of these updates reduced the uncertainty in the data compared to previous versions of the NSRDB. The Himawari and Meteosat Indian Ocean Data Coverage (IODC) satellites were added to the Geostationary Operational Environmental Satellite (GOES) and made our coverage global. While standard data from the GOES continues to be served at an hourly 4km x 4km resolution, full resolution data has also been made available to the user. The NSRDB now contains over 200Tb of data with nearly 40Tb being added annually. We provide significant flexibility for data download depending on the amount of data required by the users. In this paper we provide an update on the current status on the NSRDB.

photovoltaic systems↗

Baseline Climate Variables for Earth System Modelling

The Baseline Climate Variables for Earth System Modelling (ESM-BCVs) are defined as a list of 135 variables which have high utility for the evaluation and exploitation of climate simulations. The list reflects the most frequently used elements of the Coupled Model Intercomparison Project Phase 6 (CMIP6) archive. Successive phases of CMIP have supported strong results in science and substantially influence international climate policy formulation. This paper responds to both interest in exploiting CMIP data standards in a broader range of climate modelling activities and a need to achieve greater clarity about the significance and intention of variables in the CMIP Data Request. As Earth system modelling archives grow in scale and complexity, there are emerging problems associated with weak standardisation at the variable collection level. That is, there are good standards covering how specific variables should be archived, but this paper fills a gap in the standardisation of which variables should be archived. The ESM-BCV list is intended as a resource for ESM intercomparison projects (MIPs) developing requests to enable greater consistency among MIPs and as a reference for modelling centres to enhance consistency within MIPs. Provisional planning for the CMIP7 Data Request exploits the ESM-BCVs as a core element. The baseline variable list includes 98 variables which have modest or minor data volume footprints and could be generated systematically when simulations are produced and archived for exploitation by the World Climate Research Programme (WCRP) community. A further 35 variables are classed as “high volume” and are only suitable for production when the resource implications are justified.

Juckes, Martin [University of Oxford (United Kingd↗

Characterization of actinide abundances and isotopic compositions by HR-ICP-MS. Part 2: Results from actinide doping studies

Previously, we presented actinide isotopic and elemental data from fallout melt glass that was measured using high-resolution inductively coupled plasma mass spectrometry (HR-ICP-MS). Direct comparison between these measurements and ‘gold standard’ data obtained by multi-collector ICP-MS showed broad overlap, indicating that HR-ICP-MS is a potentially valuable technique for producing actinide elemental and isotopic data on a relatively rapid timescale. To test the usefulness of this technique further, we doped varying amounts of uranium certified reference materials (CRMs) into a rhyolitic rock standard to establish the effects of uranium concentration and isotopic composition on the accuracy of uranium isotopic analyses by HR-ICP-MS. This also enabled us to quantify peak tailing effects from 238 U on the measurement of 239 Pu and 237 Np, which in turn allows us to constrain correction factors based on measured 237 Np/ 238 U and 239 Pu/ 238 U ratios. A second doping study involved the addition of a mixed actinide standard into three samples with different matrix compositions (termed soil, city, and seawater) to assess whether sample chemistry affects the accuracy and precision of these analyses. Results suggest that this is not the case. Systematic offsets were not observed in elemental or isotopic data derived from the three matrix samples. Results indicate that useful actinide isotopic data can be obtained from whole rock solutions by HR-ICP-MS. Our findings also have implications for solid sampling techniques such as laser ablation ICP-MS, which do not require sample dissolution.

38 RADIATION CHEMISTRY, RADIOCHEMISTRY, AND NUCLEA↗

TropiRoot 1.0: Database of tropical root characteristics across environments

Tropical ecosystems contain the world's largest biodiversity of vascular plants. Yet, our understanding of tropical functional diversity and its contribution to global diversity patterns is constrained by data availability. This discrepancy underscores an urgent need to bridge data gaps by incorporating comprehensive tropical root data into global datasets. Here, we provide a database of tropical root characteristics. This new database, TropiRoot 1.0, will be instrumental in evaluating an array of hypotheses pertaining to root functional ecology and plant biogeography, both within the tropics and relative to other global biomes. The data compilation was conducted by the TropiRoot Initiative, in partnership with the Fine-Root Ecology Database (FRED) and the Global Root Trait (GRooT) database, Colorado State University (CSU) and the Smithsonian Tropical Research Institute (STRI). Literature search and data extraction were conducted between 2020 and 2024. Literature was identified using Web of Science, Scopus, and complemented using the expert knowledge of members of TropiRoot. To provide broad environmental and geographical distributions, literature searches included root characteristics (traits) across global change drivers, natural gradients, and from different continents. We adopted FRED standardized data columns and streamlined the format to enhance accessibility for data extraction across various user groups. This optimized framework resulted in a smaller, yet comprehensive datasheet. To make the database compatible with other global root trait initiatives, column identification was standardized following the codes provided by FRED. These efforts culminated in data extracted from 104 new sources, resulting in more than 8000 rows of data (either species or community data). Most of the data in TropiRoot 1.0 include root characteristics such as root biomass, morphology, root dynamics, mass fraction, architecture, anatomy, physiology, and root chemistry. This initiative represents a 30% increase in the currently available data for tropical roots in FRED. TropiRoot 1.0 contains root characteristics from 25 different countries, where seven are located in Asia, six in South America, five in Central America and the Caribbean, four in Africa, two in North America, and 1 in Oceania. Due to the volume of data, when ancillary data were available, including soil data, these data were either extracted and included in the database or its availability was recorded in an additional column. Multiple contributors checked the entries for outliers during the collation process to ensure data quality. For text-based observations, we examined all cells to ensure that their content relates to their specific categories. For numerical observations, we ordered each numerical value from least to greatest and plotted the values, checking apparent outliers against the data in their respective sources and correcting or removing incorrect or impossible values. Some data (soil and aboveground) have different columns for the same variable presented in different units, including originally published units, but root characteristics data had units converted to match those reported in FRED. By filling a gap from global databases, TropiRoot 1.0 expands our knowledge of otherwise so far underrepresented regions and our ability to assess global trends. This advancement can be used to improve tropical forest representation in vegetation models. The data are freely available and should be cited when used.

FRED↗

Post-composing ontology terms for efficient phenotyping in plant breeding

Abstract Ontologies are widely used in databases to standardize data, improving data quality, integration, and ease of comparison. Within ontologies tailored to diverse use cases, post-composing user-defined terms reconciles the demands for standardization on the one hand and flexibility on the other. In many instances of Breedbase, a digital ecosystem for plant breeding designed for genomic selection, the goal is to capture phenotypic data using highly curated and rigorous crop ontologies, while adapting to the specific requirements of plant breeders to record data quickly and efficiently. For example, post-composing enables users to tailor ontology terms to suit specific and granular use cases such as repeated measurements on different plant parts and special sample preparation techniques. To achieve this, we have implemented a post-composing tool based on orthogonal ontologies providing users with the ability to introduce additional levels of phenotyping granularity tailored to unique experimental designs. Post-composed terms are designed to be reused by all breeding programs within a Breedbase instance but are not exported to the crop reference ontologies. Breedbase users can post-compose terms across various categories, such as plant anatomy, treatments, temporal events, and breeding cycles, and, as a result, generate highly specific terms for more accurate phenotyping.

Mathematical & Computational Biology↗

The Marine and Hydrokinetic ToolKit for Data Quality Control and Analysis: Preprint

The ability to handle data is critical at all stages of marine energy (ME) development. The marine hydrokinetic toolkit (MHKiT) is an open-source marine energy software, which includes modules for ingesting, applying quality control, processing, visualizing, and managing data. MHKiT-Python and MHKiT-MATLAB provide robust and verified functions that are needed by the ME community to standardize data processing. Calculations and visualizations adhere to International Electrotechnical Commission (IEC) technical specifications and other guidelines. A resource assessment of NDBC buoy 46050 near PACWAVE is performed using MHKiT and discusses comparisons to the resource assessment provided performed by Dunkel et al.

marine energy↗

The Marine and Hydrokinetic Toolkit (Mhkit) for Data Quality Control and Analysis

The ability to handle data is critical at all stages of marine energy (ME) development. The marine hydrokinetic toolkit (MHKiT) is an open-source marine energy software, which includes modules for ingesting, applying quality control, processing, visualizing, and managing data. MHKiT-Python and MHKiT-MATLAB provide robust and verified functions that are needed by the ME community to standardize data processing. Calculations and visualizations adhere to International Electrotechnical Commission (IEC) technical specifications and other guidelines. A resource assessment of NDBC buoy 46050 near PACWAVE is performed using MHKiT and discusses comparisons to the resource assessment provided performed by Dunkel et al.

marine energy↗

Analysis of overlapping count data

Counts of a specific characteristic were obtained within regions defined on an object that was manufactured in a proprietary setting. The count regions were altered during production and resulted in misaligned or overlapping count data. A closed-formula maximum likelihood estimator (MLE) of the new region means is derived using all of the available count data and an independent Poisson model. The MLE is shown to be preferable to estimators constructed using generalized linear models for the overlapping data setting. This closed-form estimator extends to over-dispersed overlapping count data as the quasi-MLE and also performs well with correlated overlapping count data. Standard errors for the estimator are approximated and are validated with a simulation study. Additionally, the methods are extended to overlapping multinomial data. Illustrative examples of the methods are provided throughout the paper and are reproducible with the supplemental R code. Additionally, proofs of the paper’s results are also included in the supplemental material.

97 MATHEMATICS AND COMPUTING↗

Re-evaluating the prompt fission neutron spectrum of spontaneously fissioning 252 Cf

The prompt fission neutron spectrum (PFNS) of spontaneously fissioning 252 Cf is a Neutron Data Standards observable. Nearly all fission spectra of actinides were measured relative to it, using efficiencies derived from it, or analyzed with simulations validated by it. The current Standards evaluation was published by W. Mannhart in 1987. It could not be updated because the evaluation input, experimental mean values and covariances, were lost. First, we attempt to reproduce it. However, Mannhart’s evaluation can only be reproduced within its one-σ uncertainties as some of its aspects (e.g., experimental covariances, rejected data points) remain unknown. Therefore, a new evaluation is presented: We revisit all existing experimental 252 Cf(sf) PFNS data, including those published after the release of the current Standards evaluation, and re-estimate associated covariances. The newly evaluated 252 Cf(sf) PFNS differs distinctly from Mannhart’s below 300 keV and extends it to lower and higher outgoing neutron energies (500 eV–25 MeV). The new evaluated uncertainties are larger from 3–9 MeV and smaller otherwise. Spectrum averaged cross sections of importance to the International Reactor Dosimetry and Fusion File community calculated with the new spectrum are close to those calculated with Mannhart’s evaluation and agree with experimental values well within their uncertainties.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Characterizing peak electricity demand for U.S. households: an assessment of end-use loads and demand factors

Understanding household peak electricity demand is critical to evaluate the technical need for electrical infrastructure upgrades. This study characterizes peak loads for existing and new equipment using metered data from a convenience sample of 11,940 U.S. dwellings from four sources, including 911 from two sources with end-use metering. After standardized data cleaning and labeling, we derived descriptive statistics for key metrics, such as maximum demand and demand factors, and developed predictive models relating 60- to 15-min demand for the National Electrical Code (NEC). Mean 15-min maximum demand was 9.7 kW (median 9.0 kW; IQR 7.0–11.5 kW, 95% CI 9.6–9.8 kW), indicating spare capacity in 98% of homes with hypothetical 100 A panels. Maximum demand increased with floor area and number of high-demand loads. Dwelling maximum demand was driven by higher-power, longer-duration heating appliances and vehicle charging, while most user-operated appliances contributed little. Demand factors are used to account for how most devices contribute less than their rated power to maximum demand. Existing load mean demand factors (28%; median 10%; IQR 0–58%; CI 28–29%) were higher than those for new loads (21%; median 7%; IQR 0–35%; CI 20–21%), because new loads changed the timing and magnitude of maximum demand. New high-demand loads had higher than average demand factors (40–60%). Whole dwelling demand factors support the NEC's 40% assumption, but they challenge its conservative 100% treatment of new HVAC. We propose a data-driven 50% demand factor for new equipment, which would align with metered data, improve affordability, and modernize electrical codes.

Appliances↗

Reactor Containment Passive Safety Analysis: Steam Condensation in Presence of Non-condensable Gas Scaled Experiment and Modeling

This study presents steam condensation scaled experiments and semi-empirical models in presence of nitrogen (N)—a noncondensable gas (NCG), simulating air in the reactor containment—to support water-cooled small modular reactors (SMRs) passive containment cooling system (PCCS) design and analysis. Previous experimental studies on PCCS are focused on fixed and smaller tube (mostly 2-in.) geometries and specific test condition variations, bringing challenges with geometric scaling and mismatching with SMR prototypic design. To address these challenges, this study presents steam condensation test dataset obtained from three scaled test sections of 1-, 2-, and 4-in.-diameter steam condensers with an annular/jacket cooling of 2-, 3-, and 6 in.-diameter tubes, respectively. Test data were collected for steam ranges from 58 to 63 kg/hr., and NCG flow of 4.4 to 13.3 kg/hr. Annular cooling water flow was varied to obtain required testing conditions of saturated steam inlet and fully condensed outlet. Axial temperature test data of bulk cooling water, steam and condensate were collected by thermocouples for three test sections and various steam-NCG mixing/testing conditions. A standard data reduction method was adopted—utilizing iterative and nodalized mass and heat transfer calculation—to estimate axial local heat fluxes, heat transfer coefficients (HTCs), condensation rates, film thickness, and Nusselt number. Based on the obtained dataset semi-empirical model results—a ratio of experimental and Nusselt’s theoretical HTC are presented. Such results and findings are supportive of developing scaled-up testing facility, to enable model validations and accelerate next generation of reactors development and deployment

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗