Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data standard”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Challenges and Vision for Standardization of Biopolymer Data Sets for Machine Learning

Machine learning (ML) is transforming materials research, yet potential for biopolymer discovery remains constrained by fragmented data and nonstandardized reporting. Biopolymers differ significantly from synthetic polymers, requiring specialized approaches to represent their biosynthetic origins, hierarchical structures, and application-specific metrics. In this Perspective, we identify three core challenges limiting biopolymer representation: information encoding, data quality, and data sharing. We describe the most pressing issues and propose commensurate approaches to address each key challenge. Recommendations include the design and adoption of biopolymer-specific fingerprinting and representation frameworks, development of hybrid human-large language model (LLM) data extraction strategies, and expanding Findable, Accessible, Interoperable, Reusable (FAIR)-compliant repositories. We propose a robust foundation to define interoperable, high-quality data sets that capture the full context of biopolymer materials. Standardized metadata, shared ontologies, and community-driven infrastructure would enable scalable, reproducible workflows and accelerate the ML-driven development of biopolymers.

36 MATERIALS SCIENCE↗

Classification of events from α -induced reactions in the MUSIC detector via statistical and ML methods

The Multi-Sampling Ionization Chamber (MUSIC) detector is typically used to measure nuclear reaction cross sections relevant for nuclear astrophysics, fusion studies, and other applications. From the MUSIC data produced in one experiment scientists carefully extract an order of 10 3 events of interest from about 10 9 total events, where each event can be represented by an 18-dimensional vector. However, the standard data classification process is based on expert driven, manually intensive data analysis techniques that require several months to identify patterns and classify the relevant events from the collected data. Here, to address this issue, we present a method for the classification of events originating from specific α-induced reactions by combining statistical and machine learning methods that require significantly less input from the domain scientist, relative to the standard technique. Here, we applied the new method to two experimental data sets and compared our results with those obtained using traditional methods. With few exceptions, the number of events classified by our method agrees within ±20% with the results obtained using traditional methods. With the present method, which is the first of its kind for the MUSIC data, we have established the foundation for the automated extraction of physical events of interest from experiments using the MUSIC detector.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Framework for custom event sample augmentations for ATLAS analysis data

For HEP event processing, data is typically stored in column-wise synchronized containers, such as most prominently ROOT’s TTree, which have been used for several decades to store by now over 1 exabyte. These containers can combine row-wise association capabilities needed by most HEP event processing frameworks (e.g. Athena for ATLAS) with column-wise storage, which typically results in better compression and more efficient support for many analysis use-cases. One disadvantage is that these containers, TTree in the HEP use-case, require to contain the same attributes for each entry/row (representing events), which can make extending the list of attributes very costly in storage, even if those are only required for a small subsample of events. Since the initial design, the ATLAS software framework features powerful navigational infrastructure to allow storing custom data extensions for subsamples of events in separate, but synchronized containers. This allows adding event augmentations to ATLAS standard data products (such as DAOD-PHYS or PHYSLITE) avoiding duplication of those core data products, while limiting their size increase. For this functionality, the framework does not rely on any associations made by the I/O technology (i.e. ROOT), however it supports TTree friends and builds the associated index to allow for analysis outside of the ATLAS framework. A prototype based on the Long-Lived Particle search is implemented and preliminary results with this prototype will be presented. At this point, augmented data are stored within the same file as the core data. Storing them in separate files will be investigated in future, as this could provide more flexibility, e.g. certain sites may only want a subset of several augmentations or augmentations can be archived to tape once their analysis is complete.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Radiation Detection Data Competition Report

In FY2018 through FY2020, NA-22, the Defense Nuclear Nonproliferation Research and Development Program, funded a Data Science project to develop and implement statistical methodology to effectively host data competitions with the goal of leveraging the opportunity provided by crowdsourcing. By accessing and engaging expertise from a broader research community, there is an opportunity to attract innovative solutions from a variety of different research disciplines to advance the ability to solve important non-proliferation problems. This report summarizes the key results of this project after hosting two data competitions focused on urban radiation detection. The first competition was focused on attracting participants from the U.S. national laboratories, while the second, hosted by TopCoder, was open to the broader international community and awarded prize money to the top 10 competitors. At the start of the project, there was strong interest from NA-22 to explore and develop the capability to host data competitions as a means of leveraging the broader community to solve important nuclear nonproliferation problems. Having a standard data set on which to compare different approaches based on clearly defined criteria was desirable to be able to evaluate the state of solutions for important problems. Initially, it was not clear that it would even be possible logistically and bureaucratically to host a competition with an international field of competitors and to award the prize money needed to attract solutions from top competitors. Happily, a path to host the competitions was ultimately found that allowed this powerful accelerator of improvements to be leveraged.

61 RADIATION PROTECTION AND DOSIMETRY↗

What's New in the NSRDB

The National Solar Radiation Database (NSRDB) provides solar resource data across the globe at a high temporal and spatial resolution. This data is primarily used in solar energy modeling. The NSRDB is updated annually for the United States and North, Central and South America and the data is currently available from 1998-2021. In 2022 the NSRDB was updated using the latest version of the underlying Physical Solar Model (PSM). This update includes improved surface albedo and gap-filling of cloud properties. The inclusion of these updates reduced the uncertainty in the data compared to previous versions of the NSRDB. The Himawari and Meteosat Indian Ocean Data Coverage (IODC) satellites were added to the Geostationary Operational Environmental Satellite (GOES) and made our coverage global. While standard data from the GOES continues to be served at an hourly 4km x 4km resolution, full resolution data has also been made available to the user. The NSRDB now contains over 200Tb of data with nearly 40Tb being added annually. We provide significant flexibility for data download depending on the amount of data required by the users. In this paper we provide an update on the current status on the NSRDB.

photovoltaic systems↗

Baseline Climate Variables for Earth System Modelling

The Baseline Climate Variables for Earth System Modelling (ESM-BCVs) are defined as a list of 135 variables which have high utility for the evaluation and exploitation of climate simulations. The list reflects the most frequently used elements of the Coupled Model Intercomparison Project Phase 6 (CMIP6) archive. Successive phases of CMIP have supported strong results in science and substantially influence international climate policy formulation. This paper responds to both interest in exploiting CMIP data standards in a broader range of climate modelling activities and a need to achieve greater clarity about the significance and intention of variables in the CMIP Data Request. As Earth system modelling archives grow in scale and complexity, there are emerging problems associated with weak standardisation at the variable collection level. That is, there are good standards covering how specific variables should be archived, but this paper fills a gap in the standardisation of which variables should be archived. The ESM-BCV list is intended as a resource for ESM intercomparison projects (MIPs) developing requests to enable greater consistency among MIPs and as a reference for modelling centres to enhance consistency within MIPs. Provisional planning for the CMIP7 Data Request exploits the ESM-BCVs as a core element. The baseline variable list includes 98 variables which have modest or minor data volume footprints and could be generated systematically when simulations are produced and archived for exploitation by the World Climate Research Programme (WCRP) community. A further 35 variables are classed as “high volume” and are only suitable for production when the resource implications are justified.

Juckes, Martin [University of Oxford (United Kingd↗

Characterization of actinide abundances and isotopic compositions by HR-ICP-MS. Part 2: Results from actinide doping studies

Previously, we presented actinide isotopic and elemental data from fallout melt glass that was measured using high-resolution inductively coupled plasma mass spectrometry (HR-ICP-MS). Direct comparison between these measurements and ‘gold standard’ data obtained by multi-collector ICP-MS showed broad overlap, indicating that HR-ICP-MS is a potentially valuable technique for producing actinide elemental and isotopic data on a relatively rapid timescale. To test the usefulness of this technique further, we doped varying amounts of uranium certified reference materials (CRMs) into a rhyolitic rock standard to establish the effects of uranium concentration and isotopic composition on the accuracy of uranium isotopic analyses by HR-ICP-MS. This also enabled us to quantify peak tailing effects from 238 U on the measurement of 239 Pu and 237 Np, which in turn allows us to constrain correction factors based on measured 237 Np/ 238 U and 239 Pu/ 238 U ratios. A second doping study involved the addition of a mixed actinide standard into three samples with different matrix compositions (termed soil, city, and seawater) to assess whether sample chemistry affects the accuracy and precision of these analyses. Results suggest that this is not the case. Systematic offsets were not observed in elemental or isotopic data derived from the three matrix samples. Results indicate that useful actinide isotopic data can be obtained from whole rock solutions by HR-ICP-MS. Our findings also have implications for solid sampling techniques such as laser ablation ICP-MS, which do not require sample dissolution.

38 RADIATION CHEMISTRY, RADIOCHEMISTRY, AND NUCLEA↗

TropiRoot 1.0: Database of tropical root characteristics across environments

Tropical ecosystems contain the world's largest biodiversity of vascular plants. Yet, our understanding of tropical functional diversity and its contribution to global diversity patterns is constrained by data availability. This discrepancy underscores an urgent need to bridge data gaps by incorporating comprehensive tropical root data into global datasets. Here, we provide a database of tropical root characteristics. This new database, TropiRoot 1.0, will be instrumental in evaluating an array of hypotheses pertaining to root functional ecology and plant biogeography, both within the tropics and relative to other global biomes. The data compilation was conducted by the TropiRoot Initiative, in partnership with the Fine-Root Ecology Database (FRED) and the Global Root Trait (GRooT) database, Colorado State University (CSU) and the Smithsonian Tropical Research Institute (STRI). Literature search and data extraction were conducted between 2020 and 2024. Literature was identified using Web of Science, Scopus, and complemented using the expert knowledge of members of TropiRoot. To provide broad environmental and geographical distributions, literature searches included root characteristics (traits) across global change drivers, natural gradients, and from different continents. We adopted FRED standardized data columns and streamlined the format to enhance accessibility for data extraction across various user groups. This optimized framework resulted in a smaller, yet comprehensive datasheet. To make the database compatible with other global root trait initiatives, column identification was standardized following the codes provided by FRED. These efforts culminated in data extracted from 104 new sources, resulting in more than 8000 rows of data (either species or community data). Most of the data in TropiRoot 1.0 include root characteristics such as root biomass, morphology, root dynamics, mass fraction, architecture, anatomy, physiology, and root chemistry. This initiative represents a 30% increase in the currently available data for tropical roots in FRED. TropiRoot 1.0 contains root characteristics from 25 different countries, where seven are located in Asia, six in South America, five in Central America and the Caribbean, four in Africa, two in North America, and 1 in Oceania. Due to the volume of data, when ancillary data were available, including soil data, these data were either extracted and included in the database or its availability was recorded in an additional column. Multiple contributors checked the entries for outliers during the collation process to ensure data quality. For text-based observations, we examined all cells to ensure that their content relates to their specific categories. For numerical observations, we ordered each numerical value from least to greatest and plotted the values, checking apparent outliers against the data in their respective sources and correcting or removing incorrect or impossible values. Some data (soil and aboveground) have different columns for the same variable presented in different units, including originally published units, but root characteristics data had units converted to match those reported in FRED. By filling a gap from global databases, TropiRoot 1.0 expands our knowledge of otherwise so far underrepresented regions and our ability to assess global trends. This advancement can be used to improve tropical forest representation in vegetation models. The data are freely available and should be cited when used.

FRED↗

Post-composing ontology terms for efficient phenotyping in plant breeding

Abstract Ontologies are widely used in databases to standardize data, improving data quality, integration, and ease of comparison. Within ontologies tailored to diverse use cases, post-composing user-defined terms reconciles the demands for standardization on the one hand and flexibility on the other. In many instances of Breedbase, a digital ecosystem for plant breeding designed for genomic selection, the goal is to capture phenotypic data using highly curated and rigorous crop ontologies, while adapting to the specific requirements of plant breeders to record data quickly and efficiently. For example, post-composing enables users to tailor ontology terms to suit specific and granular use cases such as repeated measurements on different plant parts and special sample preparation techniques. To achieve this, we have implemented a post-composing tool based on orthogonal ontologies providing users with the ability to introduce additional levels of phenotyping granularity tailored to unique experimental designs. Post-composed terms are designed to be reused by all breeding programs within a Breedbase instance but are not exported to the crop reference ontologies. Breedbase users can post-compose terms across various categories, such as plant anatomy, treatments, temporal events, and breeding cycles, and, as a result, generate highly specific terms for more accurate phenotyping.

Mathematical & Computational Biology↗

The Marine and Hydrokinetic ToolKit for Data Quality Control and Analysis: Preprint

The ability to handle data is critical at all stages of marine energy (ME) development. The marine hydrokinetic toolkit (MHKiT) is an open-source marine energy software, which includes modules for ingesting, applying quality control, processing, visualizing, and managing data. MHKiT-Python and MHKiT-MATLAB provide robust and verified functions that are needed by the ME community to standardize data processing. Calculations and visualizations adhere to International Electrotechnical Commission (IEC) technical specifications and other guidelines. A resource assessment of NDBC buoy 46050 near PACWAVE is performed using MHKiT and discusses comparisons to the resource assessment provided performed by Dunkel et al.

marine energy↗

The Marine and Hydrokinetic Toolkit (Mhkit) for Data Quality Control and Analysis

The ability to handle data is critical at all stages of marine energy (ME) development. The marine hydrokinetic toolkit (MHKiT) is an open-source marine energy software, which includes modules for ingesting, applying quality control, processing, visualizing, and managing data. MHKiT-Python and MHKiT-MATLAB provide robust and verified functions that are needed by the ME community to standardize data processing. Calculations and visualizations adhere to International Electrotechnical Commission (IEC) technical specifications and other guidelines. A resource assessment of NDBC buoy 46050 near PACWAVE is performed using MHKiT and discusses comparisons to the resource assessment provided performed by Dunkel et al.

marine energy↗

Analysis of overlapping count data

Counts of a specific characteristic were obtained within regions defined on an object that was manufactured in a proprietary setting. The count regions were altered during production and resulted in misaligned or overlapping count data. A closed-formula maximum likelihood estimator (MLE) of the new region means is derived using all of the available count data and an independent Poisson model. The MLE is shown to be preferable to estimators constructed using generalized linear models for the overlapping data setting. This closed-form estimator extends to over-dispersed overlapping count data as the quasi-MLE and also performs well with correlated overlapping count data. Standard errors for the estimator are approximated and are validated with a simulation study. Additionally, the methods are extended to overlapping multinomial data. Illustrative examples of the methods are provided throughout the paper and are reproducible with the supplemental R code. Additionally, proofs of the paper’s results are also included in the supplemental material.

97 MATHEMATICS AND COMPUTING↗

Re-evaluating the prompt fission neutron spectrum of spontaneously fissioning 252 Cf

The prompt fission neutron spectrum (PFNS) of spontaneously fissioning 252 Cf is a Neutron Data Standards observable. Nearly all fission spectra of actinides were measured relative to it, using efficiencies derived from it, or analyzed with simulations validated by it. The current Standards evaluation was published by W. Mannhart in 1987. It could not be updated because the evaluation input, experimental mean values and covariances, were lost. First, we attempt to reproduce it. However, Mannhart’s evaluation can only be reproduced within its one-σ uncertainties as some of its aspects (e.g., experimental covariances, rejected data points) remain unknown. Therefore, a new evaluation is presented: We revisit all existing experimental 252 Cf(sf) PFNS data, including those published after the release of the current Standards evaluation, and re-estimate associated covariances. The newly evaluated 252 Cf(sf) PFNS differs distinctly from Mannhart’s below 300 keV and extends it to lower and higher outgoing neutron energies (500 eV–25 MeV). The new evaluated uncertainties are larger from 3–9 MeV and smaller otherwise. Spectrum averaged cross sections of importance to the International Reactor Dosimetry and Fusion File community calculated with the new spectrum are close to those calculated with Mannhart’s evaluation and agree with experimental values well within their uncertainties.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Characterizing peak electricity demand for U.S. households: an assessment of end-use loads and demand factors

Understanding household peak electricity demand is critical to evaluate the technical need for electrical infrastructure upgrades. This study characterizes peak loads for existing and new equipment using metered data from a convenience sample of 11,940 U.S. dwellings from four sources, including 911 from two sources with end-use metering. After standardized data cleaning and labeling, we derived descriptive statistics for key metrics, such as maximum demand and demand factors, and developed predictive models relating 60- to 15-min demand for the National Electrical Code (NEC). Mean 15-min maximum demand was 9.7 kW (median 9.0 kW; IQR 7.0–11.5 kW, 95% CI 9.6–9.8 kW), indicating spare capacity in 98% of homes with hypothetical 100 A panels. Maximum demand increased with floor area and number of high-demand loads. Dwelling maximum demand was driven by higher-power, longer-duration heating appliances and vehicle charging, while most user-operated appliances contributed little. Demand factors are used to account for how most devices contribute less than their rated power to maximum demand. Existing load mean demand factors (28%; median 10%; IQR 0–58%; CI 28–29%) were higher than those for new loads (21%; median 7%; IQR 0–35%; CI 20–21%), because new loads changed the timing and magnitude of maximum demand. New high-demand loads had higher than average demand factors (40–60%). Whole dwelling demand factors support the NEC's 40% assumption, but they challenge its conservative 100% treatment of new HVAC. We propose a data-driven 50% demand factor for new equipment, which would align with metered data, improve affordability, and modernize electrical codes.

Appliances↗

Reactor Containment Passive Safety Analysis: Steam Condensation in Presence of Non-condensable Gas Scaled Experiment and Modeling

This study presents steam condensation scaled experiments and semi-empirical models in presence of nitrogen (N)—a noncondensable gas (NCG), simulating air in the reactor containment—to support water-cooled small modular reactors (SMRs) passive containment cooling system (PCCS) design and analysis. Previous experimental studies on PCCS are focused on fixed and smaller tube (mostly 2-in.) geometries and specific test condition variations, bringing challenges with geometric scaling and mismatching with SMR prototypic design. To address these challenges, this study presents steam condensation test dataset obtained from three scaled test sections of 1-, 2-, and 4-in.-diameter steam condensers with an annular/jacket cooling of 2-, 3-, and 6 in.-diameter tubes, respectively. Test data were collected for steam ranges from 58 to 63 kg/hr., and NCG flow of 4.4 to 13.3 kg/hr. Annular cooling water flow was varied to obtain required testing conditions of saturated steam inlet and fully condensed outlet. Axial temperature test data of bulk cooling water, steam and condensate were collected by thermocouples for three test sections and various steam-NCG mixing/testing conditions. A standard data reduction method was adopted—utilizing iterative and nodalized mass and heat transfer calculation—to estimate axial local heat fluxes, heat transfer coefficients (HTCs), condensation rates, film thickness, and Nusselt number. Based on the obtained dataset semi-empirical model results—a ratio of experimental and Nusselt’s theoretical HTC are presented. Such results and findings are supportive of developing scaled-up testing facility, to enable model validations and accelerate next generation of reactors development and deployment

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Presentation: Reactor Containment Passive Safety Analysis: Steam Condensation in Presence of Non-condensable Gas Scaled Experiment and Modeling

This study presents steam condensation scaled experiments and semi-empirical models in presence of nitrogen--a noncondensable gas (NCG), simulating air in the reactor containment--to support water-cooled small modular reactors (SMRs) passive containment cooling system (PCCS) design and analysis. Previous experimental studies on PCCS are focused on fixed and smaller tube (mostly 2-in.) geometries and specific test condition variations, bringing challenges with geometric scaling and mismatching with SMR prototypic design. To address these challenges, this study presents steam condensation test dataset obtained from three scaled test sections of 1-, 2-, and 4-in.-diameter steam condensers with an annular/jacket cooling of 2-, 3-, and 6 in.-diameter tubes, respectively. Test data were collected for steam ranges from 58 to 63 kg/hr., and NCG flow of 4.4 to 13.3 kg/hr. Annular cooling water flow was varied to obtain required testing conditions of saturated steam inlet and fully condensed outlet. Axial temperature test data of bulk cooling water, steam and condensate were collected by thermocouples for three test sections and various steam-NCG mixing/testing conditions. A standard data reduction method was adopted--utilizing iterative and nodalized mass and heat transfer calculation to estimate axial local heat fluxes, heat transfer coefficients (HTCs), condensation rates, film thickness, and Nusselt number. Based on the obtained dataset semi-empirical model results--a ratio of experimental and Nusselt's theoretical HTC are presented. Such results and findings are supportive of developing scaled-up testing facility, to enable model validations and accelerate next generation of reactors development and deployment.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Livewire Data Platform: File Standards Version 1.0

This technical document is a user guide to help users of the Livewire Data Platform understand the standards and requirements for storing and sharing data on the Livewire Data Platform.

97 MATHEMATICS AND COMPUTING↗

POWTEX visits POWGEN

The high-intensity time-of-flight (TOF) neutron diffractometer POWTEX for powder and texture analysis is currently being built prior to operation in the eastern guide hall of the research reactor FRM II at Garching close to Munich, Germany. Because of the world-wide 3 He crisis in 2009, the authors promptly initiated the development of 3 He-free detector alternatives that are tailor-made for the requirements of large-area diffractometers. Herein is reported the 2017 enterprise to operate one mounting unit of the final POWTEX detector on the neutron powder diffractometer POWGEN at the Spallation Neutron Source located at Oak Ridge National Laboratory, USA. As a result, presented here are the first angular- and wavelength-dependent data from the POWTEX detector, unfortunately damaged by a 50 g shock but still operating, as well as the efforts made both to characterize the transport damage and to successfully recalibrate the voxel positions in order to yield nonetheless reliable measurements. Also described is the current data reduction process using the PowderReduceP2D algorithm implemented in Mantid [Arnold et al. (2014). Nucl. Instrum. Methods Phys. Res. A , 764 , 156–166]. The final part of the data treatment chain, namely a novel multi-dimensional refinement using a modified version of the GSAS-II software suite [Toby & Von Dreele (2013). J. Appl. Cryst. 46 , 544–549], is compared with a standard data treatment of the same event data conventionally reduced as TOF diffraction patterns and refined with the unmodified version of GSAS-II . This involves both determining the instrumental resolution parameters using POWGEN's powdered diamond standard sample and the refinement of a friendly-user sample, BaZn(NCN) 2 . Although each structural parameter on its own looks similar upon comparing the conventional (1D) and multi-dimensional (2D) treatments, also in terms of precision, a closer view shows small but possibly significant differences. For example, the somewhat suspicious proximity of the a and b lattice parameters of BaZn(NCN) 2 crystallizing in Pbca as resulting from the 1D refinement (0.008 Å) is five times less pronounced in the 2D refinement (0.038 Å). Similar features are found when comparing bond lengths and bond angles, e.g. the two N—C—N units are less differently bent in the 1D results (173 and 175°) than in the 2D results (167 and 173°). The results are of importance not only for POWTEX but also for other neutron TOF diffractometers with large-area detectors, like POWGEN at the SNS or the future DREAM beamline at the European Spallation Source.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗