Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “evaluation datasets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Validation of the Community Land Model Version 5 over the Contiguous United States (CONUS) using in situ and remote sensing data sets

The Community Land Model (CLM) is an effective tool to simulate the biophysical and biogeochemical processes and their interactions with the atmosphere. Although CLM Version 5 (CLM5) constitutes various updates in these processes, its performance in simulating energy, water and carbon cycles over the Contiguous United States (CONUS) at scales which land surface changes and hydrometeorological and hydroclimatological applications are more locally relevant is yet to be assessed. In this study, we conducted three simulations at 0.125? during 1979-2018 over the CONUS using different configurations of CLM, namely CLM5-biogeochemistry (CLM5BGC), CLM4.5BGC, and CLM5-satellite phenology (CLM5SP). We validated and compared their simulations against multiple remote-sensed and in-situ datasets. Overall, the parametric and structural updates (e.g., carbon cost for nitrogen uptake, variable soil thickness, dry surface layer) in CLM5 improve its ability in capturing terrestrial biogeochemical dynamics. The low evapotranspiration in CLM5BGC is associated with biases in simulating vegetation phenological characteristics rather than soil water limitations. The mismatch between CLM5BGC-simulated peak leaf area index and reference data can be attributed to CLM5BGC's inability in simulating phenology of trees and grasses. The differences between CLM-simulated irrigation and reference estimates can be attributed to differences between processes represented in models and in reality, and uncertainties in input and validation datasets. Evaluation against observations at small catchments suggest that hydrologic parameters needed to be calibrated to improve simulations of runoff, especially subsurface runoff. Additional efforts are needed to incorporate spatially-distributed plant phenology and physiology parameters and regional-specific agricultural management practices (e.g., planting, harvest).

Cheng, Yanyan↗

Unveiling four decades of intensifying precipitation from tropical cyclones using satellite measurements

Increases in precipitation rates and volumes from tropical cyclones (TCs) caused by anthropogenic warming are predicted by climate modeling studies and have been identified in several high intensity storms occurring over the last half decade. However, it has been difficult to detect historical trends in TC precipitation at time scales long enough to overcome natural climate variability because of limitations in existing precipitation observations. We introduce an experimental global high-resolution climate data record of precipitation produced using infrared satellite imagery and corrected at the monthly scale by a gauge-derived product that shows generally good performance during two hurricane case studies but estimates higher mean precipitation rates in the tropics than the evaluation datasets. General increases in mean and extreme rainfall rates during the study period of 1980–2019 are identified, culminating in a 12–18%/40-year increase in global rainfall rates. Overall, all basins have experienced intensification in precipitation rates. Increases in rainfall rates have boosted the mean precipitation volume of global TCs by 7–15% over 40 years, with the starkest rises seen in the North Atlantic, South Indian, and South Pacific basins (maximum 59–64% over 40 years). In terms of inland rainfall totals, year-by-year trends are generally positive due to increasing TC frequency, slower decay over land, and more intense rainfall, with an alarming increase of 81–85% seen from the strongest global TCs. As the global trend in precipitation rates follows expectations from warming sea surface temperatures (11.1%/°C), we hypothesize that the observed trends could be a result of anthropogenic warming creating greater concentrations of water vapor in the atmosphere, though retrospective studies of TC dynamics over the period are needed to confirm.

54 ENVIRONMENTAL SCIENCES↗

Validation of the Community Land Model Version 5 over the Contiguous United States (CONUS) using in-situ and remote sensing datasets

The Community Land Model (CLM) is an effective tool to simulate the biophysical and biogeochemical processes and their interactions with the atmosphere. Although CLM Version 5 (CLM5) constitutes various updates in these processes, its performance in simulating energy, water and carbon cycles over the Contiguous United States (CONUS) at scales which land surface changes and hydrometeorological and hydroclimatological applications are more locally relevant is yet to be assessed. In this study, we conducted three simulations at 0.125? during 1979-2018 over the CONUS using different configurations of CLM, namely CLM5-biogeochemistry (CLM5BGC), CLM4.5BGC, and CLM5-satellite phenology (CLM5SP). We validated and compared their simulations against multiple remote-sensed and in-situ datasets. Overall, the parametric and structural updates (e.g., carbon cost for nitrogen uptake, variable soil thickness, dry surface layer) in CLM5 improve its ability in capturing terrestrial biogeochemical dynamics. The low evapotranspiration in CLM5BGC is associated with biases in simulating vegetation phenological characteristics rather than soil water limitations. The mismatch between CLM5BGC-simulated peak leaf area index and reference data can be attributed to CLM5BGC's inability in simulating phenology of trees and grasses. The differences between CLM-simulated irrigation and reference estimates can be attributed to differences between processes represented in models and in reality, and uncertainties in input and validation datasets. Evaluation against observations at small catchments suggest that hydrologic parameters needed to be calibrated to improve simulations of runoff, especially subsurface runoff. Additional efforts are needed to incorporate spatially-distributed plant phenology and physiology parameters and regional-specific agricultural management practices (e.g., planting, harvest).

54 ENVIRONMENTAL SCIENCES↗

Livewire: A Model Platform for Data Quality Assessment and AI Readiness Across DOE Missions

High-quality, well-governed data is essential for accelerating discovery and achieving operational excellence across DOE and national laboratory missions. The Livewire Data Platform is a DOE-supported platform that offers automated assessments of data quality, standardization, provenance, and Artificial Intelligence (AI) readiness. It allows researchers and data practitioners to systematically and easily evaluate datasets against established governance criteria and prepare them for advanced analytics. Livewire addresses critical challenges in DOE's data ecosystem with integrated capabilities for metadata validation, provenance tracking, and schema alignment. This platform's automated workflows assist users in identifying data quality gaps, enhancing interoperability between datasets collected from various stakeholders, and ensuring compliance with DOE data standards, all while reducing manual curation efforts. Additionally, we will discuss its AI readiness framework, which is being developed to prepare datasets for training models, developing advanced analytic tools, and machine learning applications. Using some of the more than one hundred tabular datasets on Livewire, processed with this open-source methodology, we will demonstrate how Livewire can serve as a model for scalable, standards-driven data management. This approach provides a pathway to leverage existing and future datasets within the DOE, boosting innovation and efficiency across national laboratories.

33 - ADVANCED PROPULSION SYSTEMS↗

EI_MS_ML

The unambiguous identification of compounds from their electron ionization mass (EI-MS) spectra remains a significant unsolved problem in the field of metabolomics and analytical chemistry as a whole. Typically EI-MS spectra are compared using various mathematical operations that convert the spectral similarity or differences into a distance-like metric that roughly approximates the similarity of any two spectra. A commonly used metric for this is the cosine similarity metric which has values close to one for very similar spectra and a value of zero for very dissimilar spectra; however, no metric is perfect. Due to the prevalence of structurally-similar compounds such as isomers and the prevalence of certain fragmentation patterns across structurally-dissimilar compounds, the unambiguous assignment of EI-MS spectra compounds remains difficult. Frequently, querying an observed EI-MS spectrum against a large database such as the NIST17 library yields multiple possible assignments requiring the end user to distinguish between multiple high scoring hits, or multiple low scoring hits while keeping in mind that the correct hit may not be in the database at all. Although techniques such as orthogonal information from techniques such as chromatography can greatly aid in unambiguous assignment, this also requires more complicated experimental designs and access to more complicated analytical instrumentation. Substructures can be trivially detected and represented as strings using a previously published technique called node coloring from a known chemical structure. However, for experimentally-derived EI-MS spectra this information must be derived from the spectra itself (i.e., because we do not know what compound it represents). To achieve this, the software uses techniques from the field of machine learning and a large training dataset of EI-MS spectra corresponding to known structures annotated with substructure strings, to build models that can predict the presence of a given chemical substructure from an EI-MS spectrum directly.If these predictions are of high-quality (i.e., are unlikely to be false positives), the presence of one or more predicted substructures can be used to constrain the number of possible hits for a query spectrum. Mathematically, this restriction could be expressed in many forms, but the most straight-forward implementation is to weight the cosine similarity of a query spectrum and a plausible database match with a Tanimoto-like coefficient based on the ratio of the number of substructures predicted to the number of substructures present in the potential database hit. Determining which combination of models best reduces assignment ambiguity will be achieved using a combination of manual curation and optimization techniques such as genetic algorithms. This software will perform all the steps necessary to construct said models from a training dataset and evaluate them using a holdout dataset. Various statistical analyses can be performed to determine if this approach does decrease assignment ambiguity. For example, if this approach works, on average, the rank-order of the correct assignment for the holdout set of EI-MS spectra should decrease and the weighted cosine similarities for most of the possible matches in the database should be better than the unweighted cosine similarities. Furthermore, this same pipeline can be used on real experimental data to generate less ambiguous assignments.

Mitchell, Joshua↗

Cold-Season Precipitation Sensitivity to Microphysical Parameterizations: Hydrologic Evaluations Leveraging Snow Lidar Datasets

Abstract Cloud microphysical processes are an important facet of atmospheric modeling, as they can control the initiation and rates of snowfall. Thus, parameterizations of these processes have important implications for modeling seasonal snow accumulation. We conduct experiments with the Weather Research and Forecasting (WRF V4.3.3) Model using three different microphysics parameterizations, including a sophisticated new scheme (ISHMAEL). Simulations are conducted for two cold seasons (2018 and 2019) centered on the Colorado Rockies’ ∼750-km 2 East River watershed. Precipitation efficiencies are quantified using a drying-ratio mass budget approach and point evaluations are performed against three NRCS SNOTEL stations. Precipitation and meteorological outputs from each are used to force a land surface model (Noah-MP) so that peak snow accumulation can be compared against airborne snow lidar products. We find that microphysical parameterization choice alone has a modest impact on total precipitation on the order of ±3% watershed-wide, and as high as 15% for certain regions, similar to other studies comparing the same parameterizations. Precipitation biases evaluated against SNOTEL are 15% ± 13%. WRF Noah-MP configurations produced snow water equivalents with good correlations with airborne lidar products at a 1-km spatial resolution: Pearson’s r values of 0.9, RMSEs between 8 and 17 cm, and percent biases of 3%–15%. Noah-MP with precipitation from the PRISM geostatistical precipitation product leads to a peak SWE underestimation of 32% in both years examined, and a weaker spatial correlation than the WRF configurations. We fall short of identifying a clearly superior microphysical parameterization but conclude that snow lidar is a valuable nontraditional indicator of model performance.

54 ENVIRONMENTAL SCIENCES↗

Dataset documenting rock core evaluation from Brady’s Hot Springs well BCH-03

Nineteen short sections of core were taken from the BCH-03 borehole at Brady’s Hot Springs, Churchill County, Nevada, between 2,945 ft and the total depth of 4,885 ft. This well is in a geothermal field. Data collected includes thin section photomicrographs, Medical CT tiff stacks, microCT tiff stacks, XRD, porosity (from thin sections, He porosimeter, microCT segmentation) and P/S wave velocities.

Brady geothermal field↗

Detection of Outliers in LiDAR Data Acquired by Multiple Platforms over Sorghum and Maize

High-resolution point cloud data acquired with a laser scanner from any platform contain random noise and outliers. Therefore, outlier detection in LiDAR data is often necessary prior to analysis. Applications in agriculture are particularly challenging, as there is typically no prior knowledge of the statistical distribution of points, plant complexity, and local point densities, which are crop-dependent. The goals of this study were first to investigate approaches to minimize the impact of outliers on LiDAR acquired over agricultural row crops, and specifically for sorghum and maize breeding experiments, by an unmanned aerial vehicle (UAV) and a wheel-based ground platform; second, to evaluate the impact of existing outliers in the datasets on leaf area index (LAI) prediction using LiDAR data. Two methods were investigated to detect and remove the outliers from the plant datasets. The first was based on surface fitting to noisy point cloud data via normal and curvature estimation in a local neighborhood. The second utilized the PointCleanNet deep learning framework. Both methods were applied to individual plants and field-based datasets. To evaluate the method, an F-score was calculated for synthetic data in the controlled conditions, and LAI, the variable being predicted, was computed both before and after outlier removal for both scenarios. Results indicate that the deep learning method for outlier detection is more robust than the geometric approach to changes in point densities, level of noise, and shapes. The prediction of LAI was also improved for the wheel-based vehicle data based on the coefficient of determination (R2) and the root mean squared error (RMSE) of the residuals before and after the removal of outliers.

36 MATERIALS SCIENCE↗

Simultaneous cross-evaluation of heterogeneous E. coli datasets via mechanistic simulation

The extensive heterogeneity of biological data poses challenges to analysis and interpretation. Construction of a large-scale mechanistic model of Escherichia coli enabled us to integrate and cross-evaluate a massive, heterogeneous dataset based on measurements reported by various groups over decades. We identified inconsistencies with functional consequences across the data, including that the total output of the ribosomes and RNA polymerases described by data are not sufficient for a cell to reproduce measured doubling times, that measured metabolic parameters are neither fully compatible with each other nor with overall growth, and that essential proteins are absent during the cell cycle—and the cell is robust to this absence. Finally, considering these data as a whole leads to successful predictions of new experimental outcomes, in this case protein half-lives.

59 BASIC BIOLOGICAL SCIENCES↗

Systematic Evaluation of Atmospheric Forcing, Surface Datasets, and Mesh Effects on Kilometer-Scale Land Surface and River Modeling

Earth system models are advancing toward kilometer-scale resolution to capture local climate impacts and extremes. High-resolution land and river modeling depends on multiple factors, including mesh, surface datasets, and atmospheric forcing, but their relative effects at kilometer scales remain unquantified. We evaluated five Energy Exascale Earth System Model land and river configurations over the Mid-Atlantic region using two mesh (1/8° structured versus variable-resolution unstructured mesh), two surface datasets (default versus newly developed), and three atmospheric forcings (NLDAS2, MSWX, GSWP). Evaluation against satellite, reanalysis, and in situ benchmarks across water, energy, and carbon cycles quantifies how these factors affect model performance. Forcing selection produces the largest bias reductions (12-99% across variables), followed by surface datasets (7-75%) and mesh (up to 21%). Forcing effects vary by variable, with MSWX reducing biases for snow water equivalent, evapotranspiration, albedo, temperature, and gross primary productivity, GSWP for snow cover and runoff, and NLDAS for soil moisture and streamflow. The use of newly developed surface datasets improves gross primary productivity (58% bias reduction) and evapotranspiration but increase soil moisture and albedo biases due to current modeling limitations. Variable-resolution unstructured mesh improves the simulation of small-basin streamflow through better capturing drainage networks, though mesh minimally affects other land variables. These findings provide important guidance for high-resolution modeling development and actionable science.

Land and River modeling↗

Evaluation of LLM-Generated Kokkos Code Using Compile-Time and Run-Time Testing

Due to the growing use of large language models (LLMs) by developers and researchers, it has become essential to reliably evaluate their ability to generate code that uses specialized libraries. We explore the use of compile-time and run-time evaluation of LLM-generated Kokkos code through extending the methods used by OpenAI with the HumanEval dataset. Our evaluation framework is based on the first 40 prompts from the Kokkos138 dataset. We start by discussing two different forms of LLM prompting, using entirely plain English or providing pseudocode for added context. These two methods are used to generate Kokkos code with the Llama-3.1-8B-Instruct and CodeQwen1.5-7B-Chat models. We found that both forms of prompting led to high failure rates and difficulties with reliably parsing LLM-generated code, while prompts with pseudocode for context generally led to improved results on more complicated tests.

97 MATHEMATICS AND COMPUTING↗

The Evaluation of Machine Learning Techniques for Isotope Identification Contextualized by Training and Testing Spectral Similarity

Precise gamma-ray spectral analysis is crucial in high-stakes applications, such as nuclear security. Research efforts toward implementing machine learning (ML) approaches for accurate analysis are limited by the resemblance of the training data to the testing scenarios. The underlying spectral shape of synthetic data may not perfectly reflect measured configurations, and measurement campaigns may be limited by resource constraints. Consequently, ML algorithms for isotope identification must maintain accurate classification performance under domain shifts between the training and testing data. To this end, four different classifiers (Ridge, Random Forest, Extreme Gradient Boosting, and Multilayer Perceptron) were trained on the same dataset and evaluated on twelve other datasets with varying standoff distances, shielding, and background configurations. A tailored statistical approach was introduced to quantify the similarity between the training and testing configurations, which was then related to the predictive performance. Wilcoxon signed-rank tests revealed that the OVR-wrapped XGB significantly outperformed the other algorithms, with confidence levels of 99.0% or above for the 133Ba, 60Co, 137Cs, and 152Eu sources. The findings from this work are significant as they outline techniques to promote the development of robust ML-based approaches for isotope identification.

domain adaptation↗

Evaluation of the PV Resource Dataset in Central and North America: Preprint

A recent effort from the National Renewable Energy Laboratory (NREL) has led to a new radiative transfer model, Fast All-sky Radiation Model for Solar applications with Narrowband Irradiances on Tilted surfaces (FARMS-NIT), to efficiently compute plane-of-array (POA) irradiance. This model has been implemented by the National Solar Radiation Data Base (NSRDB) to provide photovoltaic (PV) resource over Central and North America in both narrow- and broad wavelength bands. This study conducts a comprehensive evaluation of the PV resource dataset using surface-based observations in various climate zones. The results demonstrate the PV resource has an excellent agreement with the cloudy-sky observations from thermopiles and reference cells.

41 EE - Solar Energy Technologies Office (EE-4S)↗

Deep neural network uncertainty quantification for LArTPC reconstruction

We evaluate uncertainty quantification (UQ) methods for deep learning applied to liquid argon time projection chamber (LArTPC) physics analysis tasks. As deep learning applications enter widespread usage among physics data analysis, neural networks with reliable estimates of prediction uncertainty and robust performance against overconfidence and out-of-distribution (OOD) samples are critical for their full deployment in analyzing experimental data. While numerous UQ methods have been tested on simple datasets, performance evaluations for more complex tasks and datasets are scarce. Here we assess the application of selected deep learning UQ methods on the task of particle classification using the PiLArNet monte carlo 3D LArTPC point cloud dataset. We observe that UQ methods not only allow for better rejection of prediction mistakes and OOD detection, but also generally achieve higher overall accuracy across different task settings. We assess the precision of uncertainty quantification using different evaluation metrics, such as distributional separation of prediction entropy across correctly and incorrectly identified samples, receiver operating characteristic curves (ROCs), and expected calibration error from observed empirical accuracy. We conclude that ensembling methods can obtain well calibrated classification probabilities and generally perform better than other existing methods in deep learning UQ literature.

47 OTHER INSTRUMENTATION↗

Physics-Based Method for Generating Fully Synthetic IV Curve Training Datasets for Machine Learning Classification of PV Failures

Classification machine learning models require high-quality labeled datasets for training. Among the most useful datasets for photovoltaic array fault detection and diagnosis are module or string current-voltage (IV) curves. Unfortunately, such datasets are rarely collected due to the cost of high fidelity monitoring, and the data that is available is generally not ideal, often consisting of unbalanced classes, noisy data due to environmental conditions, and few samples. In this paper, we propose an alternate approach that utilizes physics-based simulations of string-level IV curves as a fully synthetic training corpus that is independent of the test dataset. In our example, the training corpus consists of baseline (no fault), partial soiling, and cell crack system modes. The training corpus is used to train a 1D convolutional neural network (CNN) for failure classification. The approach is validated by comparing the model’s ability to classify failures detected on a real, measured IV curve testing corpus obtained from laboratory and field experiments. Results obtained using a fully synthetic training dataset achieve identical accuracy to those obtained with use of a measured training dataset. When evaluating the measured data’s test split, a 100% accuracy was found both when using simulations or measured data as the training corpus. When evaluating all of the measured data, a 96% accuracy was found when using a fully synthetic training dataset. The use of physics-based modeling results as a training corpus for failure detection and classification has many advantages for implementation as each PV system is configured differently, and it would be nearly impossible to train using labeled measured data.

Hopwood, Michael W. (ORCID:0000000161901767)↗

A high resolution, gridded product for vapor pressure deficit using Daymet

Vapor pressure deficit (VPD) is a critical variable in assessing drought conditions and evaluating plant water stress. Gridded products of global and regional VPD are not freely available from satellite remote sensing, model reanalysis, or ground observation datasets. We present two versions of the first gridded VPD product for the Continental US and parts of Northern Mexico and Southern Canada (CONUS+) at a 1 km spatial resolution and daily time step. We derived VPD from Daymet maximum daily temperature and average daily vapor pressure and scale the estimates based on (1) climate determined by the Köppen-Geiger classifications and (2) land cover determined by the International Geosphere-Biosphere Programme. Ground-based VPD data from 253 AmeriFlux sites representing different climate and land cover classifications were used to improve the Daymet-derived VPD estimates for every pixel in the CONUS+ grid to produce the final datasets. We evaluated the Daymet-derived VPD against independent observations and reanalysis data. The CONUS+ VPD datasets will aid in investigating disturbances including drought and wildfire, and informing land management strategies.

54 ENVIRONMENTAL SCIENCES↗