Engineering PapersSearch

SEARCH · Engineering Papers

Results for “content based networking”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

50 records · Page 3

An ML-based terrestrial data fusion and augmentation framework to enable advanced understanding of the terrestrial carbon and water interactions

Soil moisture is essential to the terrestrial carbon and water cycles and land–atmosphere interactions. There are various types of soil moisture data, and each type has the distinct spatiotemporal strengths and limitations, depending on the diverse applications and retrieval methodologies of different data types (Li et al., in review; The PNNL-82151 FY23 Report). However, the limitations of different soil moisture data in terms of accuracy and spatiotemporal coverage hinder our ability to further understand the soil moisture dynamics across scales. To have a gap free soil moisture data product with a fine spatiotemporal coverage and vertical profiles, we train extreme gradient boosting (XGBoost) models by using (1) in-situ soil moisture measurements from the International Soil Moisture Network (ISMN), (2) soil moisture from the ECMWF reanalysis (ERA) at the 9 km and sub-daily spatiotemporal resolution, (3) the Daymet meteorological fields, and (4) data products that characterize surface conditions, including soil texture, organic content, topography, vegetation type, and rooting depth. We use the trained XGBoost models that have consistent performance across seven soil layers, i.e., 0–5 cm, 5–10 cm, 10–20 cm, 20–40 cm, 40–60 cm, 60–100 cm, and 100–200 cm, and the gridded model predictors to generate a soil moisture data at the 1 km and daily spatiotemporal resolution for the Continental United States (CONUS) from 2001–2020. This dataset can be broadly used for Earth system model benchmark, monitoring extreme weathers, making informed decisions regarding agriculture, water resource management, climate change mitigation, and ecosystem preservation.

58 GEOSCIENCES

Machine learning for seismic low-frequency extrapolation

The cycle-skipping problem that plagues full waveform inversion (FWI) can be at least partially mitigated if low frequencies (which encode the kinematics of wave propagation in seismic data) are recorded. However, seismic sources and receivers are band-limited, so seismic data does not generally include signals down to 0 Hz. To improve our ability to solve the seismic inverse problem, one can synthesize this missing low-frequency (LF) content from the recorded high-frequency (HF) data using machine learning (ML) models. Deep learning models such as convolutional neural networks (CNNs) demonstrate impressive ability to perform low frequency extrapolation. However, such models require powerful hardware (GPU machines) and careful training. We assess the extrapolation capabilities of three different ML models that do not require GPU machines, namely, random forest, Gaussian process regression and gradient boosting, on both synthetic and real data. Experimental results on two synthetic data sets (generated from a low velocity lens embedded in a homogeneous medium, and the Marmousi model) demonstrate that FWI applied to the extrapolated data consistently improves inversion accuracy relative to FWI applied to the original data sets that do not contain low frequencies. Application of low-frequency extrapolation to real data from the Northwest Shelf of Australia demonstrates that tree-based ML models such as gradient boosting can outperform CNNs in terms of both accuracy and computational cost on non-GPU architectures.

58 GEOSCIENCES

Laboratory time series moisture manipulative experiment from sediment across the contiguous US: time series aerobic respiration and geochemistry (v2)

This dataset supports a broader study examining the effects of wetting and drying on hyporheic zone respiration across the contiguous United States (CONUS). The dataset provides data generated from a laboratory moisture manipulation experiment. The contents include time series aerobic respiration and moisture; dissolved oxygen; sediment geochemistry data; and field metadata (including qualitative information on instream and river corridor characteristics). Samples were collected as part of the WHONDRS CONUS-Scale Model-Sample Study (CM). This study was designed following ICON (integrated, coordinated, open, and networked) principles to facilitate a model-experiment (ModEx) iteration approach, leveraging crowdsourced sampling across the CONUS. The data package associated with the CM study is available at https://data.ess-dive.lbl.gov/view/doi:10.15485/1923689. CM sampling began in April 2022 and ended in October 2023. This study uses subsamples from a subset of CM samples collected between June 2022 and June 2023. The original field samples were labeled as CM_###. Subsequent subsamples for this study were labeled as EC_###. The labels from the field samples and the EC subsamples can be mapped directly based on the digits following the prefix and underscore (i.e., EC_001 is a subsample from CM_001). See the critical details section below for more details on sample naming. This data package was originally published in August 2024. It was updated in February 2026 (v2; new and modified files). See the change history section in the readme for more details. For details on how to navigate this data package, see this infographic from the River Corridor SFA https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. This dataset is comprised of one folder of raw Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS) data and one main data folder containing (1) file-level metadata; (2) data dictionary; (3) field metadata; (4) readme; (5) field protocol; and a (6) a subfolder with sediment sample data from the incubation experiment. The sample data subfolder contains (1) dissolved organic carbon (DOC, measured as non-purgeable organic carbon, NPOC); (2) total nitrogen (TN); (3) adenosine triphosphate (ATP); (4) percent carbon and nitrogen; (5) effect size; (6) iron (II); (7) gravimetric moisture; (8) respiration rates and raw dissolved oxygen values; (9) specific conductance; (10) pH; (11) temperature; (12) a summary containing median values of each data type for each treatment (wet and dry); (13) methods codes; (14) FTICR-MS methods; and (15) a subfolder of 9.4 Tesla FTICR-MS data. This folder contains three subfolders, one containing the sediment .xml data files, one containing the sediment CoreMS output files, the other containing instructions and scripts for processing the files in CoreMS (https://github.com/EMSL-Computing/CoreMS). All files are .csv, .pdf, .R, .ref, or .xml.

54 ENVIRONMENTAL SCIENCES

Identifying preferential flow from soil moisture time series: Review of methodologies

Abstract Identifying and quantifying preferential flow (PF) through soil—the rapid movement of water through spatially distinct pathways in the subsurface—is vital to understanding how the hydrologic cycle responds to climate, land cover, and anthropogenic changes. In recent decades, methods have been developed that use measured soil moisture time series to identify PF. Because they allow for continuous monitoring and are relatively easy to implement, these methods have become an important tool for recognizing when, where, and under what conditions PF occurs. The methods seek to identify a pattern or quantification that indicates the occurrence of PF. Most commonly, the chosen signature is either (1) a nonsequential response to infiltrated water, in which soil moisture responses do not occur in order of shallowest to deepest, or (2) a velocity criterion, in which newly infiltrated water is detected at depth earlier than is possible by nonpreferential flow processes. Alternative signatures have also been developed that have certain advantages but are less commonly utilized. Choosing among these possible signatures requires attention to their pertinent characteristics, including susceptibility to errors, possible bias toward false negatives or false positives, reliance on subjective judgments, and possible requirements for additional types of data. We review 77 studies that have applied such methods to highlight important information for readers who want to identify PF from soil moisture data and to inform those who aim to develop new methods or improve existing ones. Core Ideas Soil moisture data can be used to identify the occurrence of preferential flow (PF) and its initiating conditions. Various data‐analysis methods to identify PF differ in susceptibility to error, bias, and subjectivity. These methods can utilize vast amounts of data from soil moisture monitoring networks to develop understanding of when, where, and under what conditions PF occurs. Newly developed methods may lead to better accuracy and reliability, and reduce the need for subjective judgments. Plain Language Summary Preferential flow through soil occurs when a large amount of water is suddenly available, as during an intense storm. This type of flow moves rapidly through the soil in distinct narrow pathways rather than moving evenly throughout the body of soil, with major consequences for groundwater resources, ecosystems, spreading of contaminants, and other vital concerns. Methods of detecting preferential flow have been developed that utilize measurements of soil water content made by sensors installed at various depths. This measurement technology has been widely implemented, many locations now having datasets years in length, and various methods have been developed for using these to identify preferential flow. The various methods are based on different features in the soil moisture records and vary in their advantages and shortcomings. In this review, we explain and evaluate these methods, highlighting important information for their implementation to identify preferential flow from soil moisture data and for efforts to develop new methods or improve existing ones.

Nimmo, John R

Hyporheic zone, river, and groundwater metagenome resolved genomes and rpS3 genes in East River Watershed, Colorado USA Summer 2020, 2021

Here we present metagenome assembled genomes (MAGs) for the bacterial and archaeal communities from water filter collected across 8 locations along the East River Watershed, CO, and 1 nearby groundwater well. The purpose was to look for connectivity and similarities across the network and to see the impact of the groundwater. As a part of Lawrence Berkeley National Laboratory (LBNL) Watershed Science Focus Area (SFA), we assessed community composition and strain similarities between the sites and we also compared it to previous metagenomic studies within the watershed looking at floodplain (Matheus Carnevali et al. 2021) and hillslope (Lavy et al. 2019) microbiomes. Here we present metagenome assembled genomes (MAGs) for the bacterial and archaeal communities from filters across 8 locations during August 2020 and July 2021. This resulted in 32 samples. The groundwater sample was sequenced at UC Berkley's QB3. The other 31 samples were sequenced at University of Maryland. Metagenomes were assembled using four autobinners and the best bins were selected using dasTool. The genomes were dereplicated at 95% with dRep and the subset of winning genomes were manually curated based on visual inspection of taxonomic profile, GC content, coverage, and a set of 51 bacterial single copy genes (BSCG), and 38 archaeal signal copy genes (ASCG). The dataset includes a zip file of 311 genomes (HZ_River_SW_MAGS_Dereplicated_95.zip). The dataset additionally includes a zipped file of ribosomal protein small subunit 3 (rpS3) proteins from the hyporheic zone and river data (rpS3_Proteins_HZ_River.zip), a metadata file used to register associated samples with IGSNs (International Generic Sample Numbers) (samples.csv), a location metadata file (locations.csv). This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

DNA

Deriving cloud droplet number concentration from surface-based remote sensors with an emphasis on lidar measurements

Abstract. Given the importance of constraining cloud droplet number concentrations (Nd) in low-level clouds, we explore two methods for retrieving Nd from surface-based remote sensing that emphasize the information content in lidar measurements. Because Nd is the zeroth moment of the droplet size distribution (DSD), and all remote sensing approaches respond to DSD moments that are at least 2 orders of magnitude greater than the zeroth moment, deriving Nd from remote sensing measurements has significant uncertainty. At minimum, such algorithms require the extrapolation of information from two other measurements that respond to different moments of the DSD. Lidar, for instance, is sensitive to the second moment (cross-sectional area) of the DSD, while other measures from microwave sensors respond to higher-order moments. We develop methods using a simple lidar forward model that demonstrates that the depth to the maximum in lidar-attenuated backscatter (Rmax⁡) is strongly sensitive to Nd when some measure of the liquid water content vertical profile is given or assumed. Knowledge of Rmax⁡ to within 5 m can constrain Nd to within several tens of percent. However, operational lidar networks provide vertical resolutions of > 15 m, making a direct calculation of Nd from Rmax⁡ very uncertain. Therefore, we develop a Bayesian optimal estimation algorithm that brings additional information to the inversion such as lidar-derived extinction and radar reflectivity near the cloud top. This statistical approach provides reasonable characterizations of Nd and effective radius (re) to within approximately a factor of 2 and 30 %, respectively. By comparing surface-derived cloud properties with MODIS satellite and aircraft data collected during the MARCUS and CAPRICORN II campaigns, we demonstrate the utility of the methodology.

54 ENVIRONMENTAL SCIENCES

Validation of an Erythema-Weighted UV Model Using Broadband Solar Irradiance Measurements From Eleven U.S. Sites: Preprint

Erythema-weighted UV solar irradiance (UV-E) has a potential impact on human health if the recommended maximum exposure times are exceeded. In spite of this, it is not measured at most sites that measure Global Horizontal Irradiance (GHI). However, since UV-E is highly correlated with GHI and total ozone content it can be estimated from this information with sufficient accuracy to assisst in public health recommendations. The Power Model (PM) provides a simple method to estimate the erythema-weighted UV irradiance (UV-E) from measured GHI, total ozone column and air mass. In this work, the performance of the PM method is assessed using high-quality data from 11 sites in the continental U.S. (part of SURFRAD and SOLRAD networks) and total ozone estimates publicly available from the MERRA-2 re-analysis database. A three year period (2021- 2023) at 1-minute frequency is considered. The results (for time aggregations of 5 and 60 minutes) show a high Pearson's correlations (> 0.99), consistently positive mean bias deviations (below 13%) and dispersions in the 8-19% range at all sites. Relative values are expressed in terms of the corresponding measurement mean. The performance indicators remain consistent across time resolutions (5 or 60 minutes), suggesting that the model's performance is robust and not significantly affected by short-term variability (which is captured by GHI). Spatial patterns reveal higher biases and RMSD in northern and eastern locations. These values represent a significant improvement over widely used satellite-based global UV-E estimates and open the possibility of using the PM with satellite-based GHI estimates for operational UV-E mapping over the contiguous U.S. territory.

14 SOLAR ENERGY

Measure this, not that: Optimizing the cost and model-based information content of measurements

Model-based design of experiments (MBDoE) is a powerful framework for selecting and calibrating science-based mathematical models from data. Here, this work extends popular MBDoE workflows by proposing a convex mixed integer (non)linear programming (MINLP) to optimize the selection of measurements. The solver MindtPy is modified to support calculating the D-optimality objective and its gradient via an external package, scipy, using the grey-box module in Pyomo. The new approach is demonstrated in two case studies: estimating highly correlated kinetics from a batch reactor and estimating transport parameters in a large-scale rotary packed bed for CO 2 capture. Both case studies show how examining the Pareto optimal trade-offs between information content measured by A- and D-optimality versus measurement budget offers practical guidance for selecting measurements for scientific experiments.

97 MATHEMATICS AND COMPUTING

CHESS 2025: Leaf Area Index (LAI) for meadow, shrub, tree, and understory vegetation

This dataset contains Leaf Area Index (LAI) measurements made as part of the Colorado Headwaters Ecological Spectroscopy Study (CHESS) during June and July of 2025. Data were collected in the Upper Gunnison Basin, Colorado, across three study domains: the Upper East River (CRBU), Almont Triangle (ALMO), and the Upper Taylor Basin (UPTA). Field observations of LAI were collected within 72 hours of airborne data collection by the National Ecological Observatory Network’s Aerial Observation Platform (NEON AOP). The NEON AOP collected waveform LiDAR (Light Detection and Ranging) and imaging spectrometer data in 426 spectral bands from the visible to shortwave infrared. LAI measurements were collected using the LICOR LAI-2200C Plant Canopy Analyzer following protocols outlined in the instrument manual (LI-COR 2019). Sampling targeted four distinct vegetation types: meadows, shrubs, trees, and aspen forest understory. We have archived data separately by site type because different field methods were used for each. At meadow sites, measurements were made at the four corners of 1m x 1m plots, with the instrument moving inward toward the center of the plot. At shrub sites, we measured the canopies of individual shrubs. At tree sites, we made measurements within a 10m x 10m subplot centered around a focal tree, with 30 observations taken on a regular grid. At aspen understory sites, we measured overstory trees following the tree protocol and understory herbaceous vegetation following the meadow protocol. All measurements included above-canopy (A) and below-canopy (B) readings, with specific protocols for scattering correction measurements in direct-sun conditions. Data were processed using the R package `rlai` (Worsham 2025). This package includes functions to calculate LAI, gap fraction, apparent clumping factor (Ω), scattering correction, and other canopy metrics. Package contents: Full file descriptions appear in ‘flmd.csv’. Files named according to the convention ‘lai_*_summary_data_cleaned.csv’ contain summary values of LAI, apparent clumping factor (Ωapp), and scattering correction factors for each site. These are the analysis-ready products that most data users will work with. Files named ‘lai_*_metadata_cleaned.csv’ contain additional site-level observations made during field collection. We have also archived intermediate and supplementary data for users who wish to check our processing approach or apply alternative methods. ‘raw_lai_2200C.zip’ contains the raw files as read from the LI-COR instrument, with no processing applied, in TXT format. The zip archive contains subdirectories by site type, which are further subdivided by sampling area. Filenames correspond to the sampling site number. ‘intermediate_results.zip’ contains detailed output from the processing routines, in JSON format. The zip archive contains subdirectories by site type; filenames correspond to the sampling site number. ‘scattering_correction_logs.zip’ contains logfiles from the implementation of Kobayashi et al.'s (2013) scattering correction algorithm. The logfiles report values of several parameters at each iteration of the algorithm, as the model converges toward a stable solution. They are intended for users who want to verify scattering correction performance. The zip archive contains subdirectories by site type; filenames correspond to the sampling site number. ‘spot_checks.csv’ reports LAI and other values for a small number of files processed with LI-COR FV2200 software (LI-COR 2013) using the same control parameters as in our R-based approach. Additional metadata are provided in a data dictionary describing column names and definitions (dd.csv), and in a file-level metadata file (flmd.csv). All zip files can be expanded with common archive utilities. TXT, CSV, and JSON files can be ingested into R or Python computing environments or read in common text editor utilities. Geospatial information: Geospatial data for mapping measurement site locations are in the files CHESS_polygons_lai_UTM.geojson, CHESS_polygons_shrub_UTM.geojson, and CHESS_polygons_meadow_UTM.geojson in the companion geospatial package for the 2025 CHESS campaign, ‘CHESS 2025: Location data for field observations and sampling’ (Henderson et al., 2026). CHESS Project Description: The Colorado Headwaters Ecological Spectroscopy Study (CHESS) comprised a multi-week airborne remote sensing and field observation campaign in the Upper Gunnison Basin, Colorado, conducted in June and July of 2025. Airborne remote sensing was conducted by the National Ecological Observatory Network Airborne Observation Platform (NEON AOP), concurrent with a field campaign run by the Rocky Mountain Biological Laboratory (RMBL), the Lawrence Berkeley National Laboratory (LBNL) and SLAC National Accelerator Laboratory Watershed Function Science Focus Area (SFA), and NASA-JPL (Jet Propulsion Laboratory) Earth Surface Mineral Dust Source Investigation (EMIT) program. Between June 10 and July 18, 2025, the NEON AOP flight team collected high-resolution aerial imaging spectroscopy and Light Detection and Ranging (LiDAR) data over three domains: the Upper East River (CRBU), Almont Triangle (ALMO), and the Upper Taylor Basin (UPTA). In coordination with the flights, a field campaign acquired ground-truth observations, including observations of vegetation composition, foliar traits, forest demography, and subsurface properties in 18 core sampling areas within the domains. Additional surface water observations were taken at over 380 point locations. All CHESS campaign datasets can be found within the CHESS ESS-DIVE data portal: https://data.ess-dive.lbl.gov/portals/chess. Funding Acknowledgement: Field and remote-sensing data acquisition was performed under a grant from the National Aeronautics and Space Administration (80NSSC24K1005). This work was also supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231. * Todorov and Worsham are co–first authors.

2018 NEON and 2025 CHESS Campaigns

Dominant Controls on Preferential Flow and Their Implications for Future Soil Water Fluxes

Abstract Soil water flow, particularly preferential flow (PF), is a critical control on hydrological and biogeochemical processes, including groundwater recharge, contaminant transport, and carbon cycling. However, it remains challenging to predict PF occurrence across large environmental gradients. Here, we developed a deep learning (DL) model to estimate event‐scale soil water flow velocity and the probability of PF occurrence using high‐frequency soil moisture and precipitation data from 33 sites across the National Ecological Observatory Network. The model demonstrated high skill in predicting the binary occurrence of PF (91% F1‐score; 85% accuracy) but the performance was limited in predicting soil water velocity ( R 2 = 0.31). We found that precipitation characteristics (duration, volume, and intensity) were the most important predictors for soil water velocity. Among the non‐precipitation event variables, sand content showed relatively high predictive skill, though differences among non‐event climate variables were generally modest. Lower sand content was associated with increased predicted soil water velocity, a finding that highlights the role of soil structure in producing more non‐uniform flow, which contrasts with traditional uniform flow models. Projecting a reduced DL model under both moderate and high‐emissions future climate scenarios (2060–2099 Representative Concentration Pathways 4.5 and 8.5), we found ∼7.3% increase under RCP4.5 and ∼15% under RCP8.5 of soil water velocities compared to the historical simulation, while modeled likelihood of PF changed little. These findings suggest climate change is not making PF more frequent, but it is making existing PF pathways more efficient with important consequences for associated nutrient and contaminant transport under climate change. Plain Language Summary Water movement in soil is critical for water quality. While often modeled as a uniform flow process, in reality water moves rapidly through cracks and burrows in what is called “preferential flow” (PF), which limits natural filtration and can transport pollutants. We developed a deep learning model, trained on data from 33 U.S. sites, to predict when and how fast this PF occurs based on precipitation, soil, and climate data. The model showed that precipitation characteristics (duration, intensity, volume) were the most important predictors of PF. Lower soil sand content/higher clay content was associated with faster water flow, likely due to clay soils forming aggregates and cracks that water moves through rather than infiltrating uniformly. Further analyses based on climate projections suggest that the speed at which PF occurs will become more rapid under future climate scenarios compared to historical simulation. This highlights the need to represent PF in soil water models when assessing future water quality. Key Points The effect of precipitation peak intensity on soil water velocities declined with increasing precipitation intensity Antecedent soil moisture failed to predict preferential flow (PF), contrasting the high predictive power of sand content Climate predictions suggest that soil water velocities through PF paths will increase ∼15% by 2099

Li, Bonan

Tailoring Carbide Dispersed Steels: A Path to Increased Strength and Hydrogen Tolerance

The use of transition metal carbides is reported for use as a hydrogen trapping mechanism for ferritic and austenitic steel materials. The program combined computational modeling and simulations to guide experiments towards candidate metal carbide traps, both for interfacial and interior trapping. It was found that interfacial trapping is less effective than interior trapping, with the group IVB transition metal carbides being the most effect internal traps with a loss of carbon. The sub-stoichiometric rocksalt structure accommodate the hydrogen atoms in its octahedral interstices. Using percolation theory, carbon loss of approximately 25% or more was sufficient to ensure an interconnected network of vacancies for such trapping from the surface to the internal sites within the carbide. Using this as a guide, the program developed a means to provide a uniform dispersion of ZrC nanoparticles with either Fe or 304L micron-scale powders which was then consolidated by direct current sintering. Electrolytic hydrogen diffusivity studies confirmed the reduction of hydrogen diffusion in the matrix with increasing ZrC content, which was a linear response over the sample range studied (0.01 to 1.0 wt.%). The consolidated material was micro-tensile tested in either a non-hydrogen or hydrogen charge condition and compared to a control with no carbides. Additions up to 0.05 wt.% ZrC increased the yield strength with no loss in ductility in either the non-hydrogen or hydrogen tested condition. ZrC concentrations above this amount further increased the yield strength at the expense of ductility. While these samples had a lower absolute ductility value prior to failure, the relative change in ductility between the non-hydrogen and hydrogen charge states was less for the carbides than that of the control. Metal-rich ZrC nanoparticles were fabricated through a conformal coating process yielding ZrC0.66 particles that were then incorporated into a metal matrix. Notch fatigue testing in a hydrogen environment was conducted where the number of cycles to failure was found to be less in the control than that of the carbide addition. However, the spread in experimental data and the number of samples tested limits a conclusive outcome based on defects noticed in the gauge section of all the powder processed samples. The collective outcomes of this report provide further insight into the mechanisms by which carbides act as hydrogen traps; a means to process such carbides through powder metallurgy; and their associated mechanical performance in either a non-hydrogen or hydrogen-charged condition.

08 HYDROGEN

Herbaceous Feedstock 2022 State of Technology Report

The U.S. Department of Energy promotes production of advanced liquid transportation fuels from lignocellulosic biomass by funding fundamental and applied research that advances the state of technology (SOT). As part of its involvement in this mission, Idaho National Laboratory completes an annual SOT report for nth-plant and 1st-plant herbaceous biomass feedstock logistics. The purpose of the SOT is to provide the status of feedstock supply system technology development for herbaceous biomass to biofuels relative to technical targets and cost goals from specific design cases, based on data and experimental results. Although conventional feedstock supply systems form the backbone of the emerging biofuels industry, they have limitations that restrict widespread implementation on a national scale. To meet the demands of the future industry, the feedstock supply system must shift from the conventional system to what has been termed “advanced” supply systems. In advanced designs, a distributed network of aggregation and processing centers, termed “depots,” are employed near the points of biomass production (i.e., the field or forest) to reduce feedstock variability and produce feedstocks of a uniform format, moving toward biomass commoditization. The 2022 Herbaceous SOT is part of a vision of achieving an implemented advanced feedstock supply system, which produces a stable, tradable commodity at the decentralized distributed depot. It utilizes feedstock fractionation by incorporating technologies that can separate the biomass into its anatomical fractions (leaves, husks, stems and cobs) to reduce impurities and produce fractions that satisfy downstream quality considerations. By using a series of air classification steps, this strategy can reduce the extrinsic ash in corn stover and produce enriched tissue fractions that can be blended to a conversion specification or converted individually in optimized biochemical conversion campaigns. Additionally, a majority of the leaves (which do not meet the quality specification) are separated out early and can be supplied to alternate markets. The 2022 Herbaceous SOT incorporates an advanced biomass fractionation and processing system to produce pellets enriched tissues from three-pass corn stover. The resulting enriched pellets are delivered to the biorefinery individually where they can be blended to a specification or converted in campaigns where the conditions are optimized for each tissue. Unused fractions can be sent to a a midstream market or to a different conversion process that is better suited to their properties to offset the cost of the delivered feedstock. The main benefits from the proposed system can be summarized as: (1) $6.86/dry ton (2016$) lower cost for the air classification due to elimination of the requirement to discard the high ash lights fraction; (2) $1.56/dry ton lower delivered cost by selling the unsuitable leaf fraction into the feed market as a midstream co-product (assuming a selling price that is 11% higher than their cost of production); (3) 0.98% increase in carbohydrate content (from 60.16% to 61.14%); and (4) 0.97% decrease in ash content (from 6.00% to 5.03%) compared to the 2021 Herbaceous SOT. Overall, the 2022 nth-plant Herbaceous SOT predicts a modeled delivered feedstock cost of $78.64/dry ton (2016$) if it is assumed that the enriched leaf fraction is sold at its production cost; this is a slight increase of $0.43/dry ton increase from the 2021 Herbaceous SOT nth-Supply case cost. The increased cost derived from a $0.38/dry ton increase in transportation and handling cost to procure more biomass (to replace the enriched leaf fraction that was not delivered to the biorefinery. The total preprocessing cost was $0.27/dry ton higher than the 2021 result because of updates to energy consumption, purchasing price and dry matter loss data for the rotary shear ($3.00/dry ton increase) and the pelleting mill ($4.52/dry ton increase). The data utilized were generated in pilot-scale tests in the Biomass Feedstock National User Facility (BFNUF) at INL and at Forest Concepts, including tests for rotary shear and pelleting of the air classified fractions. A greenhouse gas emissions analysis was performed by Argonne National Laboratory using the most up to date version of the Greenhouse Gases, Regulated Emissions, and Energy use in Transportation model (GREET®). The analysis showed an increase of 17.34 kg CO2e/dry ton from the 2021 SOT (67.71 kg CO2e/ton in the 2021 Herbaceous SOT to 85.05 kg CO2e/ton in the 2022 Herbaceous SOT). The net increase is primarily attributed to increased energy consumption in pelleting mill.

09 BIOMASS FUELS

Custom surface reflectance, shade mask, and equivalent water thickness maps for the Colorado Headwaters Ecological Spectroscopy Study (2025)

This dataset contains land surface reflectance estimates and additional derived products generated from NEON Imaging Spectrometer (NIS) data collected in the Upper Gunnison river basin during June and July of 2025. Data was collected over three domains: the Upper East River (CRBU), Almont Triangle (ALMO), and the Upper Taylor Basin (UPTA). These products were derived from radiance and LiDAR data collected by the NEON Airborne Observation Platform (AOP) campaign funded by the Colorado Headwaters Ecological Spectroscopy Study (CHESS) (doi:10.15485/3017965). Products include per-pixel surface reflectance (rfl) and reflectance uncertainty (rfl_unc), observational data (obs), canopy equivalent water thickness (ewt), and shade masks. Atmospheric correction was performed per flightline using the ISOFIT (Imaging Spectrometer Optimal FITting) optimal estimation framework to estimate surface reflectance and the associated per-band reflectance uncertainty. Reflectance retrievals achieved a mean absolute error of 1.5% across diverse validation surfaces (see validation report.pdf). Equivalent water thickness was calculated from surface reflectance using the Beer–Lambert absorption of liquid water. Shade masks were generated based on the geometry between the sun angle, ground surface, and sensor at the time of flight. Data products are provided per-flightline and as mosaics for each domain. Flightline data products are provided as ENVI-formatted binary files (rfl, rfl_unc, ewt) and GeoTIFFs (shade). Reflectance and uncertainty mosaics are provided as tiled NetCDFs, while all other mosaicked products are provided as cloud-optimized GeoTIFFs. These formats are supported by common geospatial software (e.g., QGIS, ArcGIS, ENVI) and programmatic libraries in Python (e.g., rasterio, xarray, spectral, netCDF4) and R (e.g., terra, ncdf4). Processing workflows were designed to be equivalent to those used to generate the 2018 CHESS campaign airborne imaging spectroscopy data products (doi:10.15485/3013527). All outputs were co-registered to a common spatial grid to support time series analyses. CHESS Project Description: The Colorado Headwaters Ecological Spectroscopy Study (CHESS) comprised a multi-week airborne remote sensing and field observation campaign in the Upper Gunnison Basin, Colorado, conducted in June and July of 2025. Airborne remote sensing was conducted by the National Ecological Observatory Network Airborne Observation Platform (NEON AOP), concurrent with a field campaign run by the Rocky Mountain Biological Laboratory (RMBL), the Lawrence Berkeley National Laboratory (LBNL) and SLAC National Accelerator Laboratory Watershed Function Science Focus Area (SFA), and NASA-JPL (Jet Propulsion Laboratory) Earth Surface Mineral Dust Source Investigation (EMIT) program. Between June 10 and July 18, 2025, the NEON AOP flight team collected high-resolution aerial imaging spectroscopy and Light Detection and Ranging (LiDAR) data over three domains: the Upper East River (CRBU), Almont Triangle (ALMO), and the Upper Taylor Basin (UPTA). In coordination with the flights, a field campaign acquired ground-truth observations, including observations of vegetation composition, foliar traits, forest demography, and subsurface properties in 18 core sampling areas within the domains. Additional surface water observations were taken at over 380 point locations. All CHESS campaign datasets can be found within the CHESS ESS-DIVE data portal: https://data.ess-dive.lbl.gov/portals/chess. Funding Acknowledgment: Data acquisition was performed under a grant from the National Aeronautics and Space Administration (80NSSC24K1005). Computational research was carried out at the Jet Propulsion Laboratory, California Institute of Technology, under a contract with the National Aeronautics and Space Administration (80NM0018D0004) and was funded by EMIT Extended Mission Phase E Science.

2018 NEON and 2025 CHESS Campaigns

HydraGNN_Predictive_GFM_2026 - Ensemble of predictive graph foundation models for atomistic materials modeling

This release contains data and parameters of HydraGNN-based graph foundation models trained as a result of the work published in the pre-print "Exascale Multi-Task Graph Foundation Models for Imbalanced, Multi-Fidelity Atomistic Data" by M. Lupo Pasini et al. (https://arxiv.org/abs/2604.15380). We jointly train on 16 open first-principles datasets (544+ million structures covering 85+ elements) using a multi-task architecture with per-dataset heads and a scalable ADIOS2/DDStore data pipeline. On Frontier, we execute six large-scale DeepHyper hyperparameter optimization campaigns in FP64 and promote the top-performing message-passing models to sustained 2,048-node training, yielding a PaiNN-based lead model. The version of HydraGNN used to generate the outputs provided in this release is HydraGNN v5.0 (https://github.com/ORNL/HydraGNN/releases/tag/v5.0) The list of datasets used for the training of the graph foundation model is the following: 1) Alexandria [1] 2) ANI1x [2] 3) MPTrj [3] 4) Open Catalyst 2020 (OC20) [4] 5) Open Catalyst 2022 (OC22) [5] 6) Open Catalyst 2025 (OC25) [6] 7) Open Direct ir Capture 2023 (ODAC23) [7] 8) Open Materials 2024 (OMat24) [8] 9) Open Molecules 2025 (OMol25) [9] 10) OMol25-neutral (subset of OMol25 that contains only molecules with zero total charge) 11) OMol25-non-neutral (subset of OMol25 that contains only molecules with non-zero total charge) 12) Open Polymers 2026 (OPoly2026) [10] 13) Nabla2DFT [11] 14) QCML [12] 15) QM7X [reference 13] 16) transition1x [14] Dataset references: [1] J. Schmidt et al., “A dataset of 175k stable and metastable materials calculated with the PBEsol and SCAN functionals,” Scientific Data, vol. 9, p. 64, 2022. [2] J. S. Smith et al., “The ANI-1ccx and ANI-1x data sets, coupled-cluster and density functional theory properties for molecules,” Scientific Data, vol. 7, p. 134, 2020. [Online]. Available: https: //www.nature.com/articles/s41597-020-0473-z [3] A. Jain et al., “Commentary: The Materials Project: A materials genome approach to accelerating materials innovation,” APL Materials, vol. 1, no. 1, p. 011002, 07 2013. [Online]. Available: https://doi.org/10.1063/1.4812323 [4] L. Chanussot et al., “Open catalyst 2020 (oc20) dataset and community challenges,” ACS Catalysis, vol. 11, no. 10, pp. 6059–6072, 2021. [Online]. Available: https://doi.org/10.1021/acscatal.0c04525 [5] K. Tran et al., “Open catalyst 2022 (oc22) dataset and challenges for oxidation electrocatalysts,” ACS Catalysis, vol. 13, no. 5, pp. 3066–3084, 2023. [Online]. Available: https://doi.org/10.1021/acscatal.2c05426 [6] S. J. Sahoo et al., “The open catalyst 2025 (oc25) dataset and models for solid-liquid interfaces,” arXiv preprint arXiv:2509.17862, 2025. [Online]. Available: https://arxiv.org/abs/2509.17862 [7] A. Sriram et al., “The open DAC 2023 dataset and challenges for sorbent discovery in direct air capture,” ACS Central Science, vol. 10, no. 5, pp. 923–941, 2024. [8] L. Barroso-Luque et al., “Open materials 2024 (omat24) inorganic materials dataset and models,” 2024. [Online]. Available: https://arxiv.org/abs/2410.12771 [9] D. S. Levine et al., “The open molecules 2025 (OMol25) dataset, evaluations, and models,” 2025. [Online]. Available: https://arxiv.org/abs/2505.08762 [10] D. S. Levine et al., The open polymers 2026 (OPoly26) dataset and evaluations,” arXiv preprint arXiv:2512.23117, 2025. [Online]. Available: https://arxiv.org/abs/2512.23117 [11] K. Khrabrov et al., “Nabla2dft: A universal quantum chemistry dataset of drug-like molecules and a benchmark for neural network potentials,” in NeurIPS 2024 Datasets and Benchmarks Track, 2024. [Online]. Available: https://openreview.net/forum?id=ElUrNM9U8c [12] S. Ganscha et al., “The QCML dataset, quantum chemistry reference data from 33.5M DFT and 14.7B semi-empirical calculations,” Scientific Data, vol. 12, p. 406, 2025. [13] J. Hoja et al., “QM7-X, a comprehensive dataset of quantum-mechanical properties spanning the chemical space of small organic molecules,” Scientific Data, vol. 8, p. 43, 2021. [Online]. Available: https://www.nature.com/articles/s41597-021-00812-2 [14] M. Schreiner et al., “Transition1x - a dataset for building generalizable reactive machine learning potentials,” Scientific Data, vol. 9, p. 779, 2022. The folder "datasets_ADIOS2_format" contains the set of pre-processed datasets in Adaptable I/O System (ADIOS) format (https://www.exascaleproject.org/research-project/adios/) that have been used for the development and training of GFMs in this work. The "datasets_ADIOS2_format" directory contains 2 sub-directories, one for the version "v1" of the datasets and one for the version "v2" of the datasets. The version "v1" of the datasets provides values of the total energy as they are extracted from the original data as it was released by the respective institutions. The version "v2" of the datasets provides values of the energy that have been realigned. The realignment was performed by training a linear regression model that predicts the total energy as a function of the chemical composition of the atomistic structure, and then subtract such prediction from the original value of the total energy. Both folders "v1" and "v2" contain 16 sub-directories, each corresponding to an ADIOS2-formatted dataset The folder "DeepHyper-results" contains the configurational files and model's parameters for all the 186 HPO trials that were successfully completed by the scalable hyperparameter optimization (HPO) runs on Frontier. The content of the folder "DeepHyper-results" I structured as follows: 1) task-list.txt: list of mpnn name, jobid, and deephyper task id 2) gfm_${MPNN}_${JOBID}_0.${TASKID}: run directory with checkpoint files 3) gfm_${MPNN}: deephyper summary directory (*.csv) for each specific MPNN type 4) deephyper-experiment-${JOBID}: output and error logs for each job The file "deephyper-sorted.csv" contains the details of each HydraGNN model built and tested by HPO, obtained by merging the (*.csv) filed from each HPO run executed. Out of all the HPO trials, we selected 10 to continue the training of the respective HydraGNN models. Due to limited computational budget available in the LRN070 allocation we could not complete the training till convergence for all these 10 selected models. The folder "models" contains multiple sub-folders, one per each HydraGNN model trained. Each model sub-folder contains the parameters of each HydraGNN model, with multiple checkpoint-restarts. The list of sub-folders are as follows: 1) multidataset_hpo-BEST1-fp64 2) multidataset_hpo-BEST2-fp64 3) multidataset_hpo-BEST3-fp64 4) multidataset_hpo-BEST4-fp64 5) multidataset_hpo-BEST5-fp64 6) multidataset_hpo-BEST6-fp64 7) multidataset_hpo-BEST7-fp64 8) multidataset_hpo-BEST8-fp64 9) multidataset_hpo-BEST9-fp64 10) multidataset_hpo-BEST10-fp64 Within each one of these folders, additional auxiliary log files are provided with descriptions about how the training proceeded. The lead PaiNN-model is contained inside "multidataset_hpo-BEST6-fp64". The file "mlp_branch_weights" contains the parameters of the multi-layer perceptron (MLP) used to reconcile the predictions of the 16 output decoding heads of the HydragNN architectures. The MLP takes in input the chemical composition of the atomistic structure and predicts averaging weights to linearly mix the predictions of each output decoding head toward consolidating them into a single one. The folder "1.1billion-structure-inference" contains 1.1 billion atomistic structures randomly generated. Each structures is associated with energy and forces predicted with the lead-PaiNN model combined with the MLP model for reconciliation of the multi-branch predictions generated by the 16 output decoding heads. The folder "1.1billion-structure-inference" contains 9,300 (*.tar.gz) subdirectories, one per Frontier compute node used to execute the inference at exascale. Once uncompressed, each (*.tar.gz) subdirectory contains an ADIOS2 (*.bp) file container, where each atomistic structure is stored as a PyTorch-Geometric Data object. The file "export_dataset_environment_variables.sh" contains the environment variables that need to be set before running the HydraGNN code to reproduce the results provided in this dataset release. The code that can be used to load the ADIOS2 files, load HydraGNN models, and run inference is available at: https://github.com/ORNL/HydraGNN/releases/tag/v5.0

36 MATERIALS SCIENCE