Engineering PapersSearch

SEARCH · Engineering Papers

Results for “missing values”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Data from a throughfall exclusion experiment: Fine root dynamics, morphology, chemistry, and AMF colonization across four lowland Panamanian forests

Fine roots regulate forest nutrient, carbon, and water cycling, yet their variation within and among tropical forests remains under-characterized. We quantified root productivity, disappearance, and stocks to 1 m using minirhizotron imaging, and we measured morphology, elemental composition [root carbon (C), root nitrogen (N), root phosphorus (P)], and arbuscular mycorrhizal fungi (AMF) colonization to 20 cm using ingrowth cores and sequential coring. Sampling took place in four distinct lowland Panamanian forests (32 plots; 8 per forest) from 2018 through 2022 under control and throughfall-exclusion (drought) treatments in the Panama Rainforest Changes with Experimental Drying (PARCHED) experiment.The dataset is presented as an Excel workbook with six tabs. The first tab is the data dictionary. Tab S1 contains ingrowth-core production and mortality, morphology and soil moisture. Tab S2 contains sequential-coring standing stocks with associated morphology and soil moisture. Tab S3 contains minirhizotron row data records to 1 m depth, including per-frame root length and diameter, normalized length metrics, and session timing. Tab S4 contains AMF colonization. Tab S5 contains fine-root chemistry at 0–10 cm, reporting %P, %C, %N, and C:N for samples collected via ingrowth cores and sequential-coring standing stocks. CSV mirrors for each tab are provided, and a KML file supplies coordinates for all 32 plots.Key variables span live and dead fine-root biomass (and coarse fractions where applicable), specific root length (SRL) and area (SRA), diameter, root tissue density (RTD), soil moisture, AMF colonization, root %N, %C, %P, and C:N, along with minirhizotron root length and diameter. Depth, season, treatment, and plot/site identifiers are included to support cross-tab integration and analysis from 0–100 cm (minirhizotron) and 0–20 cm (cores).Units are reported in-column and missing values are coded as NA. No special software is required to open or use the files (Excel, CSV, and KML compatible).

54 ENVIRONMENTAL SCIENCES

A Photochemical Phosphorus-Hydrogen-Oxygen Network for Hydrogen-dominated Exoplanet Atmospheres

Due to the detection of phosphine (PH 3 ) in the solar system gas giants Jupiter and Saturn, PH 3 has long been suggested to be detectable in exosolar substellar atmospheres too. However, to date, direct detection of phosphine has proven to be elusive in exoplanet atmosphere surveys. We construct an updated phosphorus-hydrogen-oxygen (PHO) photochemical network suitable for the simulation of gas giant hydrogen-dominated atmospheres. Using this network, we examine PHO photochemistry in hot Jupiter and warm Neptune exoplanet atmospheres at solar and enriched metallicities. Our results show for HD 189733b-like hot Jupiters that HOPO, PO, and P 2 are typically the dominant P carriers at pressures important for transit and emission spectra, rather than PH 3 . For GJ1214b-like warm Neptune atmospheres our results suggest that at solar metallicity PH 3 is dominant in the absence of photochemistry, but is generally not in high abundance for all other chemical environments. At 10 and 100 times solar, small oxygenated phosphorus molecules such as HOPO and PO dominate for both thermochemical and photochemical simulations. The network is able to reproduce well the observed PH 3 abundances on Jupiter and Saturn. Despite progress in improving the accuracy of the PHO network, large portions of the reaction rate data remain with approximate, uncertain, or missing values, which could change the conclusions of the current study significantly. Improving understanding of the kinetics of phosphorus-bearing chemical reactions will be a key undertaking for astronomers aiming to detect phosphine and other phosphorus species in both rocky and gaseous exoplanetary atmospheres in the near future.

atmospheric composition

Observed and Imputed Volumetric Soil Water Content Timeseries for the New Mexico Elevation Gradient

Reliable soil water content (SWC) data are essential for understanding dryland ecosystem dynamics, but high-frequency SWC sensors often fail, creating gaps in critical datasets. To address this, we developed a Bayesian mixture model that imputes missing SWC using both linear interpolation and an ecosystem water balance model (SOILWAT2), tested across six AmeriFlux eddy covariance tower sites in the New Mexico Elevation Gradient, demonstrating its effectiveness in reconstructing SWC patterns while providing insights into the factors driving SWC variability. Daily volumetric soil water content (SWC) data are provided as csv-formatted spreadsheets for the six AmeriFlux sites (US-Seg, US-Ses, US-Wjs, US-Mpi, US-Vcp, and US-Vcs). For each site there is an observed SWC file (site_SWC_gapfill.csv) and a file that contains imputed SWC (imputed_SWC_site.csv). The observed SWC files contain temperature corrected sensor values, tower precipitation data, as well as outputs from SOILWAT2 simulations that were used to impute SWC. The imputed files contain the original observed SWC values and the imputed missing SWC values. When SWC was missing from the original data, the missing value was imputed based on the Bayesian imputation mixture model. The posterior mean of all imputed values is reported as "mean_X". When the observed SWC was NOT missing, mean_X = observed SWC value (original data). The standard deviation, 2.5th percentile and the 97.5th percentile for the imputed values are also reported in the imputed files. There are readme text files for each file type explaining the contents of each column.

54 ENVIRONMENTAL SCIENCES

Bayesian Gaussian process inference for neutron spin echo measurement

Neutron spin echo (NSE) spectroscopy provides unique access to microscopic dynamics, but its application is often constrained by low neutron flux, long acquisition times, and significant noise. Here, we present a Bayesian inference approach based on Gaussian process regression (GPR) to reconstruct high-quality spin echo signals from sparse and noisy data by exploiting correlations in reciprocal space. Benchmarks on synthetic datasets and validation with experimental NSE measurements of dendrimers show that GPR suppresses noise, interpolates missing intensity values, and accommodates irregular observations. The method improves accuracy, shortens acquisition times, and enables high-throughput and real-time studies. Beyond NSE, the framework is broadly applicable to other low signal-to-noise ratio scattering techniques, thereby extending the scope of neutron spectroscopy.

Tung, Chi-Huan [Oak Ridge National Laboratory (ORN

CryoTEN: efficiently enhancing cryo-EM density maps using transformers

Abstract Motivation Cryogenic electron microscopy (cryo-EM) is a core experimental technique used to determine the structure of macromolecules such as proteins. However, the effectiveness of cryo-EM is often hindered by the noise and missing density values in cryo-EM density maps caused by experimental conditions such as low contrast and conformational heterogeneity. Although various global and local map-sharpening techniques are widely employed to improve cryo-EM density maps, it is still challenging to efficiently improve their quality for building better protein structures from them. Results In this study, we introduce CryoTEN—a 3D UNETR++ style transformer to improve cryo-EM maps effectively. CryoTEN is trained using a diverse set of 1295 cryo-EM maps as inputs and their corresponding simulated maps generated from known protein structures as targets. An independent test set containing 150 maps is used to evaluate CryoTEN, and the results demonstrate that it can robustly enhance the quality of cryo-EM density maps. In addition, automatic de novo protein structure modeling shows that protein structures built from the density maps processed by CryoTEN have substantially better quality than those built from the original maps. Compared to the existing state-of-the-art deep learning methods for enhancing cryo-EM density maps, CryoTEN ranks second in improving the quality of density maps, while running >10 times faster and requiring much less GPU memory than them. Availability and implementation The source code and data are freely available at https://github.com/jianlin-cheng/cryoten.

Biochemistry & Molecular Biology

SetGo: Metadata Readiness for Scientific AI Datasets

Scientific datasets intended for AI use require both computational readiness for model training and metadata readiness for discovery, sharing, and reuse. The Readiness Engine for Data Integration (REDI) addresses computational readiness, but no corresponding tool evaluates whether a dataset’s metadata are sufficiently complete, governed, and standards-compliant for publication and agent-based consumption. Existing FAIR assessors operate only on published repository records, and no single system covers FAIR compliance, licensing, provenance, governance, reproducibility, and catalog readiness together. We present SetGo, an open-source Python toolkit that assesses and repairs metadata readiness across these six dimensions before a dataset is published or archived. Applied to four scientific corpora, SetGo surfaces deficiencies that general-purpose tools do not detect: ERA5 climate metadata scores 4% on ACDD 1.3 compliance; materials datasets fail OPTIMADE species-definition requirements; and PDB-derived proteomics data carries licensing terms incompatible with standard SPDX identifiers. Guided enrichment raises overall FAIR scores from 52–57% to 81–91%, and a single setgo publish command pushes to Hugging Face Hub, CKAN, or OpenMetadata with ML Commons Croissant 1.0 metadata sidecars. To support interactive and automated workflows, SetGo integrates with coding agents powered by large language models (LLMs) through a /setgo skill that enables natural-language execution of the full assess–enrich–publish loop, with user involvement limited to supplying missing metadata values.

Wilkinson, Sean [ORNL] (ORCID:0000000214437479)

Marginal Soils Index Analysis & Geospatial Data

This data package contains output files associated with Mongird et al. (in prep) organized into four dataset directories. Each dataset is described in more detail below. 1. Marginal Soils Index Analysis Description: This folder contains a csv file with land needs and availability by state, power generating technology type, and scenario in 2050 when suitable siting areas are additionally constrained to areas with increasing levels of soil marginality. Files: msi_constrained_siting_availability_2050.csv Variables: Scenario - Projected 2050 scenario name State - US state abbreviation Technology - Generating technology type solar = solar photovoltaic gas_cc_re = natural gas combined cycle (recirculating cooling) wind = onshore wind gas_cc_ccs_re = natural gas combined cycle with carbon capture sequestration (recirculating cooling) gas_cc_dry = natural gas combined cycle with (dry cooling) gas_cc_pond = natural gas combined cycle with (pond cooling) coal_conv_ccs_re = conventional coal with carbon capture sequestration (recirculating cooling) Req_Capacity_MW - The amount of rated capacity required in 2050 of the given technology type in the given state and under the given scenario from the capacity expansion plan Req_Capacity_Factor - The assumed capacity factor (fraction between 0 and 1) for the given technology type in the given state and under the given scenario by the capacity expansion plan Req_Land_km2 - The amount of land required (in km-squared) to host the required generating capacity that is capable of meeting the specified capacity factor for the given technology type in the given state and under the given scenario Req_Energy_TWh - Product of Req_Capacity_MW, Req_Capacity_Factor, and 8760/1e6 for the given technology type in the given state and under the given scenario MSI_Case - The level of MSI that siting the given technology is additionally constrained to, where >0 means siting is additionally constrained to suitable land areas that have an MSI value greater than 0 >=1 means siting is additionally constrained to suitable land areas that have an MSI value greater than or equal to 1 >=2 means siting is additionally constrained to suitable land areas that have an MSI value greater than or equal to 2 >=3 means siting is additionally constrained to suitable land areas that have an MSI value greater than or equal to 3 Soil Attribute Rasters Description: This folder contains geospatial raster files for individual soil parameters upscaled to the listed grid resolution (30m or 1 km). 1 km resolution files are a spatial average of non-missing 30m resolution values. All raster files use the USA Contiguous Albers Equal Area Conic (ESRI:102003) projection. Files: avg_cond_raster_ .tif - Average conductivity of the saturation extract across all soil horizons within a depth of 40 inches, measured in mmhos/cm max_cond_raster_ .tif - Maximum conductivity of the saturation extract across all soil horizons within a depth of 40 inches, measured in mmhos/cm min_ph_raster_ .tif- Min pH values across all soil horizons within a depth of 40 inches. avg_ph_raster_ .tif - Average pH value across all soil horizons within a depth of 40 inches. max_ph_raster_ .tif- Max pH value across all soil horizons within a depth of 40 inches. erosion_factor_raster_ .tif - Product of k-factor and percent slope flood_freq_raster_ .tif - Number of months of the year during which the area is commonly, frequently, or very frequently flooded. max_sar_raster_ .tif - Maximum sodium adsorption ratio across all horizons within a depth of 40 inches rock_frac_raster_ .tif - Fraction of the upper 6 inches of soil composed of rock fragments larger than 3 inches. temp_regime_raster_ .tif - Soil temperature regime with the following key: 0 = pergelic 1 = gelic 2 = cryic 3 = frigid 4 = isofrigid 5 = mesic 6 = isomesic 7 = thermic 8 = isothermic 9 = hyperthermic 10 =isohyperthermic Marginal Soils Index Rasters Description: This folder contains geospatial raster files of the Marginal Soils Index at the listed grid resolution (30m or 1 km). 1 km resolution files are a spatial average of 30m resolution. Both raster files use the USA Contiguous Albers Equal Area Conic (ESRI:102003) projection. A value of 0 indicates that there were no soil attributes present that indicate marginal soil. NA values indicate that data was unavailable or bodies of water. Files: marginal_soils_index_30m_raster.tif marginal_soils_index_1km_raster.tif Marginal Soils Index Resource Potential Rasters Description: This folder contains geospatial raster files of the Marginal Soils Index + Resource Potential (MSI+RP) score at 1km resolution for geothermal, solar, and wind technologies. Raster files use the USA Contiguous Albers Equal Area Conic (ESRI:102003) projection. NA values indicate that the location is not suitable for siting the given technology due to policy, environmental, socioeconomic, topological, and other constraints regardless of soil marginality level. Areas with values greater than or equal to zero represent the product of the normalized MSI value and the normalized resource potential value. Files: geothermal_msi_ep_score_raster.tif solar_msi_ep_score_raster.tif wind_msi_ep_score_raster.tif Acknowledgment IM3 is a multi-institutional effort led by Pacific Northwest National Laboratory and supported by the U.S. Department of Energy's Office of Science as part of research in MultiSector Dynamics, Earth and Environmental Systems Modeling Program. Disclaimer This material was prepared as an account of work sponsored by an agency of the United States Government. Neither the United States Government nor the United States Department of Energy, nor the Contractor, nor any or their employees, nor any jurisdiction or organization that has cooperated in the development of these materials, makes any warranty, express or implied, or assumes any legal liability or responsibility for the accuracy, completeness, or usefulness or any information, apparatus, product, software, or process disclosed, or represents that its use would not infringe privately owned rights. Reference herein to any specific commercial product, process, or service by trade name, trademark, manufacturer, or otherwise does not necessarily constitute or imply its endorsement, recommendation, or favoring by the United States Government or any agency thereof, or Battelle Memorial Institute. The views and opinions of authors expressed herein do not necessarily state or reflect those of the United States Government or any agency thereof. PACIFIC NORTHWEST NATIONAL LABORATORYoperated byBATTELLEfor theUNITED STATES DEPARTMENT OF ENERGYunder Contract DE-AC05-76RL01830

Agriculture

Marginal Soils Index Analysis & Geospatial Data

This data package contains output files associated with the Mongird et al. paper entitled "Can US power grid expansion avoid prime agricultural lands?" and is organized into four dataset directories. Each dataset is described in more detail below. 1. Marginal Soils Index Analysis Description: This folder contains a csv file with land needs and availability by state, power generating technology type, and scenario in 2050 when suitable siting areas are additionally constrained to areas with increasing levels of soil marginality. Files: msi_constrained_siting_availability_2050.csv Variables: Scenario - Projected 2050 scenario name State - US state abbreviation Technology - Generating technology type solar = solar photovoltaic gas_cc_re = natural gas combined cycle (recirculating cooling) wind = onshore wind gas_cc_ccs_re = natural gas combined cycle with carbon capture sequestration (recirculating cooling) gas_cc_dry = natural gas combined cycle with (dry cooling) gas_cc_pond = natural gas combined cycle with (pond cooling) coal_conv_ccs_re = conventional coal with carbon capture sequestration (recirculating cooling) Req_Capacity_MW - The amount of rated capacity required in 2050 of the given technology type in the given state and under the given scenario from the capacity expansion plan Req_Capacity_Factor - The assumed capacity factor (fraction between 0 and 1) for the given technology type in the given state and under the given scenario by the capacity expansion plan Req_Land_km2 - The amount of land required (in km-squared) to host the required generating capacity that is capable of meeting the specified capacity factor for the given technology type in the given state and under the given scenario Req_Energy_TWh - Product of Req_Capacity_MW, Req_Capacity_Factor, and 8760/1e6 for the given technology type in the given state and under the given scenario MSI_Case - The level of MSI that siting the given technology is additionally constrained to, where >0 means siting is additionally constrained to suitable land areas that have an MSI value greater than 0 >=1 means siting is additionally constrained to suitable land areas that have an MSI value greater than or equal to 1 >=2 means siting is additionally constrained to suitable land areas that have an MSI value greater than or equal to 2 >=3 means siting is additionally constrained to suitable land areas that have an MSI value greater than or equal to 3 Soil Attribute Rasters Description: This folder contains geospatial raster files for individual soil parameters upscaled to the listed grid resolution (30m or 1 km). 1 km resolution files are a spatial average of non-missing 30m resolution values. All raster files use the USA Contiguous Albers Equal Area Conic (ESRI:102003) projection. Files: avg_cond_raster_ .tif - Average conductivity of the saturation extract across all soil horizons within a depth of 40 inches, measured in mmhos/cm max_cond_raster_ .tif - Maximum conductivity of the saturation extract across all soil horizons within a depth of 40 inches, measured in mmhos/cm min_ph_raster_ .tif- Min pH values across all soil horizons within a depth of 40 inches. avg_ph_raster_ .tif - Average pH value across all soil horizons within a depth of 40 inches. max_ph_raster_ .tif- Max pH value across all soil horizons within a depth of 40 inches. erosion_factor_raster_ .tif - Product of k-factor and percent slope flood_freq_raster_ .tif - Number of months of the year during which the area is commonly, frequently, or very frequently flooded. max_sar_raster_ .tif - Maximum sodium adsorption ratio across all horizons within a depth of 40 inches rock_frac_raster_ .tif - Fraction of the upper 6 inches of soil composed of rock fragments larger than 3 inches. temp_regime_raster_ .tif - Soil temperature regime with the following key: 0 = pergelic 1 = gelic 2 = cryic 3 = frigid 4 = isofrigid 5 = mesic 6 = isomesic 7 = thermic 8 = isothermic 9 = hyperthermic 10 =isohyperthermic Marginal Soils Index Rasters Description: This folder contains geospatial raster files of the Marginal Soils Index at the listed grid resolution (30m or 1 km). 1 km resolution files are a spatial average of 30m resolution. Both raster files use the USA Contiguous Albers Equal Area Conic (ESRI:102003) projection. A value of 0 indicates that there were no soil attributes present that indicate marginal soil. NA values indicate that data was unavailable or bodies of water. Files: marginal_soils_index_30m_raster.tif marginal_soils_index_1km_raster.tif Marginal Soils Index Resource Potential Rasters Description: This folder contains geospatial raster files of the Marginal Soils Index + Resource Potential (MSIxRP) score at 1km resolution for geothermal, solar, and wind technologies. Raster files use the USA Contiguous Albers Equal Area Conic (ESRI:102003) projection. NA values indicate that the location is not suitable for siting the given technology due to policy, environmental, socioeconomic, topological, and other constraints regardless of soil marginality level. Areas with values greater than or equal to zero represent the product of the normalized MSI value and the normalized resource potential value. Files: geothermal_msi_rp_score_raster.tif solar_msi_rp_score_raster.tif wind_msi_rp_score_raster.tif Acknowledgment IM3 is a multi-institutional effort led by Pacific Northwest National Laboratory and supported by the U.S. Department of Energy's Office of Science as part of research in MultiSector Dynamics, Earth and Environmental Systems Modeling Program. Disclaimer This material was prepared as an account of work sponsored by an agency of the United States Government. Neither the United States Government nor the United States Department of Energy, nor the Contractor, nor any or their employees, nor any jurisdiction or organization that has cooperated in the development of these materials, makes any warranty, express or implied, or assumes any legal liability or responsibility for the accuracy, completeness, or usefulness or any information, apparatus, product, software, or process disclosed, or represents that its use would not infringe privately owned rights. Reference herein to any specific commercial product, process, or service by trade name, trademark, manufacturer, or otherwise does not necessarily constitute or imply its endorsement, recommendation, or favoring by the United States Government or any agency thereof, or Battelle Memorial Institute. The views and opinions of authors expressed herein do not necessarily state or reflect those of the United States Government or any agency thereof. PACIFIC NORTHWEST NATIONAL LABORATORYoperated byBATTELLEfor theUNITED STATES DEPARTMENT OF ENERGYunder Contract DE-AC05-76RL01830

Agriculture

Predicting Missing Regions in Charged Particle Tracks Using a Sparse 3D Convolutional Neural Network

The 2x2 Demonstrator is a prototype of ND-LAr, the liquid argon time-projection chamber of the Deep Underground Neutrino Experiment’s Near Detector complex. Both the 2x2 Demonstrator and ND-LAr are modular detectors that will have pixelated charge readouts and inactive regions wherein there is no sensitivity to charge deposition and light signals that arise from charged particle interactions with liquid argon. In the 2x2, these inactive regions are located in between the active detector modules, which introduces the challenge of inferring what charge signals ought to look like in these regions. This study explores the use of a Sparse 3D Convolutional Neural Network (ConvNet) to infer missing regions in charged particle tracks. Hits corresponding to energy depositions are voxelized into a three-dimensional grid for each track. Voxels that fall into predefined inactive regions are removed to simulate the lack of detector output. The model is trained to infer the topology of the missing track voxels, with the ultimate goal of inferring the missing charge or energy values in these voxels as well. Results indicate that this approach shows promise in prediction of missing track regions with some accuracy.

Utaegbulam, Hilary

Hot or Not? An Evaluation of Methods for Identifying Hot Moments of Nitrous Oxide Emissions From Soils

Abstract Effectively quantifying hot moments of nitrous oxide (N 2 O) emissions from agricultural soils is critical for managing this potent greenhouse gas. However, we are challenged by a lack of standard approaches for identifying hot moments, including (a) determining thresholds above which emissions are considered hot moments, and (b) considering seasonal variation in the magnitude and frequency distribution of net N 2 O fluxes. We used one year of hourly N 2 O flux measurements from 16 autochambers that varied in flux magnitude and frequency distribution in a conventionally tilled maize field in central Illinois, USA, to compare three approaches to identify hot moment thresholds: standard deviations (SD) above the mean, 1.5x the interquartile range (IQR), and isolation forest (IF) identification of anomalous values. We also compared these approaches on seasonally subdivided data (early, late, and non‐growing seasons) versus the whole year. Our analyses revealed that 1.5x IQR method best identified N 2 O hot moments. In contrast, using 2 or 4 SD both yielded hot moment threshold values too high, and IF yielded threshold values too low, leading to missed N 2 O hot moments or low net N 2 O fluxes mischaracterized as hot moments, respectively. Furthermore, seasonally subdividing the data set not only facilitated identification of smaller hot moments in the late‐ and non‐growing seasons when N 2 O hot moments were generally smaller but it also increased hot moment threshold values in the early growing season when N 2 O hot moments were larger. Consequently, of the methods evaluated here, we recommend using the 1.5x IQR method on whole year data sets to identify N 2 O hot moments.

Stuchiner, Emily R. [Institute for Sustainability,

Including frameworks of public health ethics in computational modelling of infectious disease interventions

Decisions on public health interventions to control infectious diseases are often informed by computational models. Interpreting the predicted outcomes of a public health decision requires not only high-quality modelling but also an ethical framework for assessing the benefits and harms associated with different options. The design and specification of ethical frameworks matured independently of computational modelling, so many values recognized as important for ethical decision-making are missing from computational models. We demonstrate a proof-of-concept approach to incorporate multiple public health values into the evaluation of a simple computational model for vaccination against a pathogen such as SARS-CoV-2. By examining a bounded space of alternative prioritizations of three values relevant to public health ethics (aggregate clinical burden, equity in clinical burden, equity in adverse effects from vaccination), we identify value trade-offs, where the outcomes of optimal strategies differ depending on the ethical framework. This work demonstrates an approach to incorporating diverse values into decision criteria used to evaluate outcomes of models of infectious disease interventions.

"Mathematical Biology"

Predicting Missing Regions in Charged Particle Tracks Using a Sparse 3D Convolutional Neural Network

The 2x2 Demonstrator is a prototype detector for the Deep Underground Neutrino Experiment (DUNE)'s Near Detector. Both the 2x2 Demonstrator and the Near Detector itself will have inactive regions wherein there is no sensitivity to charge deposition and light signals that arise from charged particle interactions with liquid argon. In the 2x2, these inactive regions are positioned in-between the active detector modules, which introduces the challenge of inferring what charge signals ought to look like in these regions. This study explores the use of a Sparse 3D Convolutional Neural Network (ConvNet) to infer missing regions in charged particle tracks. Hits corresponding to energy depositions are voxelized into a three-dimensional (3D) grid for each track. Inactive regions within the tracks are replaced with a dense, rectangular 3D grid of voxels, ensuring consistent step sizes in X, Y, and Z directions. Voxels in these dense regions are initialized with an energy value of -1, indicating nonphysical energy or charge. The model is trained to predict which voxels should activate as part of the track and which should not, with the goal of eventually inferring the missing charge or energy values in these voxels. Results indicate that the model accurately predicts track voxels within 1 unit in X, Y, or Z directions and effectively identifies non-track voxels, despite some overprediction. The approach shows promise in prediction of missing track regions with some accuracy.

Utaegbulam, Hilary

Daily, 30 m Resolution NDSI Data for the East River Watershed, CO for 2000-2020

This dataset contains daily Normalized Difference Snow Index (NDSI) values at 30 m spatial resolution for the East River watershed in Colorado, USA. The temporal range of these data includes water years 2001-2020. These data were created using the Spatial and Temporal Adaptive Reflectance Fusion Model (STARFM). This model fuses low spatial and high temporal resolution data from MODIS (500 m, daily) with high spatial and low temporal resolution data from Landsat (30 m, 16 days) to create a 30m synthetic daily snow product. This product allows for the analysis of historical snow covered area trends in the East River Watershed at fine spatiotemporal resolutions where it was not available previously. This research was performed as a part of the Department of Energy’s Subsurface Biogeochemical Research Program with the primary intent of better understanding the timing and spatial patterns of water delivery to the Critical Zone in mountain watersheds. Each .zip file contains one "water year" of data (October 1 - September 30; i.e., water year 2010 starts October 1, 2010 and ends September 30, 2011). Each zip file contains the following: STARFM daily Normalized Difference Snow Index (NDSI) fusion data files in GeoTiff format with one layer for each day between Landsat data acquisition dates (i.e., for dates of Landsat acquisition, the Landsat image is included for that date). The study area is located in an area of Landsat path overlap, so Landsat dates acquisitions are every 7-9 days. Landsat NDSI files containing the high spatial (30m), low temporal (7-9 days due to Landsat path overlap) resolution data used as input to STARFM in GeoTiff format with one layer for each day. Dates for which no Landsat data were obtained are included as NoData layers. MODIS NDSI files containing the high temporal (daily), low spatial (500m) resolution data used as input to STARFM in GeoTiff format with one layer for each day. Please note the MODIS data were resampled to 30m pixels for input into the STARFM model. The data have a scale factor of 10,000 and a no data value of -32767. The projection of all datasets is WGS 84 (EPSG: 4326), which has a latitude/longitude based degree resolution of 0.0002694946 X 0.0002694946, and approximates to the 30 m spatial resolution mentioned above. The Layer Index files in .csv format. They contain information for each layer in the above GeoTiff files regarding the corresponding date for each layer, the fraction of pixels in the image that contain valid data (missing data is due to either cloud cover or poor data quality; these values are not percent snow cover). Dates of Landsat overpass are indicated in these files. If no Landsat data were able to be obtained due to cloud cover or lack of Landsat Tier 1 data available on Google Earth Engine, this is also noted.

EARTH SCIENCE > CRYOSPHERE > SNOW/ICE

Multi-Artifact Analysis of Self-Admitted Technical Debt in Scientific Software

Context: Self-admitted technical debt (SATD) occurs when developers acknowledge shortcuts in code. In scientific software (SSW), such debt poses unique risks to the validity and reproducibility of results. Objective: This study aims to identify, categorize, and evaluate scientific debt, a specialized form of SATD in SSW, and assess the extent to which traditional SATD categories capture these domain-specific issues. Method: We conduct a multi-artifact analysis across code comments, commit messages, pull requests, and issue trackers from 23 open-source SSW projects. We construct and validate a curated dataset of scientific debt, develop a multi-source SATD classifier to guide SATD management, and conduct a practitioner validation to assess the practical relevance of scientific debt. Results: Our classifier performs strongly across 900,358 artifacts from 23 SSW projects. SATD is most prevalent in pull requests and issue trackers, underscoring the value of multi-artifact analysis. Models trained on traditional SATD often miss scientific debt, emphasizing the need for its explicit detection in SSW. Practitioner validation confirmed that scientific debt is both recognizable and useful in practice. Conclusions: Scientific debt represents a unique form of SATD in SSW that that is not adequately captured by traditional categories and requires specialized identification and management. Our dataset, classification analysis, and practitioner validation results provide the first formal multi-artifact perspective on scientific debt, highlighting the need for tailored SATD detection approaches in SSW.

Melin, Eric [Boise State University]

Paper and Plastics in Landfills: The Missed Opportunity

Landfilling paper and plastic waste represents missed opportunities for resource and energy recovery and significant loses in market value and disposal costs. More paper and plastics are landfilled than previously thought which highlights opportunities for improvement.

disposal cost

COMPASS-FME Synoptic Site Tree Greenhouse Gas Concentrations

These data are tree stem greenhouse gas concentrations collected from tree gas wells at some of the COMPASS-FME (Coastal Observations, Mechanisms, and Predictions Across Systems and Scales; see https://compass.pnnl.gov/) 'synoptic' sites in the Chesapeake Bay region: Moneystump (MSM), Goodwin Islands (GWI), and GCReW (GCW). The sap flow monitoring trees at these sites in the Upland (UP) and Transition (TR) zones were cored and had gas wells installed at breast height. There were also some dead standing trees cored, gas well installed, and sampled at MSM and GWI. The GCW UP samples overlap with the TEMPEST experiment control plot, so the GCW UP data was pulled from the TEMPEST page and included here. These data provide crucial information about possible pathways for the greenhouse gas (carbon dioxide and methane, CO2 and CH4 respectively) production and emission (or in the case of CH4, perhaps taken up from) the atmosphere.All data are plain text CSV (comma separated value) files and require no special software to read.Updated 2025-10-09 to fix two missing dates (lines 77 and 78 in the data file).

54 ENVIRONMENTAL SCIENCES

Filling the Gaps: A Bayesian Mixture Model for Imputing Missing Soil Water Content Data

ABSTRACT Soil water content (SWC) data are central to evaluating how soil moisture varies over time and space and influences critical plant and ecosystem functions, especially in water‐limited drylands. However, sensors that record SWC at high frequencies often malfunction, leading to incomplete timeseries and limiting our understanding of dryland ecosystem dynamics. We developed an analytical approach to impute missing SWC data, which we tested at six eddy flux tower sites along an elevation gradient in the southwestern United States. We impute missing data as a mixture of linearly interpolated SWC between the observed endpoints of a missing data gap and SWC simulated by an ecosystem water balance model (SOILWAT2). Within a Bayesian framework, we allowed the relative utility (mixture weight) of each component (linearly interpolated vs. SOILWAT2) to vary by depth, site and gap characteristics. We explored “fixed” weights versus “dynamic” weights that vary as a function of cumulative precipitation, average temperature, and time since the start of the gap. Both models estimated missing SWC data well ( R 2 = 0.70–0.88 vs. 0.75–0.91 for fixed vs. dynamic weights, respectively), but the utility of linearly interpolated versus SOILWAT2 values depended on site and depth. SOILWAT2 was more useful for more arid sites, shallower depths, longer and warmer gaps and gaps that received greater precipitation. Overall, the mixture model reliably gap‐fills SWC, while lending insight into processes governing SWC dynamics. This approach to impute missing data could be adapted to accommodate more than two mixture components and other types of environmental timeseries.

Ogle, Kiona [School of Informatics, Computing, and

Generative learning for slow manifolds and bifurcation diagrams

In dynamical systems characterized by separation of time scales, the approximation of so called “slow manifolds”, on which the long term dynamics lie, is a useful step for model reduction. Initializing on such slow manifolds is a useful step in modeling, since it circumvents fast transients, and is crucial in multiscale algorithms (like the equation-free approach) alternating between fine scale (fast) and coarser scale (slow) simulations. In a similar spirit, when one studies the infinite time dynamics of systems depending on parameters, the system attractors (e.g., its steady states) lie on bifurcation diagrams (curves for one-parameter continuation, and more generally, on manifolds in state parameter space. Sampling these manifolds gives us representative attractors (here, steady states of ODEs or PDEs) at different parameter values. Algorithms for the systematic construction of these manifolds (slow manifolds, bifurcation diagrams) are required parts of the “traditional” numerical nonlinear dynamics toolkit. In more recent years, as the field of Machine Learning develops, conditional score-based generative models (cSGMs) have been demonstrated to exhibit remarkable capabilities in generating plausible data from target distributions that are conditioned on some given label. It is tempting to exploit such generative models to produce samples of data distributions (points on a slow manifold, steady states on a bifurcation surface) conditioned on (consistent with) some quantity of interest (QoI, observable). In this work, we present a framework for using cSGMs to quickly (a) initialize on a low-dimensional (reduced-order) slow manifold of a multi-time-scale system consistent with desired value(s) of a QoI (a “label”) on the manifold, and (b) approximate steady states in a bifurcation diagram consistent with a (new, out-of-sample) parameter value. This conditional sampling can help uncover the geometry of the reduced slow-manifold and/or approximately “fill in” missing segments of steady states in a bifurcation diagram. Finally, the quantity of interest, which determines how the sampling is conditioned, is either known a priori or identified using manifold learning-based dimensionality reduction techniques applied to the training data.

Dynamical systems