Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “evaluation datasets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Impacts of benchmarking choices on inferred model skill of the Arctic–Boreal terrestrial carbon cycle

Abstract Land surface models require continuous validation against observations to improve and reduce simulation uncertainty. However, inferred model performance can be heavily influenced by subjective choices made in the selection and application of observational data products. A key area often misrepresented by models is the Arctic–Boreal region, which is a potential tipping point region in Earth’s climate system due to large permafrost carbon stocks that are vulnerable to release with climate warming. We use the International Land Model Benchmarking (ILAMB) framework to evaluate how the model skill of TRENDY-v9 models varies based on the choice of observational-based benchmark and how benchmarks are applied in model evaluation. This analysis uses global datasets integrated into ILAMB and new, regionally-specific observational products from the Arctic–Boreal Vulnerability Experiment. Our results cover the overall time period of 1979–2019 and show that model scores can vary substantially depending on the data product applied, with higher model scores indicating better model performance against observations. The lowest model scores occur when benchmarked against regional, compared to global, datasets. We also evaluate observed and modeled functional relationships between ecosystem respiration and air temperature and between gross primary production and precipitation. Here, we find that the magnitude and shape of the responses are strongly impacted by the choice of observational dataset and the approach used to construct the functional relationship benchmark. These results suggest that model evaluation studies could conclude a false sense of model skill if only using a single benchmark data product or if not applying regional data products when performing a regional model analysis. Collectively, our findings highlight the influence of benchmarking choices on model evaluation and point to the need for benchmarking guidelines when assessing model skill.

Poe, Jeralyn (ORCID:0000000318495278)↗

Assessing Low-Temperature Geothermal Play Types: Relevant Data and Play Fairway Analysis Methods

The U.S. Department of Energy (DOE) Geothermal Technologies Office (GTO) is supporting the Geothermal Heating and Cooling Geospatial Datasets and Analysis project conducted by the National Renewable Energy Laboratory (NREL) as part of a broader effort to demonstrate the multi-faceted value of integrating geothermal power and geothermal heating and cooling (GHC) technologies into national decarbonization plans and community energy plans. Currently, there is a need to establish baseline low-temperature geothermal resource datasets and evaluate methods for deploying these technologies to provide the basis for supporting private sector investment. This project is focused on collecting baseline datasets, updating conceptual models, and creating Play Fairway Analysis (PFA) workflows for low-temperature (<150 degrees Celsius) geothermal resources of different geothermal play types (i.e., sedimentary basin, orogenic belts, and radiogenic geothermal play types) that could be used for geothermal heating and cooling (GHC), combined heat and power (CHP), and other geothermal direct uses (GDU) applications. Low-temperature geothermal resources are defined as reservoirs - natural or engineered - with temperatures <150 degrees Celsius. While the focus in the NREL effort is on GHC, resources at the upper end of this temperature range can also be used for small-scale power generation. This project does not include Ground Source Heat Pumps (GSHPs) technologies because they can be effectively developed almost anywhere. Low-temperature geothermal resources have not been studied as extensively as higher- to medium-temperature geothermal resources, but there is recent interest in improving understanding of these types of resources with an uptick of interest in geothermal technologies for decarbonizing heating and cooling systems. In addition, Enhanced Geothermal Systems (EGS) and other emerging technologies for exploiting petrothermal resources have opened the possibility of utilizing deep sedimentary basin systems, where porous media provide permeability and high temperatures can be reached at great depths. This project takes the approach of classifying low- temperature geothermal resources by geothermal play type (GPT). We defined and characterized three major classes of low-temperature GPT: sedimentary basins, orogenic systems, and radiogenic systems. We develop methodologies for evaluating and analyzing the potential for these resources building off the PFA approach to de-risking geothermal exploration and characterization. The proposed PFA approach for low-temperature geothermal resources includes: 1) identifying relevant data (e.g., datasets such bottom-hole temperatures from oil and gas wells, heat flow data, Quaternary faults and stress field data, geophysical data, etc.); 2) grouping and weighting of relevant datasets into PFA criteria (e.g., geological, risk, and economic criteria); 3) uncertainty quantification; 4) developing favorability or common risk maps for low-temperature geothermal resources to identify potential locations for more focused data collection; and 5) estimating electric power generation and heating potential at those locations using the GeoRePORT Resource Size Assessment Tool (RSAT). This project should facilitate future deployment of GHC, CHP, and GDU by providing data, tools, and a workflow applicable to low-temperature geothermal resources. Increased deployment of GHC and GDU will help achieve national and local decarbonization goals.

15 GEOTHERMAL ENERGY↗

Using remote sensing to quantify the additional climate benefits of California forest carbon offset projects

Abstract Nature‐based climate solutions are a vital component of many climate mitigation strategies, including California's, which aims to achieve carbon neutrality by 2045. Most carbon offsets in California's cap‐and‐trade program come from improved forest management (IFM) projects. Since 2012, various landowners have set up IFM projects following the California Air Resources Board's IFM protocol. As many of these projects approach their 10th year, we now have the opportunity to assess their effectiveness, identify best practices, and suggest improvements toward future protocol revisions. In this study, we used remote sensing‐based datasets to evaluate the carbon trends and harvest histories of 37 IFM projects in California. Despite some current limitations and biases, these datasets can be used to quantify carbon accumulation and harvest rates in offset project lands relative to nearby similar “control” lands before and after the projects began. Five lines of evidence suggest that the carbon accumulated in offset projects to date has generally not been additional to what might have otherwise occurred: (1) most forests in northwestern California have been accumulating carbon since at least the mid‐1980s and continue to accumulate carbon, whether enrolled in offset projects or not; (2) harvest rates were high in large timber company project lands before IFM initiation, suggesting they are earning carbon credits for forests in recovery; (3) projects are often located on lands with higher densities of low‐timber‐value species; (4) carbon accumulation rates have not yet increased on lands that enroll as offset projects, relative to their pre‐enrollment levels; and (5) harvest rates have not decreased on most project lands since offset project initiation. These patterns suggest that the current protocol should be improved to robustly measure and reward additionality. In general, our framework of geospatial analyses offers an important and independent means to evaluate the effectiveness of the carbon offsets program, especially as these data products continue improving and as offsets receive attention as a climate mitigation strategy.

59 BASIC BIOLOGICAL SCIENCES↗

Challenges in predicting protein-protein interactions of understudied viruses: Arenavirus-human interactions

Understanding protein-protein interactions (PPIs) between viruses and host organisms is crucial for uncovering infection mechanisms and identifying potential therapeutic targets. The ability to generalize PPI predictive models across understudied viruses presents a significant challenge. In this work, we use arenavirus-human PPIs to illustrate the difficulties associated with model generalization, which are compounded by a lack of both positive and negative data. We employ a Transfer Learning approach to investigate arenavirus-human PPIs by utilizing models trained on better-studied virus-human and human-human PPIs. Additionally, we curate and assess four types of negative sampling datasets to evaluate their impact on model performance. Despite the overall high accuracies (93–99 %) and AUPRC scores (0.8–0.9) appearing promising, further analysis indicates that these performance metrics can be misleading due to data leakage, data bias, and overfitting, especially concerning under-represented viral proteins. We reveal these gaps and assess the impact of data imbalance using standard k-fold cross-validation and Independent Blind Testing with a Balanced Dataset, resulting in a drop in accuracy below 50 %. We propose a viral protein-specific evaluation framework that categorizes viral proteins into majority and minority classes based on their representation in the dataset, enabling comparison of model performance across these groups using balanced accuracies. This framework offers a more robust evaluation of model generalizability, addressing biases inherent in standard evaluation techniques and paving the way for more reliable PPI prediction models for understudied viruses.

59 BASIC BIOLOGICAL SCIENCES↗

Development of observation-based global multilayer soil moisture products for 1970 to 2016

Abstract. Soil moisture (SM) datasets are critical to understanding the global water, energy, and biogeochemical cycles and benefit extensive societal applications. However, individual sources of SM data (e.g., in situ and satellite observations, reanalysis, offline land surface model simulations, Earth system model – ESM – simulations) have source-specific limitations and biases related to the spatiotemporal continuity, resolutions, and modeling and retrieval assumptions. Here, we developed seven global, gap-free, long-term (1970–2016), multilayer (0–10, 10–30, 30–50, and 50–100 cm) SM products at monthly 0.5∘ resolution (available at https://doi.org/10.6084/m9.figshare.13661312.v1; Wang and Mao, 2021) by synthesizing a wide range of SM datasets using three statistical methods (unweighted averaging, optimal linear combination, and emergent constraint). The merged products outperformed their source datasets when evaluated with in situ observations (mean bias from −0.044 to 0.033 m3 m−3, root mean square errors from 0.076 to 0.104 m3 m−3, Pearson correlations from 0.35 to 0.67) and multiple gridded datasets that did not enter merging because of insufficient spatial, temporal, or soil layer coverage. Three of the new SM products, which were produced by applying any of the three merging methods to the source datasets excluding the ESMs, had lower bias and root mean square errors and higher correlations than the ESM-dependent merged products. The ESM-independent products also showed a better ability to capture historical large-scale drought events than the ESM-dependent products. The merged products generally showed reasonable temporal homogeneity and physically plausible global sensitivities to observed meteorological factors, except that the ESM-dependent products underestimated the low-frequency temporal variability in SM and overestimated the high-frequency variability for the 50–100 cm depth. Based on these evaluation results, the three ESM-independent products were finally recommended for future applications because of their better performances than the ESM-dependent ones. Despite uncertainties in the raw SM datasets and fusion methods, these hybrid products create added value over existing SM datasets because of the performance improvement and harmonized spatial, temporal, and vertical coverages, and they provide a new foundation for scientific investigation and resource management.

54 ENVIRONMENTAL SCIENCES↗

Data, model inputs, and analysis scripts associated with a manuscript on stream intermittency controls across spatial scales in Pacific Northwest watersheds

NOTE: The manuscript associated with this data package is currently in review. The data may be revised based on reviewer feedback. Upon manuscript acceptance, this data package will be updated with the final dataset and additional metadata. This data package is associated with the manuscript "Hydroclimatic Memory and Watershed Template Shape Stream Intermittency: Multi-scale Attribution Using Process-based Simulation and Explainable ML" by Niroula et al. (2026), submitted to Water Resources Research (WRR). The study investigates the dominant controls on stream intermittency across local, reach, and watershed scales using a coupled process-based simulation and explainable machine-learning framework. Long-term daily simulations from the Advanced Terrestrial Simulator (ATS) were used to generate wetness states and ponded-depth responses over river-corridor cells. These ATS outputs were then aggregated across scales and used to train XGBoost (eXtreme Gradient Boosting) models. SHAP (SHapley Additive exPlanations) was applied to quantify the relative importance of hydroclimatic forcings, watershed template attributes, and antecedent-memory effects in shaping intermittency behavior. The analysis is carried out for three contrasting Pacific Northwest watersheds: Oak Creek (OCW), American River Watershed (ARW), and H.J. Andrews (HJA). Across these testbeds, the package contains ATS-ready watershed inputs, ATS run configuration and selected output files, model-evaluation data products, intermittency-analysis datasets, machine-learning target-feature tables, SHAP outputs, and notebooks used to organize, analyze, and visualize results. At a high level, the package documents a workflow in which ATS provides the physically based simulation backbone and explainable machine learning is used as a post-processing attribution tool. The contents are intended to support interpretation of the manuscript figures and results, provide context for how intermittency metrics were generated at multiple scales, and preserve the key artifacts needed to understand and reuse the analysis workflow. The package contains a high-level directory summary file (`summary.txt`) and four main content folders (1) `evaluation_plots` contains evaluation figures and supporting evaluation datasets; (2) `intermittency_plots` contains intermittency-focused analysis notebook and prepared datasets; (3) `ml-training-and-shap_values_plots` contains ML training inputs, SHAP outputs, and figure-generation notebooks; and (4) `watershed_mesh_and_ats_input` contains ATS model setup materials, forcing inputs, geometry, and selected run files. More specifically, the `evaluation_plots` folder contains the notebook used for ATS evaluation plotting and site-specific evaluation datasets. These include evapotranspiration and water-balance products for three watersheds, as well as an Oak Creek field-measurement discharge file. The `intermittency_plots` folder contains the notebook used for intermittency analysis and the prepared datasets used to analyze intermittent and non-intermittent wetness behavior across the study watersheds. The `ml-training-and-shap_values_plots` folder contains notebooks and outputs for the machine-learning and explainability workflow. This includes the main XGBoost and SHAP notebook(s), a beeswarm plotting notebook, target-feature tables for machine-learning training, SHAP summary tables, and per-sample SHAP value archives. The `watershed_mesh_and_ats_input` folder contains ATS-related watershed inputs and supporting materials. This includes mesh and shape products, ATS-readable LAI and meteorological forcing inputs, selected ATS spinup and transient-run files, and a watershed workflow example notebook. Subdirectories are organized by watershed where applicable.All files are .cpg (codepage files), .csv (comma-separated values), .dbf (database files), .exo (Exodus mesh format), .h5 (HDF5 format), .ipynb (Jupyter notebooks), .pkl (Python pickle), .prj (projection files), .sh (shell scripts), .shp (shapefile geometry), .shx (shapefile index), .txt (text files), or .xml (markup data).

Advanced Terrestrial Simulator↗

Machine learning for the redox potential prediction of molecules in organic redox flow battery

Here, organic redox flow batteries (ORFB) are recognized as an innovative technology for the large-scale storage of renewable energy. The redox potential of organic redox-active molecules plays a vital role in their performance. Advanced screening techniques like high-throughput experiment and machine learning (ML) have significantly enhanced organic material performance and transformed the field of ORFB. However, the scarcity of experimental data poses a considerable challenge for ML model development in this domain. In our study, we developed lightweight graph-based Gaussian process regression (GPR) models with GPU-accelerated marginalized graph kernel and hybrid kernel to predict the redox potentials of organic redox-active molecules for ORFBs, specifically focusing on small datasets. To evaluate model accuracy, we created a new experimental database of organic redox-active molecules by the data from hundreds of published papers and assembled previous computational datasets. We also considered some key parameters, such as pH conditions and solvent type, to assess their impact on redox potential prediction. Our GPR model predicted redox potentials with high accuracy across all datasets using minimal training data. The study provides powerful tools for molecule screening and design and delivers valuable guidance on designing training datasets for costly experiments.

25 ENERGY STORAGE↗

A knowledge-informed large language model framework for U.S. nuclear power plant shutdown initiating event classification for probabilistic risk assessment

Identifying and classifying shutdown initiating events (SDIEs) is critical for developing shutdown probabilistic risk assessment for nuclear power plants. Existing computational approaches cannot achieve satisfactory performance due to the challenges of unavailable large, labeled datasets, imbalanced event types, and label noise. To address these challenges, we propose a hybrid pipeline that integrates a knowledge-informed machine learning model to prescreen non-SDIEs and a large language model (LLM) to classify SDIEs into four types. In the prescreening stage, we proposed a set of 44 SDIE text patterns that consist of the most salient keywords and phrases from six SDIE types. Text vectorization based on the SDIE patterns generates feature vectors that are highly separable by using a simple binary classifier. The second stage builds Bidirectional Encoder Representations from Transformers (BERT)-based LLM, which learns generic English language representations from self-supervised pretraining on a large dataset and adapts to SDIE classification by fine-tuning it on an SDIE dataset. The proposed approaches are evaluated on a dataset with 10,928 events using precision, recall ratio, F 1 score, and average accuracy. In conclusion, the results demonstrate that the prescreening stage can exclude more than 97% non-SDIEs, and the LLM achieves an average accuracy of 95.1% for SDIE classification.

99 - GENERAL AND MISCELLANEOUS↗

Mesoscale evaluation of AMPS using AWARE radar observations of a wind and precipitation event over the Ross Island region of Antarctica

Surface, upper-air, and radar observations are used to assess the performance of the Antarctic Mesoscale Prediction System (AMPS) in simulating the mesoscale aspects of a wind and precipitation event over the Ross Island region of Antarctica that spanned January 16–20, 2016. The observations, collected during the Atmospheric Radiation Measurement (ARM) West Antarctic Radiation Experiment (AWARE), provide a unique dataset for evaluating AMPS, especially the radar observations that facilitate a three-dimensional depiction of winds and precipitation. Comparisons of AMPS forecast data with surface meteorology, balloon-sounding, and profiling radar observations at and above sites near McMurdo Station reveal a mixture of similarities and differences. A generally southerly flow is evident at low levels in both the AMPS simulations and observed Doppler radial velocities. AMPS winds are comparable to those observed at the surface and aloft in terms of magnitude, direction, and timing but the strongest simulated southerly flow is displaced eastward relative to the observations. AMPS-simulated reflectivity over the broader Ross Island region is more limited in areal extent and smaller in magnitude than observed by a scanning Doppler radar. Three episodes of surface precipitation are observed near McMurdo Station over the five-day event with peak rates of ~3 mm h -1 and a total accumulation of ~22 mm. However, AMPS produces no surface precipitation at that location over the five-day event due to a low-level dry bias in the forecasts. Herein, the results show the first observationally based three-dimensional understanding of meteorology in the Ross Island region.

AMPS↗

Re-evaluating probable maximum precipitation estimates: sensitivity to transposition domains and storm rotation using modern datasets

This study examines the sensitivity of Probable Maximum Precipitation (PMP) estimates to key methodological decisions embedded in the legacy approach adopted in the U.S. National Weather Service Hydrometeorological Reports No. 51 and No. 52. Although widely used for infrastructure design and risk regulation, fundamental aspects of PMP estimation—such as storm sample size, transposition domain, maximization procedures, and storm rotation—remain poorly constrained and lack formal guidance. Using the Red Rock watershed in Iowa as a case study, and leveraging the 2002–2023 NOAA Analysis of Record for Calibration (AORC) precipitation dataset, we systematically evaluate how each methodological choice, individually and in combination, influences PMP estimates. Our findings demonstrate that PMP is not a fixed physical upper bound but rather a modeling construct shaped heavily by user-defined assumptions. Notably, PMP values derived from modern gridded rainfall datasets can be substantially higher than the legacy estimate used in the original spillway design for Red Rock Dam. Decisions regarding storm sample size, domain extent, climatological window, and particularly storm rotation all contributed to higher PMP estimates. Storm rotation alone—a loosely constrained element in the current PMP practice—can amplify PMP by more than 25%. These results reveal the lack of standardized bounds in current PMP workflows and the need for systematic sensitivity and uncertainty analysis. As PMP estimation shifts toward probabilistic approaches, incorporating physically meaningful storm attributes will be key to developing more transparent, defensible methods for dam safety and climate-resilient infrastructure.

Probable maximum precipitation↗

Leveraging large language models to address data scarcity in machine learning for graphene synthesis

Machine learning in experimental materials science faces significant challenges due to the scarcity of data, which are costly and time-consuming to generate, particularly when relying on in-house experiments. Literature data mining offers a potential solution but introduces issues like mixed data quality, inconsistent formats, and non-uniform reporting of synthesis parameters, resulting in partially missing and heterogeneous features across the dataset. Here, we propose data imputation and feature engineering methods that employ pre-trained large language models (LLMs) to enhance machine learning performance on scarce, heterogeneous datasets, demonstrated on graphene CVD synthesis data and the ML-HydPARK hydrogen storage dataset. GPT models perform data imputation via tailored prompting and semantic normalization of inconsistently reported features through embeddings, for example, to harmonize the complex nomenclature of CVD substrates. Beyond yielding more diverse and richer feature representations than traditional methods such as K-nearest neighbors (KNN) and Multivariate Imputation by Chained Equations (MICE), LLM-based data imputation is evaluated against dataset characteristics and prompting strategies. We vary the level of autonomy granted to the LLM, from generic prompting that leverages pre-trained knowledge for autonomous data generation to data-informed prompting that constrains outputs using target-specific information, and demonstrate which level of autonomy yields superior imputation performance across datasets and feature types. The proposed data engineering methods markedly improve downstream performance; for example, in graphene layer number classification using a support vector machine (SVM), binary accuracy increases from 39% to 65% and ternary accuracy from 52% to 72%. Fine-tuning experiments on both datasets show that combining our proposed LLM-based data imputation and feature encoding methods with numerical machine learning predictors outperforms standalone fine-tuned LLM predictors in data-scarce settings. The proposed strategies emphasize data enhancement techniques rather than refining learning architectures or regularizing loss functions, offering a broadly applicable framework for improving machine learning performance on scarce, inhomogeneous datasets.

Chemical vapor deposition↗

High bias machine learning for antineutrino-based safeguards for small reactors

The statistical methods used for antineutrino detection will need to be improved to effectively monitor the inventory of next-generation nuclear reactors. In this sensitivity study, we evaluate machine learning models compared to previously used statistical approaches to identify diversion scenarios in a simulated Advanced Fast Reactor (AFR)-100. A chi-square goodness-of-fit technique, which individually compares the simulated antineutrino yields to the expected antineutrino yield, resulted in precise but low diversion detection probability. Various support vector machine (SVM) models were applied with diverse training datasets to evaluate the robustness of the method towards unexpected or “unseen” diversion scenarios. Furthermore, our results indicate that while the SVM models significantly improved the detection probability of near-field antineutrino-based safeguards, up to a probability of ~0.04, for the simulated small reactor, the detection system still needs improvements to reach the 0.2 detection limit established by the International Atomic Energy Agency.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

Machine learning and atomic layer deposition: Predicting saturation times from reactor growth profiles using artificial neural networks

In this work, we explore the application of deep neural networks to the optimization of atomic layer deposition (ALD) processes. In particular, we focus on a one-shot optimization problem, where we try to predict the optimal dose time that leads to saturation everywhere in the reactor based on thickness values measured at different points of an ALD reactor after a single trial growth. In order to tackle this problem, we introduce a dataset designed to train neural networks to predict saturation times based on these inputs for a cross-flow ALD reactor. Here, we then explore the predictive ability of artificial neural networks of different depths and sizes using a separate testing dataset to evaluate their accuracies. The results obtained show that networks trained using stochastic gradient descent methods can accurately predict saturation times without requiring any additional information on the surface kinetics. This provides a viable approach to minimize the number of experiments required to optimize new ALD processes in a known reactor, and it highlights the way machine learning can be leveraged for thin film growth and manufacturing. While the datasets and training procedure depend on the reactor geometry, the trained neural networks provide a general surrogate model connecting thickness values and trial dose times with optimal saturation times that can be reused for different ALD processes within the same reactor.

36 MATERIALS SCIENCE↗

Application of artificial intelligence methods in the international roughness index prediction of rigid and composite pavements: a systematic review

The International Roughness Index (IRI) is a widely adopted metric for quantifying pavement roughness, directly influencing vehicle safety, ride comfort, and overall roadway performance. In recent years, the use of Machine Learning (ML) models for IRI prediction has gained momentum, with the goal of improving the allocation of maintenance and rehabilitation resources by enabling accurate assessments of pavement conditions. Most prior reviews, however, have concentrated on flexible pavements, leaving a notable gap regarding rigid and composite pavements. To address this gap, the present study conducts a systematic review of Artificial Intelligence (AI) methods applied to IRI prediction for rigid and composite pavements. Literature published between 2004 and 2025 is synthesized to highlight prevailing trends, methodological contributions, and directions for future research. Particular attention is given to the types of models employed, the datasets used for training and validation, and the role of input variables and data-processing strategies. Across the included studies, ensemble learning methods (especially gradient boosting variants such as XGBoost), artificial neural networks, and hybrid architectures frequently achieved high predictive skill, with several models reporting test-set coefficients of determination approaching 0.9–0.96, indicating strong potential for capturing the influence of traffic, pavement structure, and climatic factors. Since these results are obtained from heterogeneous datasets and evaluation protocols, they are interpreted qualitatively rather than as strict cross-study rankings. Analysis of input variables revealed that pavement age and initial IRI were included in 91% (21 of 23) and 78% (18 of 23) of studies, respectively. Climatic variables such as the freezing index appeared in 57% (13 of 23), while traffic-related factors were considered in 65% (15 of 23). The findings underscore the importance of standardized, high-quality datasets, such as those from the Long-Term Pavement Performance (LTPP) program, along with data consistency, model interpretability, computational efficiency, and replicability in enhancing IRI prediction. Future research should focus on incorporating input variable selection techniques to identify the most influential predictors, thereby improving accuracy and robustness. Integrating these approaches with advanced non-linear data-driven models, coupled with robust hyperparameter optimization, holds considerable promise for strengthening the reliability of IRI prediction and supporting resilient pavement management strategies.

42 ENGINEERING↗

Self-supervised and multi-fidelity learning for extended predictive soil spectroscopy

Infrared spectroscopy is a cost-effective, non-destructive, and environmentally benign technology that is increasingly recognized as an important solution for meeting the global demand for soil data. While both near-infrared (NIR) and mid-infrared (MIR) diffuse reflectance spectroscopy enable rapid estimation of soil properties, they present a significant trade-off: NIR offers superior scalability and lower operational costs, whereas MIR provides higher analytical fidelity by capturing fundamental molecular vibrations. In this study, we propose a self-supervised, multi-fidelity learning framework designed to bridge this gap. Our approach leverages large-scale MIR spectral libraries to learn a compact, transferable latent representation, into which NIR spectra are subsequently aligned for downstream prediction. The workflow consists of pretraining a latent model on a large MIR library, adapting the representation using a smaller paired NIR–MIR dataset, and evaluating generalization on an independent external test set. Across a range of chemical and physical soil properties, we found that MIR-derived embeddings improved prediction accuracy relative to baseline models that used raw MIR inputs. Predictions derived from the spectrum conversion (NIR to MIR) task did not match the performance of the original MIR spectra but were similar or superior to predictive performance of NIR-only models, suggesting the unified spectral latent space can effectively leverage the larger and more diverse MIR dataset for prediction of soil properties not well represented in current NIR libraries.

54 ENVIRONMENTAL SCIENCES↗

Hidden Features: How Subsurface and Landscape Heterogeneity Govern Hydrologic Connectivity and Stream Chemistry in a Montane Watershed

ABSTRACT Hydrologic connectivity is defined as the connection among stores of water within a watershed and controls the flux of water and solutes from the subsurface to the stream. Hydrologic connectivity is difficult to quantify because it is goverened by heterogeniety in subsurface storage and permeability and responds to seasonal changes in precipitation inputs and subsurface moisture conditions. How interannual climate variability impacts hydrologic connectivity, and thus stream flow generation and chemistry, remains unclear. Using a rare, four‐year synoptic stream chemistry dataset, we evaluated shifts in stream chemistry and stream flow source of Coal Creek, a montane, headwater tributary of the Upper Colorado River. We leveraged compositional principal component analysis and end‐member mixing to evaluate how seasonal and interannual variation in subsurface moisture conditions impacts stream chemistry. Overall, three main findings emerged from this work. First, three geochemically distinct end members were identified that constrained stream flow chemistry: reach inflows, and quick and slow flow groundwater contributions. Reach inflows were impacted by historic base and precious metal mine inputs. Bedrock fractures facilitated much of the transport of quick flow groundwater and higher‐storage subsurface features (e.g., alluvial fans) facilitated the transport of slow flow groundwater. Second, the contributions of different end members to the stream changed over the summer. In early summer, stream flow was composed of all three end members, while in late summer, it was composed predominantly of reach inflows and slow flow groundwater. Finally, we observed minimal differences in proportional composition in stream chemistry across all four years, indicating seasonal variability in subsurface moisture and spatial heterogeneity in landscape and geologic features had a greater influence than interannual climate fluctuation on hydrologic connectivity and stream water chemistry. These findings indicate that mechanisms controlling solute transport (e.g., hydrologic connectivity and flow path activation) may be resilient (i.e., able to rebound after perturbations) to predicted increases in climate variability. By establishing a framework for assessing compositional stream chemistry across variable hydrologic and subsurface moisture conditions, our study offers a method to evaluate watershed biogeochemical resilience to variations in hydrometeorological conditions.

Johnson, Keira [College of Earth, Ocean, and Atmos↗

Projections of future climate for U.S. national assessments: past, present, future

Climate assessments consolidate our understanding of possible future climate conditions as represented by climate projections, which are largely based on the output of global climate models. Over the past 30 years, the scientific insights gained from climate projections have been refined through model structural improvements, emerging constraints on climate feedbacks, and increased computational efficiency. Within the same period, the process of assessing and evaluating information from climate projections has become more defined and targeted to inform users. As the size and audience of climate assessments has expanded, the framing, relevancy, and accessibility of projections has become increasingly important. This paper reviews the use of climate projections in national climate assessments (NCA) while highlighting challenges and opportunities that have emerged over time. Reflections and lessons learned address the continuous process to understand the broadening assessment audience and evolving user needs. Insights for future NCA development include (1) identifying benchmarks and standards for evaluating downscaled datasets, (2) expanding efforts to gather research gaps and user needs to inform how climate projections are presented in the assessment (3) providing practitioner guidance on the use, interpretation, and reporting of climate projections and uncertainty to better inform decision-making.

Assessment↗