Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Scarce data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

A Representation Fusion Framework for Decoupling Diagnostic Information in Multimodal Learning

Modern medicine increasingly relies on multimodal data, ranging from clinical notes to imaging and genomics, to guide diagnosis and treatment. However, integrating these heterogeneous data sources in a principled and interpretable manner remains a major challenge. We present MODES (Multi-mOdal Disentangled Embedding Space), a representation fusion framework that explicitly separates shared and modality-specific factors of variation, offering a structured latent space for multimodal information that improves both prediction and interpretability. By leveraging pre-trained unimodal foundation models, MODES mitigates the dependency on extensive paired datasets, crucial in data-scarce clinical settings. We introduce a masking strategy that optimizes representation dimensionality by eliminating low-information dimensions, to achieve compact, information-rich representations. Our framework demonstrates superior performance in predicting diagnoses and phenotypes compared to unimodal and conventional fusion models. MODES also enables robust diagnostic inference in missing data scenarios, offering an opportunity toward interpretable and efficient multimodal diagnostics in personalized healthcare.

60 APPLIED LIFE SCIENCES↗

Predicting weather impacts on corn production in a data-limited region using a transfer learning approach

The stability of food supply and prices may depend more on annual changes in yields from year-to-year variability in weather than on longer-term average changes from changing climatic conditions. However, the absence of high-quality data on crop yields at fine spatial resolutions in many regions of the world makes it challenging to statistically model their response to interannual variability in weather patterns. Therefore, there is a need for empirical methods that can project annual crop yield changes even in limited data regions. Here, we propose a transfer learning algorithm that uses high spatial resolution data from one region to project yields in another region with more limited data. The goal of our work is to understand what data types can be beneficial for transferring learning from a source region to a very different target region with more limited data. We utilize Long Short-Term Memory to develop a transfer learning model that is trained on historical county-level corn yield in the United States and predicts district-level corn yield variations in India. Even using smaller amounts of data in India, simulating a data-scarce region, we achieve an average root mean square error of 0.48 bu acre−1 in predicting interannual yield variations. Using Shapley values to interpret results, we explore the contribution of the different weather parameters to interannual yield variability and find a larger influence of precipitation-related variables. Our study demonstrates the usefulness of this method for transferring models of weather impacts on crop yields trained on a data-rich country to one with more limited data. It suggests the potential of applying the transfer learning model to mitigate the need for extensive raw data globally.

Vishwakarma, Srishti [ORNL] (ORCID:000000031674419↗

BuildingQA: A Benchmark for Natural Language Question Answering over Building Knowledge Graphs

Graph-based representations of building metadata using ontologies like Brick are vital for smart building applications, but querying them remains a challenge for practitioners. Knowledge Graph Question Answering (KGQA) systems, meant to retrieve answers from natural language questions, traditionally require large-scale training data, making them ill-suited for the specialized and data-scarce building domain. The advent of Large Language Models (LLMs) offers a paradigm shift, enabling zero-shot natural language querying without building/domain-specific training. Yet, there is no standardized benchmark for building-specific KGQA which can guide and validate research in this area. To address this gap, our work makes three primary contributions. First, we introduce the BuildingQA Benchmark Dataset, constructed through a multi-stage process of collecting practitioner data, augmenting it with LLMs for linguistic diversity, and curating a final set of 188 questions across 4 buildings. Second, we characterize the benchmark's complexity and ambiguity, introducing a novel method to quantify its "lexical gap" and providing a four-stage diagnostic framework for analyzing how systems fail. Third, we benchmark zero-shot LLM-powered KGQA systems to establish baseline performance and analyze their failure modes. Our evaluation reveals that top-performing systems achieve a maximum F1 score of only 0.38. This result does not indicate a failure of these powerful systems, but rather underscores the unique challenges posed by our benchmark. It demonstrates a critical performance gap, showing that current methods successful on general KGs struggle with the specific lexical and structural nuances of the building domain. BuildingQA1 thus provides the benchmark dataset and foundational analysis needed to drive the development of novel, domain-aware methods required to unlock the use of semantic data in buildings.

Mulayim, Ozan Baris↗

Physics-informed latent neural operator for real-time predictions of time-dependent parametric PDEs

Deep operator network (DeepONet) has shown significant promise as surrogate models for systems governed by partial differential equations (PDEs), enabling accurate mappings between infinite-dimensional function spaces. However, when applied to systems with high-dimensional input-output mappings arising from large numbers of spatial and temporal collocation points, these models often require heavily overparameterized networks, leading to long training times. Latent DeepONet addresses some of these challenges by introducing a two-step approach: first learning a reduced latent space using a separate model, followed by operator learning within this latent space. While efficient, this method is inherently data-driven and lacks mechanisms for incorporating physical laws, limiting its robustness and generalizability in data-scarce settings. Here, in this work, we propose PI-Latent-NO, a physics-informed latent neural operator framework that integrates governing physics directly into the learning process. Our architecture features two coupled DeepONets trained end-to-end: a Latent-DeepONet that learns a low-dimensional representation of the solution, and a Reconstruction-DeepONet that maps this latent representation back to the physical space. By embedding PDE constraints into the training via automatic differentiation, our method eliminates the need for labeled training data and ensures physics-consistent predictions. The proposed framework is both memory and compute-efficient, exhibiting near-constant scaling with problem size and demonstrating significant speedups over traditional physics-informed operator models. We validate our approach on a range of parametric PDEs, showcasing its accuracy, scalability, and suitability for real-time prediction in complex physical systems.

Latent representations↗

Derivation of physical equations for high-speed laser welding using large language models

It is challenging to formulate complex physical phenomena that occur in a manufacturing process, particularly when the available data are limited, rendering conventional data-driven approaches ineffective. This study aims to predict humping onset in high-speed laser welding by introducing a novel framework, namely text-to-equations generative pre-trained transformer (T2EGPT). This method leverages the capabilities of large language models (LLMs), in combination with sparse experimental data and enriched literature data, to derive an interpretable and generalizable equation for predicting humping initiation. By capturing key correlations among physical parameters, T2EGPT generates a compact and dimensionless expression that accurately predicts hump formation. The equation reveals that humping arises from the interplay between inertia-driven backward melt flow and capillary-driven surface stabilization, where inertial forces drive molten metal backward and capillary forces resist surface deformation. Furthermore, compared to traditional data-driven models, T2EGPT demonstrates enhanced predictive accuracy and cross-material transferability. More broadly, this study highlights the potential of LLMs to integrate textual information with data-driven discovery, enabling the extraction of physical laws in data-scarce scientific domains.

36 MATERIALS SCIENCE↗

Climate-extreme modeling framework for sustainable flood management in the Arabian Peninsula

Evaluating extreme precipitation events (EPEs) is essential for building climate-resilient water management strategies, but it remains a major challenge in ungauged basins. Using the 26,070 km 2 Wadi al-Rummah basin in central Saudi Arabia as a case study, we developed an alternative, reliable, cost-effective satellite-based framework that combines empirically derived EPE thresholds, imagery-calibrated 2D hydrodynamic modeling, GRACE water-storage diagnostics, and bias-corrected CMIP6 projections to assess flood hazards and recharge potential under current and future climate scenarios in ungauged basins. The integrated approach and the resulting findings followed four key steps: (1) Identified a 22.5 mm EPE threshold, the 80th percentile of 3-day GPM/IMERG rainfall (2000–2024), aligned with flood-triggering events (Nov 2018: 23–28 mm; Apr, 2023: 42 mm); (2) Developed and calibrated a RiverFlow2D model using Sentinel-2 and PlanetScope imagery for the November 2018 flood, accurately reproducing flood depth and extent (RMSE ≤0.31 m; fuzzy-Dice ≥0.91), and estimating runoff (41 %), infiltration (25 %), and evaporation (34 %); (3) Independently validated the model with the April 2023 event (RMSE ≤0.35 m; fuzzy-Dice ≥0.86); (4) Conducted climate projections (2025–2100) from five bias-corrected NEX-GDDP CMIP6 models that revealed a 34 % increase in EPE intensity under SSP2-4.5 and 48 % under SSP5-8.5 scenarios, relative to 20th-century baselines. Our findings indicate that while intensifying extremes in the 21st century increase flood risk, the results highlight the potential for episodic recharge if effective retention strategies are employed, and offer a transferable model for climate-informed planning in data-scarce arid regions.

CMIP6↗

Advancing stream temperature prediction with a generalizable large-sample framework across CONUS river reaches

Accurately predicting stream temperature in ungauged basins remains a critical challenge for water resource management, thermoelectric power plant cooling, and ecosystem conservation. Large-sample machine learning models trained on hundreds of well-monitored river basins have shown remarkable performance; however, such models have yet to be developed solely using forcing data that can be readily extracted to simulate stream temperatures anywhere in the contiguous United States (CONUS). In this study, we present a scalable, large-sample deep learning framework using Long Short-Term Memory (LSTM) networks to simulate daily stream temperatures in ungauged basins across the CONUS. The framework leverages both modeled reanalysis of meteorological and streamflow inputs as well as static attributes available for all 2.7 million CONUS river reaches in the National Hydrography Dataset Plus (NHDPlusV2). By generating dynamical inputs from predefined thermally relevant upstream contributing areas, rather than the entire upstream basin, the model also offers improvements in very large basins where full-basin averaging can dilute the most important influences on stream temperature. Evaluated across 300 basins, the model achieves a median Mean Absolute Error (MAE) of 1.1 °C and a Nash-Sutcliffe Efficiency (NSE) of 0.95 on temporally and spatially distinct test folds—comparable to models trained exclusively using meteorological and streamflow observational data. The flexible, high-performing framework generalizes to any unmonitored river reach without significant regulation or unnatural thermal input immediately upstream, substantially expanding predictive capabilities in data-scarce regions.

Hydrology↗

Why Firn Quakes

Snow dampens sounds, but anecdotal reports concisely describe audible propagating collapse events—firnquakes—in Antarctic and Arctic snowfields. We propose combining granular and continuum mechanics to form a testable theory for conditioning, triggering, and propagation of firnquakes consistent with scarce data. A central condition for collapse events is unconsolidated firn at depth. As firn grains compact, stresses are transmitted along force chains which carry the overburden and transition into a continuous medium by pressure sintering. This granular legacy creates solid-like supports of denser layers that keep the material below unconsolidated. Dynamic amplification triggers local brittle failure of the supports, which induces a cascade of collapse propagation. Using bulk density from ice cores as proxy for stiffness, we find the flexural wave speed by collapsing supports matches the recorded firnquake velocities on the order of 100 m/s. Our theory is to be tested in firn sheets and other compacting granular materials.

Voigtländer, A. [Lawrence Berkeley National Labora↗

Differentiable modelling to unify machine learning and physical models for geosciences

Process-based modelling offers interpretability and physical consistency in many domains of geosciences but struggles to leverage large datasets efficiently. Machine-learning methods, especially deep networks, have strong predictive skills yet are unable to answer specific scientific questions. Here, in this Perspective, we explore differentiable modelling as a pathway to dissolve the perceived barrier between process-based modelling and machine learning in the geosciences and demonstrate its potential with examples from hydrological modelling. ‘Differentiable’ refers to accurately and efficiently calculating gradients with respect to model variables or parameters, enabling the discovery of high-dimensional unknown relationships. Differentiable modelling involves connecting (flexible amounts of) prior physical knowledge to neural networks, pushing the boundary of physics-informed machine learning. It offers better interpretability, generalizability, and extrapolation capabilities than purely data-driven machine learning, achieving a similar level of accuracy while requiring less training data. Additionally, the performance and efficiency of differentiable models scale well with increasing data volumes. Under data-scarce scenarios, differentiable models have outperformed machine-learning models in producing short-term dynamics and decadal-scale trends owing to the imposed physical constraints. Differentiable modelling approaches are primed to enable geoscientists to ask questions, test hypotheses, and discover unrecognized physical relationships. Future work should address computational challenges, reduce uncertainty, and verify the physical significance of outputs.

58 GEOSCIENCES↗

Towards the understanding of the genuine three-body interaction for p–p–p and p–p–Λ

Three-body nuclear forces play an important role in the structure of nuclei and hypernuclei and are also incorporated in models to describe the dynamics of dense baryonic matter, such as in neutron stars. So far, only indirect measurements anchored to the binding energies of nuclei can be used to constrain the three-nucleon force, and if hyperons are considered, the scarce data on hypernuclei impose only weak constraints on the three-body forces. In this work, we present the first direct measurement of the p–p–p and p–p–Λ systems in terms of three-particle correlation functions carried out for pp collisions at $\sqrt{s}$=13 TeV. Three-particle cumulants are extracted from the correlation functions by applying the Kubo formalism, where the three-particle interaction contribution to these correlations can be isolated after subtracting the known two-body interaction terms. A negative cumulant is found for the p–p–p system, hinting to the presence of a residual three-body effect while for p–p–Λ the cumulant is consistent with zero. This measurement demonstrates the accessibility of three-baryon correlations at the LHC.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

ML-based Micro-CT SOFC Microstructure Models (from Kent 2026 Microstructural Augmentation paper)

Overview -------------------------- This repository contains datasets from the manuscript **"Enhanced Generalizability to Deep-Learning Quantification of 3D Microstructural Characteristics through Microstructurally Aware Augmentation of Scarce Data"** (*William F. Kent, Rochan Bajpai, Rachel C. Kurchin, William K. Epting, Harry W. Abernathy, Paul A. Salvador. Submitted 2026*). The methods are also described in the dissertation **Data Intensive Analysis of Solid Oxide Cell Microstructures** (*Doctoral dissertation, Carnegie Mellon University, 2025*). The datasets here are trained convolutional neural network (CNN) models for predicting key microstructural properties of solid oxide cell (SOC) electrodes from low-res, 2-channel 3D images, as well as some helpful code. The parameters for input images are provided in the paper. Sample data is provided in the file `Combined_anode_aug_dual_1k_examples` - that particular data was used to train `anode_all_aug.pth` and will work most accurately with that model. Please familiarize yourself with all caveats on accuracy and applicability, as detailed in the associated paper. Usage -------------------------- The basic usage is as follows, assuming `model_fn` is the path to the .pth file, and `X` is 2-channel input image(s) of the proper dimensions (either one image of shape `[2,12,24,24]`, or a batch of N input images of shape `[N,2,12,24,24]`): from CNN_inferencer import load_model_for_inference model = load_model_for_inference(model_fn) y_predicted = model(X) The model object automatically handles input scaling and output de-scaling based on the way the models were trained - in other words, pass in a 2-channel micro-CT image, and it will output microstructural property values in real units. ## Other model object attributes Note that model has useful attributes other than its forward pass model(X). * `model.output_descaler` - returns the output descaler object. Model does the de-scaling when generating inferences, but you may want to re-use this de-scaler on other values to e.g. compare predictions to ground truth from already-scaled training data. * `model.prop_names` - Gives the property names of the predicted y values, in order. Only exists if there's an output scaler as part of the model object, which there will be in the models provided here. ## Usage with sample data Here is a short script to use with the included sample data. from CNN_inferencer import display_predictions, load_model_for_inference, calculate_mape, parity_plot import h5py import numpy as np model_fn = 'anode_all_aug.pth' data_fn = 'Combined_anode_aug_dual_1k_examples.h5' N_samples = 200 figure_outdir = '.' model = load_model_for_inference(model_fn) with h5py.File(data_fn,'r') as f: XX = f['X'] #These are the 2-channel 3D images yy = f['y'] #These are the ground-truth microstructural properties, but they have been scaled for training - need to de-scale below N = XX.shape[0] #How many images total in the input data file #Run inferences on N_samples random samples from XX. #Run in a batch, much more efficient than one at a time. ii = np.random.choice(N,N_samples,replace=False) ii.sort() y_pred = model(XX[ii]) #Get the original/true (but normalized/scaled) values from the training dataset... #Because they were normalized, they are not in real units yet. So let's also de-scale them using model.output_scaler. y_true = model.output_scaler.transform(yy[ii]) #Let's display actual values for just 5 random ones for i in np.random.choice(N_samples,5,replace=False): display_predictions(y_true[i], y_pred[i], model.prop_names) #Make parity plots for each property (ground truth vs predicted values) #Also label each plot with the mean abs. percent error (MAPE) of the predicted values for i,key in enumerate(model.prop_names): mape = calculate_mape(y_true[:,i], y_pred[:,i]) parity_plot(y_true[:,i], y_pred[:,i], figure_outdir, key, extra_title=f' ({mape:.2f}% MAPE)')

3D microstructure↗

Satellite Embedding-Based Population Imputation for Areas with Missing Building Footprint Data: A Computer Vision-Based Approach

High-resolution population modeling is important for supporting effective decision-making across diverse sectors. LandScan Mosaic generates population estimates at the level of individual buildings and aggregates them to 3 arc-second grids, and this approach performs well in regions where building footprint data are comprehensive and reliable. However, large portions of the globe still suffer from incomplete, sparse, or entirely missing building stock datasets, creating a structural limitation for strictly building-based population models. To address this research gap, this study proposes a computer vision-based framework that employs Google Earth Engine satellite embeddings and UNet, which allows us to directly impute grid-level population estimates in building-data-deficient areas. Applied to Taiwan as a case study, the framework achieved strong predictive performance with R$^{2}$ of 0.89, RMSE of 18.70, and MAE of 8.41, outperforming traditional machine learning approaches. Notably, the proposed framework effectively addressed building false-positive errors inherent in Global Human Settlement Layer (GHSL) data, correctly identifying uninhabited areas that were erroneously classified as populated. The framework also offers significant advantages for global population mapping, particularly in terms of scalability and temporal consistency, thereby extending the coverage and accuracy of high-resolution population products in data-scarce regions worldwide. Urban planners, decision makers, and related stakeholders can obtain granular population distributions to support more accurate and targeted infrastructure investment, service delivery, resource allocation, and risk assessment decisions.

97 MATHEMATICS AND COMPUTING↗

The suitability of differentiable, physics-informed machine learning hydrologic models for ungauged regions and climate change impact assessment

As a genre of physics-informed machine learning, differentiable process-based hydrologic models (abbreviated as δ or delta models) with regionalized deep-network-based parameterization pipelines were recently shown to provide daily streamflow prediction performance closely approaching that of state-of-the-art long short-term memory (LSTM) deep networks. Meanwhile, δ models provide a full suite of diagnostic physical variables and guaranteed mass conservation. Here, we ran experiments to test (1) their ability to extrapolate to regions far from streamflow gauges and (2) their ability to make credible predictions of long-term (decadal-scale) change trends. We evaluated the models based on daily hydrograph metrics (Nash–Sutcliffe model efficiency coefficient, etc.) and predicted decadal streamflow trends. For prediction in ungauged basins (PUB; randomly sampled ungauged basins representing spatial interpolation), δ models either approached or surpassed the performance of LSTM in daily hydrograph metrics, depending on the meteorological forcing data used. They presented a comparable trend performance to LSTM for annual mean flow and high flow but worse trends for low flow. For prediction in ungauged regions (PUR; regional holdout test representing spatial extrapolation in a highly data-sparse scenario), δ models surpassed LSTM in daily hydrograph metrics, and their advantages in mean and high flow trends became prominent. In addition, an untrained variable, evapotranspiration, retained good seasonality even for extrapolated cases. The δ models' deep-network-based parameterization pipeline produced parameter fields that maintain remarkably stable spatial patterns even in highly data-scarce scenarios, which explains their robustness. Combined with their interpretability and ability to assimilate multi-source observations, the δ models are strong candidates for regional and global-scale hydrologic simulations and climate change impact assessment.

54 ENVIRONMENTAL SCIENCES↗

Use of Very High-Resolution Optical Data for Landslide Mapping and Susceptibility Analysis Along the Karnali Highway, Nepal

The Karnali highway is a vital transport link and the only primary roadway that connects the remote Karnali region to the lowlands in Mid-Western Nepal. Every year there are reports of landslides blocking the road, making this area largely inaccessible. However, little effort has focused on systematically identifying landslides and landslide-prone areas along this highway. In this study, landslides were mapped with an object-based approach from very high-resolution optical satellite imagery obtained by the DigitalGlobe constellation in 2012 and PlanetScope in 2018. Landslides ranging from 10 to 30,496 sq.m were detected within a 3 km buffer along the highway. Most of the landslides were located at lower elevations (between 500–1500 m) and on steep south-facing slopes. Landslides tended to cluster closer to the highway, near drainage channels and away from faults. Landslides were also most prevalent within the Kuncha Formation geologic class, and the forested and agricultural land cover classes. A susceptibility map was then created using a logistic regression methodology to highlight patterns in landslide activity. The landslide susceptibility map showed a good prediction rate with an area under the curve (AUC) of 0.90. A total of 33% of the study arealies in high/very high susceptibility zones. The map highlighted the lower elevated areas between Bangesimal and Manma towns with the Kuncha Formation geologic class as being the most hazardous. The banks of the Karnali River, its tributaries and areas near the highway were also highly susceptible to landslides. The results highlight the potential of very high-resolution optical imagery for documenting detailed spatial information on landslide occurrence, which enables susceptibility assessment in remote and data scarce regions such as the Karnali highway.

Pukar Amatya↗

Tonlé Sap Food Security and Agriculture: Evaluating the Effects of Land Use and Hydrological Change on Ecosystem Vitality using Remotely-Sensed Data in the Tonlé Sap Lake Basin

Tonlé Sap Lake, the largest lake in Southeast Asia, is a critical source of fish and freshwater resources for the region. The health of this freshwater system is under pressure from accelerating dam construction, intensifying agriculture, deforestation, and changing climate patterns, forcing tradeoffs between immediate food security and the long-term vitality and productivity of the ecosystem. Efficient freshwater system monitoring is crucial to navigating these challenges. In collaboration with Conservation International, the Cambodian Ministry of Water Resources and Meteorology, and the Tonlé Sap Authority, we developed and tested remotely-sensed proxies for sub-indicators of the Freshwater Health Index (FHI), which is typically calculated using in situ datasets. We used landcover datasets derived from Landsat 5 Thematic Mapper (TM), Landsat 7 Enhanced Thematic Mapper Plus (ETM+), Landsat 8 Operational Land Imager (OLI), PROBA-V Vegetation sensor (VGT), Sentinel-2 Multispectral Imager (MSI), Advanced Very High Resolution Radiometer (AVHRR), and Envisat Medium Resolution Imaging Spectrometer (MERIS) as inputs to calculate land cover naturalness and bank modification. Additionally, we created a lake-level time series using a collection of altimetry data sources to estimate deviation from natural flow. We observed a decrease in landcover naturalness and a breakdown in the volume and regularity of annual lake levels from 2000-2020, reflecting increased pressure on water supply and agricultural productivity. At least 8% of forested areas in the basin were lost and rice harvest intensity increased over the course of the study period. These results will help our partners make informed decisions regarding freshwater management. Furthermore, our remotely-sensed FHI analysis can be replicated in other regions, providing decision makers with a snapshot of freshwater health in data-scarce environments.

Marco Vallejos↗

Soil Moisture Estimation in South Asia via Assimilation of SMAP Retrievals

A soil moisture retrieval assimilation framework is implemented across South Asia in an attempt to improve regional soil moisture estimation as well as to provide a consistent regional soil moisture dataset. This study aims to improve the spatiotemporal variability of soil moisture estimates by assimilating Soil Moisture Active Passive (SMAP) near-surface soil moisture retrievals into a land surface model. The Noah-MP (v4.0.1) land surface model is run within the NASA Land Information System software framework to model regional land surface processes. NASA Modern-Era Retrospective Analysis for Research and Applications (MERRA2) and Global Precipitation Measurement (GPM) Integrated Multi-satellitE Retrievals (IMERG) provide the meteorological boundary conditions to the land surface model. Assimilation is carried out using both cumulative distribution function (CDF)-corrected (DA-CDF) and uncorrected SMAP retrievals (DA-NoCDF). CDF matching is applied to correct the statistical moments of the SMAP soil moisture retrieval relative to the land surface model. Comparison of assimilated and model-only soil moisture estimates with publicly available in situ measurements highlights the relative improvement in soil moisture estimates by assimilating SMAP retrievals. Across the Tibetan Plateau, DA-NoCDF reduced the mean bias and RMSE by 8.4 % and 9.4 %, even though assimilation only occurred during less than 10 % of the study period due to frozen (or partially frozen) soil conditions. The best goodness-of-fit statistics were achieved for the IMERG DA-NoCDF soil moisture experiment. The general lack of publicly available in situ measurements across irrigated areas limited a domain-wide direct model validation. However, comparison with regional irrigation patterns suggested correction of biases associated with an unmodeled hydrologic phenomenon (i.e., anthropogenic influence via irrigation) as a result of SMAP soil moisture retrieval assimilation. The greatest sensitivity to assimilation was observed in cropland areas. Improvements in soil moisture potentially translate into improved spatiotemporal patterns of modeled evapotranspiration, although limited influence from soil moisture assimilation was observed on modeled processes within the carbon cycle such as gross primary production. Improvement in fine-scale modeled estimates by assimilating coarse-scale retrievals highlights the potential of this approach for soil moisture estimation over data-scarce regions.

Jawairia Ahmad↗

Utilizing Earth Observations to Understand Landscape Patterns and Assist in Wildlife Management in Iona National Park, Angola

Following the end of the Angolan Civil War (1975-2002), human habitation in Iona National Park has grown exponentially, as has the livestock population. An ongoing drought beginning in 2017 has brought people, livestock, and wildlife into increasing competition for resources within the park. This study used Earth observation data, primarily Landsat and Sentinel imagery, to examine landscape trends to improve wildlife preservation approaches in Iona National Park, Angola. In collaboration with the NGO African Parks, we developed a robust land use and land cover (LULC) classification model using remote sensing data to augment sparse ground-based data in this arid land region. We used Google Earth Engine and a random forest classifier to map vegetation types, water bodies, and potential wildlife habitats. This analysis resulted in a high spatial resolution LULC time-series between 1984-2023, highlighting critical periods of socioecological change over the past 40 years. These results increased the partner’s ability to make scientifically grounded decisions about resource allocation and conservation priorities. This analysis supports the feasibility of applying remote sensing techniques coupled with machine learning models in dry regions, where standard survey methods are frequently limited by accessibility and resource availability. However, we identified limitations in ground-truth data and the difficulty of recognizing certain vegetation types in arid areas. Despite these limitations, the study demonstrated Earth observations' ability to transform wildlife management techniques in distant and data-scarce locations, providing a reproducible foundation for similar ecosystems around the world.

Emmanuel Aklie↗