Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data scarce”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Hydrologic Regionalization under Data Scarcity: Implications for Streamflow Prediction

Continuous streamflow prediction is crucial in many applications of water resources planning and management. However, streamflow prediction is challenging, particularly in data-scarce regions. Here, we demonstrate an approach to regionalize the flow duration curve for predicting daily streamflow in the data-scare region of the central Himalayas. We developed a regression-based model to estimate streamflow at various segments of a flow duration curve by incorporating basin characteristics and climate variables. This study analyzes the sensitivities of proximity and characteristics between the donor (gauged) and receptor (ungauged) basins for time-series streamflow prediction. Our results show that regionalization techniques perform better in low to medium flows over high flows. Our findings are significant in the central Himalayan regional context to inform operational and management decisions in water sector projects like hydropower plants, which generally rely on low-to-medium streamflow information. Although the quantitative results are region-specific, the approach and insights are generalizable to the Himalayan region.

54 ENVIRONMENTAL SCIENCES↗

Knowledge-guided learning with curated prior genetic biomarkers for robust model interpretation

Abstract Motivation Knowledge-guided learning offers effective and robust model training strategies in data-scarce settings by incorporating established domain knowledge, thereby enhancing generalization, robustness, and interpretability. By contrast, conventional deep learning approaches rely purely on data-driven learning, which can limit robust model interpretability, particularly in high-dimensional settings with limited size samples. In computational biology, knowledge-guided learning has primarily leveraged network- and structural-based knowledge, leading to biologically interpretable representations and enhanced predictive performance compared to conventional approaches. However, curated biomarkers, one of the most accessible forms of biological knowledge, remain largely unexplored within knowledge-guided paradigms. Results In this study, we propose a model-agnostic training paradigm, Biomarker-driven Explainable Prior-guided Learning (BioExPL), that can be applied to any neural networks that incorporates curated prior knowledge. BioExPL enforces neural networks to reflect curated biomarker priors in their latent representations through a novel knowledge-alignment loss. BioExPL consistently demonstrated significantly improved predictive performance and enhanced model interpretability with minimized computational overhead in simulation studies and intensive experiments on multiple cancer datasets. BioExPL not only integrates prior curated knowledge into the model but also accurately identifies unknown associated signals additionally. BioExPL is model-agnostic and domain-independent, enabling its integration into diverse neural network architectures. Availability and implementation The open-source is publicly available at: https://github.com/datax-lab/BioExPL.

Baek, Beomsu [Department of Computer Science, Univ↗

Designing alloys with process-mapping AI pre-trained on empirical knowledge

<span style="font-family: Calibri, sans-serif; font-size: 12pt;">Accelerated materials design should match the recent trends in the product development cycles. Materials data analytics can be used to significantly shorten development time of specialized alloys needed for next generation energy applications. However, it faces a challenge of scarce data available for training ML models. Incorporation of the domain knowledge into deep-learning graph structure via fuzzy pre-training and causal process imitation presents a viable approach to developing accurate data-driven models and reliable alloy design tools, with limited datasets. Artificial Intelligence (AI) was used in this study to incorporate such knowledge in the domain-specific computational tool, pyroMind. The tool provides not only novel design ideas but also their interpretation via physics and engineering concepts.</span>

Romanov, Vyacheslav↗

A Representation Fusion Framework for Decoupling Diagnostic Information in Multimodal Learning

Modern medicine increasingly relies on multimodal data, ranging from clinical notes to imaging and genomics, to guide diagnosis and treatment. However, integrating these heterogeneous data sources in a principled and interpretable manner remains a major challenge. We present MODES (Multi-mOdal Disentangled Embedding Space), a representation fusion framework that explicitly separates shared and modality-specific factors of variation, offering a structured latent space for multimodal information that improves both prediction and interpretability. By leveraging pre-trained unimodal foundation models, MODES mitigates the dependency on extensive paired datasets, crucial in data-scarce clinical settings. We introduce a masking strategy that optimizes representation dimensionality by eliminating low-information dimensions, to achieve compact, information-rich representations. Our framework demonstrates superior performance in predicting diagnoses and phenotypes compared to unimodal and conventional fusion models. MODES also enables robust diagnostic inference in missing data scenarios, offering an opportunity toward interpretable and efficient multimodal diagnostics in personalized healthcare.

60 APPLIED LIFE SCIENCES↗

Predicting weather impacts on corn production in a data-limited region using a transfer learning approach

The stability of food supply and prices may depend more on annual changes in yields from year-to-year variability in weather than on longer-term average changes from changing climatic conditions. However, the absence of high-quality data on crop yields at fine spatial resolutions in many regions of the world makes it challenging to statistically model their response to interannual variability in weather patterns. Therefore, there is a need for empirical methods that can project annual crop yield changes even in limited data regions. Here, we propose a transfer learning algorithm that uses high spatial resolution data from one region to project yields in another region with more limited data. The goal of our work is to understand what data types can be beneficial for transferring learning from a source region to a very different target region with more limited data. We utilize Long Short-Term Memory to develop a transfer learning model that is trained on historical county-level corn yield in the United States and predicts district-level corn yield variations in India. Even using smaller amounts of data in India, simulating a data-scarce region, we achieve an average root mean square error of 0.48 bu acre−1 in predicting interannual yield variations. Using Shapley values to interpret results, we explore the contribution of the different weather parameters to interannual yield variability and find a larger influence of precipitation-related variables. Our study demonstrates the usefulness of this method for transferring models of weather impacts on crop yields trained on a data-rich country to one with more limited data. It suggests the potential of applying the transfer learning model to mitigate the need for extensive raw data globally.

Vishwakarma, Srishti [ORNL] (ORCID:000000031674419↗

BuildingQA: A Benchmark for Natural Language Question Answering over Building Knowledge Graphs

Graph-based representations of building metadata using ontologies like Brick are vital for smart building applications, but querying them remains a challenge for practitioners. Knowledge Graph Question Answering (KGQA) systems, meant to retrieve answers from natural language questions, traditionally require large-scale training data, making them ill-suited for the specialized and data-scarce building domain. The advent of Large Language Models (LLMs) offers a paradigm shift, enabling zero-shot natural language querying without building/domain-specific training. Yet, there is no standardized benchmark for building-specific KGQA which can guide and validate research in this area. To address this gap, our work makes three primary contributions. First, we introduce the BuildingQA Benchmark Dataset, constructed through a multi-stage process of collecting practitioner data, augmenting it with LLMs for linguistic diversity, and curating a final set of 188 questions across 4 buildings. Second, we characterize the benchmark's complexity and ambiguity, introducing a novel method to quantify its "lexical gap" and providing a four-stage diagnostic framework for analyzing how systems fail. Third, we benchmark zero-shot LLM-powered KGQA systems to establish baseline performance and analyze their failure modes. Our evaluation reveals that top-performing systems achieve a maximum F1 score of only 0.38. This result does not indicate a failure of these powerful systems, but rather underscores the unique challenges posed by our benchmark. It demonstrates a critical performance gap, showing that current methods successful on general KGs struggle with the specific lexical and structural nuances of the building domain. BuildingQA1 thus provides the benchmark dataset and foundational analysis needed to drive the development of novel, domain-aware methods required to unlock the use of semantic data in buildings.

Mulayim, Ozan Baris↗

Physics-informed latent neural operator for real-time predictions of time-dependent parametric PDEs

Deep operator network (DeepONet) has shown significant promise as surrogate models for systems governed by partial differential equations (PDEs), enabling accurate mappings between infinite-dimensional function spaces. However, when applied to systems with high-dimensional input-output mappings arising from large numbers of spatial and temporal collocation points, these models often require heavily overparameterized networks, leading to long training times. Latent DeepONet addresses some of these challenges by introducing a two-step approach: first learning a reduced latent space using a separate model, followed by operator learning within this latent space. While efficient, this method is inherently data-driven and lacks mechanisms for incorporating physical laws, limiting its robustness and generalizability in data-scarce settings. Here, in this work, we propose PI-Latent-NO, a physics-informed latent neural operator framework that integrates governing physics directly into the learning process. Our architecture features two coupled DeepONets trained end-to-end: a Latent-DeepONet that learns a low-dimensional representation of the solution, and a Reconstruction-DeepONet that maps this latent representation back to the physical space. By embedding PDE constraints into the training via automatic differentiation, our method eliminates the need for labeled training data and ensures physics-consistent predictions. The proposed framework is both memory and compute-efficient, exhibiting near-constant scaling with problem size and demonstrating significant speedups over traditional physics-informed operator models. We validate our approach on a range of parametric PDEs, showcasing its accuracy, scalability, and suitability for real-time prediction in complex physical systems.

Latent representations↗

Observations on the deformation of metal microspheres in shock-driven polymer flows

We report that solid particles can be fragmented by a fast-moving fluid if their velocity difference is great enough, such as during the atmospheric entry of meteoroids or the shock compression of engineered particulate composites. The extent of particle deformation and breakup in such systems is poorly understood because the necessary extreme conditions make observation difficult and data scarce. To meet this need, experiments combining ultrafast synchrotron-based radiography with plate impact loading were performed at the dynamic compression sector at the advanced photon source. Metal microspheres of several densities and strengths (Au, Ta, and W) were placed inside a polymer matrix. A planar shock wave was then produced in the polymer by the impact of a gun-launched flyer plate. X-ray images of the resulting flow were collected at ~150ns intervals. These images document the progression of particle deformation across a range of flow conditions and particle materials. They show that the extent of deformation is sensitive to the ratio of drag stress to particle strength. The deforming particle's shape is determined by the initial shock–particle interaction, fluid stagnation pressure, and vorticity, each acting on its own timescale. A set of scaling relationships is presented to capture these observations and enable comparison with prior hydrodynamic data. The result is a framework for predicting the conditions under which strong particles are severely deformed by a shock-driven flow.

36 MATERIALS SCIENCE↗

Modeling regional precipitation over the Indus River basin of Pakistan using statistical downscaling

Complex processes govern spatiotemporal distribution of precipitation within the high-mountainous headwater regions (commonly known as the upper Indus basin (UIB)), of the Indus River basin of Pakistan. Reliable precipitation simulations particularly over the UIB present a major scientific challenge due to regional complexity and inadequate observational coverage. Here, we present a statistical downscaling approach to model observed precipitation of the entire Indus basin, with a focus on UIB within available data constraints. Taking advantage of recent high altitude (HA) observatories, we perform precipitation regionalization using K-means cluster analysis to demonstrate effectiveness of low-altitude stations to provide useful precipitation inferences over more uncertain and hydrologically important HA of the UIB. We further employ generalized linear models (GLM) with gamma and Tweedie distributions to identify major dynamic and thermodynamic drivers from a reanalysis dataset within a robust cross-validation framework that explain observed spatiotemporal precipitation patterns across the Indus basin. Final statistical models demonstrate higher predictability to resolve precipitation variability over wetter southern Himalayans and different lower Indus regions, by mainly using different dynamic predictors. The modeling framework also shows an adequate performance over more complex and uncertain trans-Himalayans and the northwestern regions of the UIB, particularly during the seasons dominated by the westerly circulations. However, the cryosphere-dominated trans-Himalayan regions, which largely govern the basin hydrology, require relatively complex models that contain dynamic and thermodynamic circulations. Furthermore, we also analyzed relevant atmospheric circulations during precipitation anomalies over the UIB, to evaluate physical consistency of the statistical models, as an additional measure of reliability. Overall, our results suggest that such circulation-based statistical downscaling has the potential to improve our understanding towards distinct features of the regional-scale precipitation across the upper and lower Indus basin. Additionally, such understanding should help to assess the response of this complex, data-scarce, and climate-sensitive river basin amid future climatic changes, to serve communal and scientific interests.

54 ENVIRONMENTAL SCIENCES↗

A functional global sensitivity measure and efficient reliability sensitivity analysis with respect to statistical parameters

Sensitivity analysis and reliability assessment are two important aspects of structural and system safety. Epistemic uncertainty with respect to probabilistic model of input parameters due to lack of knowledge is present in many scarce-data applications and complicates the characterization of uncertainty in model response. In this article, we present two importance measures to evaluate the impact of distribution parameters on the probability distribution function (PDF) of the output and the failure probability. The epistemic uncertainty associated with the distribution parameters is modeled as random variables. Additionally, a modified extended polynomial chaos expansion (MEPCE) approach is introduced in which aleatory and epistemic random variables are modeled and propagated simultaneously while allowing the separate assessment for any single epistemic variable. A MEPCE-based kernel density estimation (KDE) construction provides a composite map from each epistemic variable to the response PDF. The functional global sensitivity index of the PDF with respect to the distribution parameters is thus derived, as a function of output, which is both more informative and more efficient than standard scalar sensitivity measures. Reliability sensitivity indices can be readily evaluated by integrating the global sensitivity index function over the failure zone. Three illustrative examples are used to demonstrate the proposed methodology.

42 ENGINEERING↗

Derivation of physical equations for high-speed laser welding using large language models

It is challenging to formulate complex physical phenomena that occur in a manufacturing process, particularly when the available data are limited, rendering conventional data-driven approaches ineffective. This study aims to predict humping onset in high-speed laser welding by introducing a novel framework, namely text-to-equations generative pre-trained transformer (T2EGPT). This method leverages the capabilities of large language models (LLMs), in combination with sparse experimental data and enriched literature data, to derive an interpretable and generalizable equation for predicting humping initiation. By capturing key correlations among physical parameters, T2EGPT generates a compact and dimensionless expression that accurately predicts hump formation. The equation reveals that humping arises from the interplay between inertia-driven backward melt flow and capillary-driven surface stabilization, where inertial forces drive molten metal backward and capillary forces resist surface deformation. Furthermore, compared to traditional data-driven models, T2EGPT demonstrates enhanced predictive accuracy and cross-material transferability. More broadly, this study highlights the potential of LLMs to integrate textual information with data-driven discovery, enabling the extraction of physical laws in data-scarce scientific domains.

36 MATERIALS SCIENCE↗

Climate-extreme modeling framework for sustainable flood management in the Arabian Peninsula

Evaluating extreme precipitation events (EPEs) is essential for building climate-resilient water management strategies, but it remains a major challenge in ungauged basins. Using the 26,070 km 2 Wadi al-Rummah basin in central Saudi Arabia as a case study, we developed an alternative, reliable, cost-effective satellite-based framework that combines empirically derived EPE thresholds, imagery-calibrated 2D hydrodynamic modeling, GRACE water-storage diagnostics, and bias-corrected CMIP6 projections to assess flood hazards and recharge potential under current and future climate scenarios in ungauged basins. The integrated approach and the resulting findings followed four key steps: (1) Identified a 22.5 mm EPE threshold, the 80th percentile of 3-day GPM/IMERG rainfall (2000–2024), aligned with flood-triggering events (Nov 2018: 23–28 mm; Apr, 2023: 42 mm); (2) Developed and calibrated a RiverFlow2D model using Sentinel-2 and PlanetScope imagery for the November 2018 flood, accurately reproducing flood depth and extent (RMSE ≤0.31 m; fuzzy-Dice ≥0.91), and estimating runoff (41 %), infiltration (25 %), and evaporation (34 %); (3) Independently validated the model with the April 2023 event (RMSE ≤0.35 m; fuzzy-Dice ≥0.86); (4) Conducted climate projections (2025–2100) from five bias-corrected NEX-GDDP CMIP6 models that revealed a 34 % increase in EPE intensity under SSP2-4.5 and 48 % under SSP5-8.5 scenarios, relative to 20th-century baselines. Our findings indicate that while intensifying extremes in the 21st century increase flood risk, the results highlight the potential for episodic recharge if effective retention strategies are employed, and offer a transferable model for climate-informed planning in data-scarce arid regions.

CMIP6↗

Advancing stream temperature prediction with a generalizable large-sample framework across CONUS river reaches

Accurately predicting stream temperature in ungauged basins remains a critical challenge for water resource management, thermoelectric power plant cooling, and ecosystem conservation. Large-sample machine learning models trained on hundreds of well-monitored river basins have shown remarkable performance; however, such models have yet to be developed solely using forcing data that can be readily extracted to simulate stream temperatures anywhere in the contiguous United States (CONUS). In this study, we present a scalable, large-sample deep learning framework using Long Short-Term Memory (LSTM) networks to simulate daily stream temperatures in ungauged basins across the CONUS. The framework leverages both modeled reanalysis of meteorological and streamflow inputs as well as static attributes available for all 2.7 million CONUS river reaches in the National Hydrography Dataset Plus (NHDPlusV2). By generating dynamical inputs from predefined thermally relevant upstream contributing areas, rather than the entire upstream basin, the model also offers improvements in very large basins where full-basin averaging can dilute the most important influences on stream temperature. Evaluated across 300 basins, the model achieves a median Mean Absolute Error (MAE) of 1.1 °C and a Nash-Sutcliffe Efficiency (NSE) of 0.95 on temporally and spatially distinct test folds—comparable to models trained exclusively using meteorological and streamflow observational data. The flexible, high-performing framework generalizes to any unmonitored river reach without significant regulation or unnatural thermal input immediately upstream, substantially expanding predictive capabilities in data-scarce regions.

Hydrology↗

Why Firn Quakes

Snow dampens sounds, but anecdotal reports concisely describe audible propagating collapse events—firnquakes—in Antarctic and Arctic snowfields. We propose combining granular and continuum mechanics to form a testable theory for conditioning, triggering, and propagation of firnquakes consistent with scarce data. A central condition for collapse events is unconsolidated firn at depth. As firn grains compact, stresses are transmitted along force chains which carry the overburden and transition into a continuous medium by pressure sintering. This granular legacy creates solid-like supports of denser layers that keep the material below unconsolidated. Dynamic amplification triggers local brittle failure of the supports, which induces a cascade of collapse propagation. Using bulk density from ice cores as proxy for stiffness, we find the flexural wave speed by collapsing supports matches the recorded firnquake velocities on the order of 100 m/s. Our theory is to be tested in firn sheets and other compacting granular materials.

Voigtländer, A. [Lawrence Berkeley National Labora↗

Buffering the impacts of extreme climate variability in the highly engineered Tigris Euphrates river system

More extreme and prolonged floods and droughts, commonly attributed to global warming, are affecting the livelihood of major sectors of the world’s population in many basins worldwide. While these events could introduce devastating socioeconomic impacts, highly engineered systems are better prepared for modulating these extreme climatic variabilities. Herein, we provide methodologies to assess the effectiveness of reservoirs in managing extreme floods and droughts and modulating their impacts in data-scarce river basins. Our analysis of multiple satellite missions and global land surface models over the Tigris-Euphrates Watershed (TEW; 30 dams; storage capacity: 250 km 3 ), showed a prolonged (2007–2018) and intense drought (Average Annual Precipitation [AAP]: < 400 km 3 ) with no parallels in the past 100 years (AAP during 1920–2020: 538 km 3 ) followed by 1-in-100-year extensive precipitation event (726 km 3 ) and an impressive recovery (113 ± 11 km 3 ) in 2019 amounting to 50% of losses endured during drought years. Dam reservoirs captured water equivalent to 40% of those losses in that year. Additional studies are required to investigate whether similar highly engineered watersheds with multi-year, high storage capacity can potentially modulate the impact of projected global warming-related increases in the frequency and intensity of extreme rainfall and drought events in the twenty-first century.

54 ENVIRONMENTAL SCIENCES↗

Differentiable modelling to unify machine learning and physical models for geosciences

Process-based modelling offers interpretability and physical consistency in many domains of geosciences but struggles to leverage large datasets efficiently. Machine-learning methods, especially deep networks, have strong predictive skills yet are unable to answer specific scientific questions. Here, in this Perspective, we explore differentiable modelling as a pathway to dissolve the perceived barrier between process-based modelling and machine learning in the geosciences and demonstrate its potential with examples from hydrological modelling. ‘Differentiable’ refers to accurately and efficiently calculating gradients with respect to model variables or parameters, enabling the discovery of high-dimensional unknown relationships. Differentiable modelling involves connecting (flexible amounts of) prior physical knowledge to neural networks, pushing the boundary of physics-informed machine learning. It offers better interpretability, generalizability, and extrapolation capabilities than purely data-driven machine learning, achieving a similar level of accuracy while requiring less training data. Additionally, the performance and efficiency of differentiable models scale well with increasing data volumes. Under data-scarce scenarios, differentiable models have outperformed machine-learning models in producing short-term dynamics and decadal-scale trends owing to the imposed physical constraints. Differentiable modelling approaches are primed to enable geoscientists to ask questions, test hypotheses, and discover unrecognized physical relationships. Future work should address computational challenges, reduce uncertainty, and verify the physical significance of outputs.

58 GEOSCIENCES↗

Towards the understanding of the genuine three-body interaction for p–p–p and p–p–Λ

Three-body nuclear forces play an important role in the structure of nuclei and hypernuclei and are also incorporated in models to describe the dynamics of dense baryonic matter, such as in neutron stars. So far, only indirect measurements anchored to the binding energies of nuclei can be used to constrain the three-nucleon force, and if hyperons are considered, the scarce data on hypernuclei impose only weak constraints on the three-body forces. In this work, we present the first direct measurement of the p–p–p and p–p–Λ systems in terms of three-particle correlation functions carried out for pp collisions at $\sqrt{s}$=13 TeV. Three-particle cumulants are extracted from the correlation functions by applying the Kubo formalism, where the three-particle interaction contribution to these correlations can be isolated after subtracting the known two-body interaction terms. A negative cumulant is found for the p–p–p system, hinting to the presence of a residual three-body effect while for p–p–Λ the cumulant is consistent with zero. This measurement demonstrates the accessibility of three-baryon correlations at the LHC.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗