Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “limited data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Evaluation of hydrograph separation techniques with uncertain end-member composition

Hydrograph separation is one of many approaches used to analyse shifts in source water contributions to stream flow resulting from climate change in remote watersheds. Understanding these shifts is vital, as shifts in source water contributions to a stream can shape water management decisions. Because remote watersheds are often inaccessible and have poorly characterized contributing water sources, or end-members, it is critical to understand the implications of using different hydrograph separation techniques in these data-limited environments. To explore the uncertainty associated with different techniques, results from two hydrograph separation techniques, mass balance and principle component analysis, were compared using 3 years of aqueous geochemical data from the East River watershed located in the Elk Mountains of Central Colorado. Solute concentrations of the end-members were characterized by both a limited set of direct chemical measurements of different sources and detailed seasonal instream chemistry to examine the influences of uncertain end-member compositions in a data-limited environment. Annual volumetric end-member contributions to stream flow had relatively good agreement across separation techniques. Large variations in time were observed in the hydrograph separations, depending on the end-member type, and estimated flow contributions varied between the selected solutes. End-member concentrations characterized by stream chemistry showed several limitations including a reduced number of distinguishable end-members and differences in timing of flow contributions. Here the results highlight the benefits of using multiple hydrograph separation techniques by providing a ‘weight-of-evidence’ approach to environments with limited end-member concentration data.

54 ENVIRONMENTAL SCIENCES↗

Remotely Sensed High‐Resolution Soil Moisture and Evapotranspiration: Bridging the Gap Between Science and Society

This paper reviews the current state of high‐resolution remotely sensed soil moisture (SM) and evapotranspiration (ET) products and modeling, and the coupling relationship between SM and ET. SM downscaling approaches for satellite passive microwave products leverage advances in artificial intelligence and high‐resolution remote sensing using visible, near‐infrared, thermal‐infrared, and synthetic aperture radar sensors. Remotely sensed ET continues to advance in spatiotemporal resolutions from MODIS to ECOSTRESS to Hydrosat and beyond. These advances enable a new understanding of bio‐geo‐physical controls and coupled feedback mechanisms between SM and ET reflecting the land cover and land use at field scale (3–30 m, daily). Still, the state‐of‐the‐science products have their challenges and limitations, which we detail across data, retrieval algorithms, and applications. We describe the roles of these data in advancing 10 application areas: drought assessment, food security, precision agriculture, soil salinization, wildfire modeling, dust monitoring, flood forecasting, urban water, energy, and ecosystem management, ecohydrology, and biodiversity conservation. We discuss that future scientific advancement should focus on developing open‐access, high‐resolution (3–30 m), sub‐daily SM and ET products, enabling the evaluation of hydrological processes at finer scales and revolutionizing the societal applications in data‐limited regions of the world, especially the Global South for socio‐economic development.

54 ENVIRONMENTAL SCIENCES↗

The Global Spectra-Trait Initiative: A database of paired leaf spectroscopy and functional traits associated with leaf photosynthetic capacity

Accurate assessment of leaf functional traits is crucial for a diverse range of applications from crop phenotyping to parameterizing global climate models. Leaf reflectance spectroscopy offers a promising avenue to advance ecological and agricultural research by complementing traditional, time-consuming gas exchange measurements. However, the development of robust hyperspectral models for predicting leaf photosynthetic capacity and associated traits from reflectance data has been hindered by limited data availability across species and environments. Here we introduce the Global Spectra-Trait Initiative (GSTI), a collaborative repository of paired leaf hyperspectral and gas exchange measurements from diverse ecosystems. The GSTI repository currently encompasses over 7500 observations from 397 species and 41 sites gathered from 36 published and unpublished studies, thereby offering a key resource for developing and validating hyperspectral models of leaf photosynthetic capacity. The GSTI database is developed on GitHub (https://github.com/plantphys/gsti, last access: 4 January 2026) and published to ESS-DIVE https://doi.org/10.15485/2530733, Lamour et al., 2025). It includes gas exchange data, derived photosynthetic parameters, and key leaf traits often associated with traditional gas exchange measurements such as leaf mass per area and leaf elemental composition. By providing a standardized repository for data sharing and analysis, we present a critical step towards creating hyperspectral models for predicting photosynthetic traits and associated leaf traits for terrestrial plants.

Lamour, Julien [Université of Toulouse (France); U↗

Basin-Scale Structural Features Database

The Basin-Scale Structural Features database provides spatial datasets of faults, fractures, folds, and earthquakes compiled from public, authoritative sources (e.g., U.S. Geological Survey and State Geological Surveys) and aggregated into derivative forms to support subsurface assessments. Recognizing that characterizing basin-scale structural features requires interpreting data that are often ambiguous or lack key information, the source data were evaluated using a knowledge-data framework and geospatial fuzzy logic method (Justman et al., 2020) to represent both measured (observed) and predicted (inferred or potential) structural features as derivative datasets. This workflow employs conceptual models for known structural features and predicted structural features, incorporating geospatial data to estimate potential, even with limited data. The aim is to aid and support an understanding of basin-scale features and identify potential gaps in data and knowledge. As of 4/30/2025, the database includes resources for nine sedimentary basins: Appalachian, Denver, U.S. Gulf Coast, Illinois, Michigan, Permian, Sacramento, San Joquin and Williston. The database is organized by basin and then data category: 1) Faults, fractures, folds, 2) Earthquakes, 3) Topographic, 4) Structural contours and isopachs, 5) Geophysical, and 6) Structural feature density assessment maps.

basin scale↗

Evaluating the limitations of Bayesian metabolic control analysis

AbstractBayesian Metabolic Control Analysis (BMCA) has emerged as a promising framework for inferring metabolic control coefficients in data-limited scenarios by integrating Bayesian inference with linlog rate laws. However, its predictive accuracy and limitations remain underexplored. This study systematically evaluates BMCA’s ability to infer elasticity values, flux control coefficients (FCCs), and concentration control coefficients (CCCs) under varying data availability conditions using three synthetic metabolic network models. Our findings highlight the strengths and weaknesses of BMCA, guiding its application in metabolic engineering and emphasizing the need for methodological refinements.Author summaryUnderstanding how enzymes control metabolic pathways is crucial for optimizing biomanufacturing and synthetic biology applications. Bayesian Metabolic Control Analysis (BMCA) is a promising computational method that integrates Bayesian inference with metabolic control analysis to estimate key control parameters, even in cases with limited experimental data. However, the accuracy and limitations of BMCA remain unclear. In this study, we systematically evaluate BMCA using three synthetic metabolic networks to determine how different types of physiological data impact its predictive performance. We find that BMCA requires flux and enzyme concentration data for accurate predictions, while external metabolite concentrations contribute little. Additionally, BMCA fails to predict elasticity values beyond a magnitude of 1.5 and reliably infer allosteric regulation, even when strong regulatory interactions exist. In addition, BMCA does not accurately rank metabolic control points, which may limit its utility in identifying key enzymes in engineered pathways. Our work provides practical insights into when and how BMCA can be applied, guiding future research in metabolic modeling and control analysis.

Shin, Janis (ORCID:0000000216572455)↗

Coarse-to-fine Task-driven Inpainting for Geoscience Images

The processing and recognition of geoscience images have wide applications. Most of existing researches focus on understanding the high-quality geoscience images by assuming that all the images are clear. However, in many real-world cases, the geoscience images might contain occlusions during the image acquisition. This problem actually implies the image inpainting problem in computer vision and multimedia. As far as we know, all the existing image inpainting algorithms learn to repair the occluded regions for a better visualization quality, they are excellent for natural images but not good enough for geoscience images, and they never consider the following geoscience task when developing inpainting methods. Here, this paper aims to repair the occluded regions for a better geoscience task performance and advanced visualization quality simultaneously, without changing the current deployed deep learning based geoscience models. Because of the complex context of geoscience images, we propose a coarse-to-fine encoder-decoder network with the help of designed coarse-to-fine adversarial context discriminators to reconstruct the occluded image regions. Due to the limited data of geoscience images, we propose a MaskMix based data augmentation method, which augments inpainting masks instead of augmenting original images, to exploit the limited geoscience image data. The experimental results on three public geoscience datasets for remote sensing scene recognition, cross-view geolocation and semantic segmentation tasks respectively show the effectiveness and accuracy of the proposed method. The code is available at: https://github.com/HMS97/Task-driven-Inpainting.

97 MATHEMATICS AND COMPUTING↗

An open-access simulated earthquake ground-motion database for an M7 Hayward Fault earthquake in the San Francisco Bay Region

Comprehensive understanding of earthquake ground motions, particularly in the near-fault region of large-magnitude events, is limited by gaps in strong-motion data. This challenge is prominent in areas with high seismic hazard but infrequent large earthquakes where data is sparse and difficult to interpret. These data limitations lead to uncertainties in the development of site-specific ground motions, which are crucial for engineering risk assessments. To address these challenges, physics-based regional-scale ground-motion simulations have been developed. With the emergence of exaflop-scale computing ecosystems, it is now possible to simulate regional earthquake processes at unprecedented fidelity and generate the large number of fault rupture realizations necessary to characterize both intra- and inter-event ground-motion variability. This article introduces a new database of simulated earthquake ground motions, created for applications in earthquake engineering, earthquake planning, and emergency response. The inaugural version of the database features simulated ground motions for a magnitude 7 Hayward Fault earthquake in the San Francisco Bay Region (SFBR), using the EarthQuake SIMulation (EQSIM) simulation framework and the Graves–Pitarka kinematic rupture model. The aim is to provide high-fidelity, spatially dense, three-component motions generated on the Department of Energy’s (DOE) newest generation of graphics processing unit (GPU)-accelerated supercomputers. These motions are being made openly available to the engineering, scientific, and disaster planning communities. In addition, this work develops protocols for the efficient dissemination of these large data sets and emphasizes community engagement to build confidence in their application. This article discusses the methodology behind the data, underlying software verification and validation, scalable data management, and a user interface for data access. The goal is to facilitate widespread use and elicit expert feedback to maximize the utility and exploitation of simulated motions. While the initial focus is on the San Francisco Region, simulations for additional regions will be added as the DOE program progresses.

Simulated ground-motion database↗

Uncertainty quantification in scientific machine learning: Methods, metrics, and comparisons

Neural networks (NNs) are currently changing the computational paradigm on how to combine data with mathematical laws in physics and engineering in a profound way, tackling challenging inverse and ill-posed problems not solvable with traditional methods. However, quantifying errors and uncertainties in NN-based inference is more complicated than in traditional methods. This is because in addition to aleatoric uncertainty associated with noisy data, there is also uncertainty due to limited data, but also due to NN hyperparameters, overparametrization, optimization and sampling errors as well as model misspecification. Although there are some recent works on uncertainty quantification (UQ) in NNs, there is no systematic investigation of suitable methods towards quantifying the total uncertainty effectively and efficiently even for function approximation, and there is even less work on solving partial differential equations and learning operator mappings between infinite-dimensional function spaces using NNs. In this work, we present a comprehensive framework that includes uncertainty modeling, new and existing solution methods, as well as evaluation metrics and post-hoc improvement approaches. Further, to demonstrate the applicability and reliability of our framework, we present an extensive comparative study in which various methods are tested on prototype problems, including problems with mixed input-output data, and stochastic problems in high dimensions. In the Appendix, we include a comprehensive description of all the UQ methods employed. Further, to help facilitate the deployment of UQ in Scientific Machine Learning research and practice, we present and develop in [1] an open-source Python library (github.com/Crunch-UQ4MI/neuraluq), termed NeuralUQ, that is accompanied by an educational tutorial and additional computational experiments.

11 physics-informed neural networks↗

Synchro-Waveform-Based Event Identification Using Multi-Task Time-Frequency Transform Networks

Influenced by the transient dynamics and reduced inertia characteristics of high-penetration renewable energy systems, power system events frequently exhibit distinct characteristics such as high-frequency components including wide-band oscillations and hyper-harmonics. This makes standard systems face challenges including significant latency and reduced accuracy due to limited data resolution. However, current methods face significant limitations, including insufficient pattern capture ability, low noise immunity, limited feature learning, and restricted localization capabilities, thereby hindering real-time performance. To tackle this issue, this paper proposed a novel synchro-waveform-based event identification approach via a Multi-task Time-frequency Transform Network (MTTNet). Initially, a Time-frequency Transform Block (TTB) is developed to extract both local and global information. The TTB leverages both Fourier and S-transforms to derive comprehensive time-frequency information from synchro-waveforms. Subsequently, a multi-task learning strategy is employed to identify the type and distinguish localization of events. Integrating the TTB and multi-task learning, the MTTNet is designed for synchro-waveform-based event identification, incorporating an adaptive weighting strategy and simplified computation for the S-transform. Two different datasets, comprising simulated and actual synchro-waveforms, are collected from the IEEE 123 bus system and a real-world high-penetration renewable energy system using a universal grid analyzer. Extensive experiments on various conditions are carried out. In conclusion, results demonstrated that the MTTNet consistently surpasses both basic and advanced baselines, with maximum improvements of 13.24% and 9.86%, respectively, while reducing the calculation burden by 15-19 times to achieve real-time event identification.

Event identification↗

Jobs, jobs, jobs: what’s an analyst to do?

Analysts and economists often face the task of using employment metrics to characterize industries of interest. Some key challenges can be understanding where to find employment metrics, the differences in various employment metrics, and when each metric should be used. This article analyzes a variety of publicly available employment data for the United States and compares these data. A detailed description of the intricacies of each data source is provided, which covers factors such as regionality, industry breakout, periodicity, and the types of jobs included. This article provides several case study examples, using the oil and gas extraction, coal mining, and chemical manufacturing sectors to portray challenges data users may face when developing employment estimates that suit their needs. Data users should be aware of a variety of data sources to understand alternative analysis options when data limitations are present and to determine which data source best meets their needs. Instances may occur in which information from one dataset may be used to help impute missing values.

99 GENERAL AND MISCELLANEOUS↗

An Application of a Modified Beta Factor Method for the Analysis of Software Common Cause Failures

This paper presents an approach for modeling software common cause failures (CCFs) within digital instrumentation and control (I&C) systems. CCFs consist of a concurrent failure between two or more components due to a shared failure cause and coupling mechanism. This work emphasizes the importance of identifying software-centric attributes related to the coupling mechanisms necessary for simultaneous failures of redundant software components. The groups of components that share coupling mechanisms are called common cause component groups (CCCGs). Most CCF models rely on operational data as the basis for establishing CCCG parameters and predicting CCFs. This work is motivated by two primary concerns: (1) a lack of operational and CCF data for estimating software CCF model parameters; and (2) the need to model single components as part of multiple CCCGs simultaneously. A hybrid approach was developed to account for these concerns by leveraging existing techniques: a modified beta factor model allows single components to be placed within multiple CCCGs, while a second technique provides software-specific model parameters for each CCCG. This hybrid approach provides a means to overcome the limitations of conventional methods while offering support for design decisions under the limited data scenario.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

AI-Driven Crack Detection for Remanufacturing Cylinder Heads Using Deep Learning and Engineering-Informed Data Augmentation

Detecting cracks in cylinder heads traditionally relies on manual inspection, which is time-consuming and susceptible to human error. As an alternative, automated object detection utilizing computer vision and machine learning models has been explored. However, these methods often face challenges due to a lack of sufficiently annotated training data, limited image diversity, and the inherently small size of cracks. Addressing these constraints, this paper introduces a novel automated crack-detection method that enhances data availability through a synthetic data generation technique. Unlike general data augmentation practices, our method involves copying cracks from one location to another, guided by both random and informed engineering decisions about likely crack formations due to cyclic thermomechanical loads. The innovative aspect of our approach lies in the integration of domain-specific engineering knowledge into the synthetic generation process, which substantially improves detection accuracy. We evaluate our method’s effectiveness using two metrics: the F2 score, which emphasizes recall to prioritize detecting all potential cracks, and mean average precision (MAP), a standard measure in object detection. Experimental results demonstrate that, without engineering insights, our method increases the F2 score from 0.40 to 0.65, while maintaining a stable MAP. Incorporating detailed engineering knowledge further enhances the F2 score to 0.70 and improves MAP to 0.57, representing increases of 63% and 43%, respectively. These results confirm that our approach not only mitigates the limitations of traditional data augmentation but also significantly advances the reliability and precision of crack detection in industrial settings.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Bi-fidelity variational auto-encoder for uncertainty quantification

Quantifying the uncertainty of quantities of interest (QoIs) from physical systems is a primary objective in model validation. However, achieving this goal entails balancing the need for computational efficiency with the requirement for numerical accuracy. To address this trade-off, we propose a novel bi-fidelity formulation of variational auto-encoders (BF-VAE) designed to estimate the uncertainty associated with a QoI from low-fidelity (LF) and high-fidelity (HF) samples of the QoI. Here, this model allows for the approximation of the statistics of the HF QoI by leveraging information derived from its LF counterpart. Specifically, we design a bi-fidelity auto-regressive model in the latent space which is integrated within the VAE’s probabilistic encoder–decoder structure. An effective algorithm is proposed to maximize the variational lower bound of the HF log-likelihood in the presence of limited HF data, resulting in the synthesis of HF realizations with a reduced computational cost. Additionally, we introduce the concept of the bi-fidelity information bottleneck (BF-IB) to provide an information-theoretic interpretation of the proposed BF-VAE model. Our numerical results demonstrate that the BF-VAE leads to considerably improved accuracy, as compared to a VAE trained using only HF data, when limited HF data is available.

42 ENGINEERING↗

Multitask Machine Learning of Collective Variables for Enhanced Sampling of Rare Events

Computing accurate reaction rates is a central challenge in computational chemistry and biology because of the high cost of free energy estimation with unbiased molecular dynamics. In this work, a data-driven machine learning algorithm is devised to learn collective variables with a multitask neural network, where a common upstream part reduces the high dimensionality of atomic configurations to a low dimensional latent space and separate downstream parts map the latent space to predictions of basin class labels and potential energies. Here, the resulting latent space is shown to be an effective low-dimensional representation, capturing the reaction progress and guiding effective umbrella sampling to obtain accurate free energy landscapes. This approach is successfully applied to model systems including a 5D Müller Brown model, a 5D three-well model, the alanine dipeptide in vacuum, and an Au(110) surface reconstruction unit reaction. It enables automated dimensionality reduction for energy controlled reactions in complex systems, offers a unified and data-efficient framework that can be trained with limited data, and outperforms single-task learning approaches, including autoencoders.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

State Estimation for Distribution Networks with Asynchronous Sensors Using Stochastic Descent: Preprint

This paper investigates the problem of state estimation for distribution networks with asynchronous sensors comprising of a mix of smart meters and phasor measurement units (PMUs) with multiple sampling and reporting rates. We consider two independent scenarios of state estimation and tracking, with either voltages or currents as states. With these two sets, we investigate estimation under (a) full data, assuming all measurements are available and (b) limited data, where an online algorithmic approach is adopted to estimate the possibly time-varying states by processing measurements as and when available. The proposed algorithm, inspired by the classical Stochastic Gradient Descent (SGD) approach updates the states based on the previous estimate and the newly available measurements. Finally, we demonstrate the estimation and tracking efficacy through numerical simulations on the IEEE-37 test network, while also highlighting how estimation with currents as states leads to faster convergence.

asynchronous sensors↗

Deriving Stable Peak Models to Fit Complex XPS Data From Cu Contaminated Pt Electrocatalysts

X-ray Photoelectron Spectroscopy spectra peak models, designed to partition photoemission signals emanating from different elements or chemical states within an atom, are fitted to data limited to an energy interval over which inelastically scattered photoemission signal can be estimated. While the choice of background approximation and line shapes of components to the peak model requires careful consideration, the energy interval used to define the data to which the peak model is optimized has a significant impact on the final peak model. The relationship between the background intensity and data intensity at the start and end of the energy interval dictates the line shapes used in the peak model. In this work, we devise a method to peak fit a complex overlapping Cu 3p and Pt 4f XPS peak structure to perform the elemental quantification. We first use an Al 2s peak to illustrate how background curves approach data at the limits of the energy interval over which the background is defined, influencing the analysis of XPS spectra. Next, we demonstrate the nature of interactions between specific line shapes (Voigt and pseudo-Voigt profiles) suitable for photoemission peaks and a specific background curve (Shirley) and a peak model is presented that includes components to the peak model that accommodates background intensity during fitting of the peak model to data. The peak model allowed for quantification of the contributions of Pt 4f peaks emanating from the substrate that exhibits strong asymmetry in the presence of the inhomogeneously distributed Cu species, mostly of Lorentzian character.

XPS↗

Graphene SHDMC Data

The data used to produce all of the figures and tables in the manuscript titled "Highly Accurate Many-Body Theory Reaches 2D Materials" can be found here. This data set includes: -SHDMC results for graphene -selected CI with and without re-normalized second-order perturbation (rPT2) theory corrections for graphene -data demonstrating that SHDMC displays an exponential rate of convergence -data used for sCI + rPT2 complete basis set extrapolation -data used to extrapolate SHDMC energies to the infinite basis limit -data used to demonstrate compactness of SHDMC wavefunction

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Natural Language Processing-Enhanced Nuclear Industry Operating Experience Data Analysis: Aggregation and Interpretation of Multi-Report Analysis Results

Industry-wide operating experience is a critical source of raw data for reliability and risk model parameter estimations for nuclear power plants. A large portion of operating experience data are failure events stored as reports that contain unstructured data, such as narratives. In current practice, a failure report is usually reviewed and manually coded by analysts. The coding is based on extracting several event characteristics such as system name, component type, sub-part type, failure mode, and failure cause. Event narratives are mostly used to help understand events and extract their characteristics. In this line of research, we aim to maximize the usage of event narratives by leveraging natural language processing (NLP) methods to automatically convert an event narrative to a causal graph. This research has promise to improve physical understanding of failure initiation and propagation and to facilitate use of non-failure data (e.g., near-misses and degradations) to complement the limited data pool of failures. In our previous work, we developed an NLP tool and applied it to analyze a number of licensee event reports submitted by U.S. nuclear power plants to the Nuclear Regulatory Commission. In this paper, we will report our recent research progress in aggregating the results of multiple reports, developing network model(s), and drawing statistical insights.

99 GENERAL AND MISCELLANEOUS↗