Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “imputation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

87 records · Page 5

Physics-Informed Deep Learning for Reconstruction of Spatial Missing Climate Information in the Antarctic

Understanding the influence of the Antarctic on the global climate is crucial for the prediction of global warming. However, due to very few observation sites, it is difficult to reconstruct the rational spatial pattern by filling in the missing values from the limited site observations. To tackle this challenge, regional spatial gap-filling methods, such as Kriging and inverse distance weighted (IDW), are regularly used in geoscience. Nevertheless, the reconstructing credibility of these methods is undesirable when the spatial structure has massive missing pieces. Inspired by image inpainting, we propose a novel deep learning method that demonstrates a good effect by embedding the physics-aware initialization of deep learning methods for rapid learning and capturing the spatial dependence for the high-fidelity imputation of missing areas. We create the benchmark dataset that artificially masks the Antarctic region with ratios of 30%, 50% and 70%. The reconstructing monthly mean surface temperature using the deep learning image inpainting method RFR (Recurrent Feature Reasoning) exhibits an average of 63% and 71% improvement of accuracy over Kriging and IDW under different missing rates. With regard to wind speed, there are still 36% and 50% improvements. In particular, the achieved improvement is even better for the larger missing ratio, such as under the 70% missing rate, where the accuracy of RFR is 68% and 74% higher than Kriging and IDW for temperature and also 38% and 46% higher for wind speed. In addition, the PI-RFR (Physics-Informed Recurrent Feature Reasoning) method we proposed is initialized using the spatial pattern data simulated by the numerical climate model instead of the unified average. Compared with RFR, PI-RFR has an average accuracy improvement of 10% for temperature and 9% for wind speed. When applied to reconstruct the spatial pattern based on the Antarctic site observations, where the missing rate is over 90%, the proposed method exhibits more spatial characteristics than Kriging and IDW.

54 ENVIRONMENTAL SCIENCES↗

STAIR 2.0: A Generic and Automatic Algorithm to Fuse Modis, Landsat, and Sentinel-2 to Generate 10 m, Daily, and Cloud-/Gap-Free Surface Reflectance Product

Remote sensing datasets with both high spatial and high temporal resolution are critical for monitoring and modeling the dynamics of land surfaces. However, no current satellite sensor could simultaneously achieve both high spatial resolution and high revisiting frequency. Therefore, the integration of different sources of satellite data to produce a fusion product has become a popular solution to address this challenge. Many methods have been proposed to generate synthetic images with rich spatial details and high temporal frequency by combining two types of satellite datasets—usually frequent coarse-resolution images (e.g., MODIS) and sparse fine-resolution images (e.g., Landsat). In this paper, we introduce STAIR 2.0, a new fusion method that extends the previous STAIR fusion framework, to fuse three types of satellite datasets, including MODIS, Landsat, and Sentinel-2. In STAIR 2.0, input images are first processed to impute missing-value pixels that are due to clouds or sensor mechanical issues using a gap-filling algorithm. The multiple refined time series are then integrated stepwisely, from coarse- to fine- and high-resolution, ultimately providing a synthetic daily, high-resolution surface reflectance observations. We applied STAIR 2.0 to generate a 10-m, daily, cloud-/gap-free time series that covers the 2017 growing season of Saunders County, Nebraska. Moreover, the framework is generic and can be extended to integrate more types of satellite data sources, further improving the quality of the fusion product. View Full-Text

47 OTHER INSTRUMENTATION↗

A Random Forest Approach to Identifying Young Stellar Object Candidates in the Lupus Star-forming Region

The identification and characterization of stellar members within a star-forming region are critical to many aspects of star formation, including formalization of the initial mass function, circumstellar disk evolution, and star formation history. Previous surveys of the Lupus star-forming region have identified members through infrared excess and accretion signatures. We use machine learning to identify new candidate members of Lupus based on surveys from two space-based observatories: ESA’s Gaia and NASA’s Spitzer. Astrometric measurements from Gaia's Data Release 2 and astrometric and photometric data from the Infrared Array Camera on the Spitzer Space Telescope, as well as from other surveys, are compiled into a catalog for the random forest (RF) classifier. The RF classifiers are tested to find the best features, membership list, non-membership identification scheme, imputation method, training set class weighting, and method of dealing with class imbalance within the data. We list 27 candidate members of the Lupus star-forming region for spectroscopic follow-up. Most of the candidates lie in Clouds V and VI, where only one confirmed member of Lupus was previously known. These clouds likely represent a slightly older population of star formation.

79 ASTRONOMY AND ASTROPHYSICS↗

Functional Data Analysis for Extracting the Intrinsic Dimensionality of Spectra: Application to Chemical Homogeneity in the Open Cluster M67

High-resolution spectroscopic surveys of the Milky Way have entered the Big Data regime and have opened avenues for solving outstanding questions in Galactic archeology. However, exploiting their full potential is limited by complex systematics, whose characterization has not received much attention in modern spectroscopic analyses. In this work, we present a novel method to disentangle the component of spectral data space intrinsic to the stars from that due to systematics. Using functional principal component analysis on a sample of 18,933 giant spectra from APOGEE, we find that the intrinsic structure above the level of observational uncertainties requires ≈10 functional principal components (FPCs). Our FPCs can reduce the dimensionality of spectra, remove systematics, and impute masked wavelengths, thereby enabling accurate studies of stellar populations. To demonstrate the applicability of our FPCs, we use them to infer stellar parameters and abundances of 28 giants in the open cluster M67. We employ Sequential Neural Likelihood, a simulation-based Bayesian inference method that learns likelihood functions using neural density estimators, to incorporate non-Gaussian effects in spectral likelihoods. By hierarchically combining the inferred abundances, we limit the spread of the following elements in M67: Fe ≲ 0.02 dex; C ≲ 0.03 dex; O, Mg, Si, Ni ≲ 0.04 dex; Ca ≲ 0.05 dex; N, Al ≲ 0.07 dex (at 68% confidence). Our constraints suggest a lack of self-pollution by core-collapse supernovae in M67, which has promising implications for the future of chemical tagging to understand the star formation history and dynamical evolution of the Milky Way.

79 ASTRONOMY AND ASTROPHYSICS↗

Anomaly Detection and Approximate Similarity Searches of Transients in Real-time Data Streams

Abstract We present Lightcurve Anomaly Identification and Similarity Search ( LAISS ), an automated pipeline to detect anomalous astrophysical transients in real-time data streams. We deploy our anomaly detection model on the nightly Zwicky Transient Facility (ZTF) Alert Stream via the ANTARES broker, identifying a manageable ∼1–5 candidates per night for expert vetting and coordinating follow-up observations. Our method leverages statistical light-curve and contextual host galaxy features within a random forest classifier, tagging transients of rare classes ( spectroscopic anomalies), of uncommon host galaxy environments ( contextual anomalies), and of peculiar or interaction-powered phenomena ( behavioral anomalies). Moreover, we demonstrate the power of a low-latency (∼ms) approximate similarity search method to find transient analogs with similar light-curve evolution and host galaxy environments. We use analogs for data-driven discovery, characterization, (re)classification, and imputation in retrospective and real-time searches. To date, we have identified ∼50 previously known and previously missed rare transients from real-time and retrospective searches, including but not limited to superluminous supernovae (SLSNe), tidal disruption events, SNe IIn, SNe IIb, SNe I-CSM, SNe Ia-91bg-like, SNe Ib, SNe Ic, SNe Ic-BL, and M31 novae. Lastly, we report the discovery of 325 total transients, all observed between 2018 and 2021 and absent from public catalogs (∼1% of all ZTF Astronomical Transient reports to the Transient Name Server through 2021). These methods enable a systematic approach to finding the “needle in the haystack” in large-volume data streams. Because of its integration with the ANTARES broker, LAISS is built to detect exciting transients in Rubin data.

79 ASTRONOMY AND ASTROPHYSICS↗

A graph neural network (GNN) approach to basin-scale river network learning: the role of physics-based connectivity and data fusion

Abstract. Rivers and river habitats around the world are under sustained pressure from human activities and the changing global environment. Our ability to quantify and manage the river states in a timely manner is critical for protecting the public safety and natural resources. In recent years, vector-based river network models have enabled modeling of large river basins at increasingly fine resolutions, but are computationally demanding. This work presents a multistage, physics-guided, graph neural network (GNN) approach for basin-scale river network learning and streamflow forecasting. During training, we train a GNN model to approximate outputs of a high-resolution vector-based river network model; we then fine-tune the pretrained GNN model with streamflow observations. We further apply a graph-based, data-fusion step to correct prediction biases. The GNN-based framework is first demonstrated over a snow-dominated watershed in the western United States. A series of experiments are performed to test different training and imputation strategies. Results show that the trained GNN model can effectively serve as a surrogate of the process-based model with high accuracy, with median Kling–Gupta efficiency (KGE) greater than 0.97. Application of the graph-based data fusion further reduces mismatch between the GNN model and observations, with as much as 50 % KGE improvement over some cross-validation gages. To improve scalability, a graph-coarsening procedure is introduced and is demonstrated over a much larger basin. Results show that graph coarsening achieves comparable prediction skills at only a fraction of training cost, thus providing important insights into the degree of physical realism needed for developing large-scale GNN-based river network models.

54 ENVIRONMENTAL SCIENCES↗

Improved Particle Heat Transfer by way of Bimodal Particle Distributions for High Temperature Solar Thermal Energy

High temperature solar thermal facilities are looking to increase operating temperatures through novel heat transfer media, one such being solid particles. These particles operating at high temperatures will require transferring their thermal energy into another working fluid like supercritical carbon dioxide which can be used in advanced power cycles. Achieving high heat transfer between the particles and supercritical carbon dioxide is essential to high efficiency and low-cost operation. Therefore, optimizing the thermal conductivity of these particles is one potential way to ensure high performance. Traditionally, unimodal particle distributions have been employed in high temperature particle solar power plants. However, ambient temperature testing of bimodal particle distributions has revealed a superior thermal conductivity when compared to its unimodal counterpart at the same temperature. This data was obtained by certified, off-the-shelf instruments that can effectively simulate the conditions a particle would be exposed to in a high temperature solar thermal system. Data obtained in this way suggests that the increased thermal conductively imputed by a bimodal particle distribution is significant at working temperatures in solar facilities. Furthermore, the thermal conductivity of these bimodal particle distributions peaks when the best combination of large and small particles is applied. At high temperatures, binary particle distributions are compared to monodispersed distributions of larger particles where heat transfer is more prolific due to the increased surface radiation. Various thermal conductivity, porosity and heat exchanger models are explored in conjunction with data acquired up to 700 C.

Stout, Dallin (ORCID:0009000294586091)↗

AUTONOMIE VID

Autonomie Vehicle Information Database (VID) offers a comprehensive list of vehicle specifications since 1990. The database details more than 65,000 vehicles with hundreds of attributes. The database is the result of the development of a general automated data collection framework as well as the development of building blocks for processing, cleaning, integrating and analyzing complex data. The data has undergone several layers of outlier detections processes, machine learning based imputations methods have been used to deal with missing data problems, and new fields have been created according to the rules of feature engineering. Thanks to this streamlined data pipelines, the resulting processed aggregated data should deliver a unique level of information to the user in which the content can be efficiently maintained and updated.

Moswd, Ayman↗

Fission Product Gas Monitoring During Fuel Drying Operations - 20094

The goal of the project was to devise and implement a method of accurately qualifying any noble gas emitted during fuel drying operations (part of the dry fuel storage process) in order to provide the data needed to validate the off-site dose calculations. The imputes of the project was the NRC publishing Information Notice 18-01, 'Noble Fission Gas Releases During Spent Fuel Cask Loading Operations' in February 2018. The information notice describes some industry events that occurred during vacuum drying operations as part of a dry fuel campaign. The solution was devised by Fort Calhoun Station staff in partnership with Mirion Technologies technical staff. The monitoring was accomplished by directing all the effluent from the fuel storage cask vacuum drying machine through a sample chamber containing a compact ISOCS characterized CZT based gamma spectrometer connected to a Mirion Data Analyst. This resulted near real-time quantification of the radionuclides of concern. (authors)

07 ISOTOPE AND RADIATION SOURCES↗

Design Choices in Anomaly Detection for Industrial Control Systems: Insights from Gas Pipeline Data

Industrial control systems (ICS) remain vulnerable to increasingly sophisticated cyberattacks, yet evaluating anomaly detection models in these environments is challenging due to temporal dependencies, missing-not-at-random patterns, and extremely imbalanced datasets. These factors make common practices—especially random data splits and naïve imputation—prone to severe temporal leakage, which can inflate reported performance and obscure real-world limitations. In this work, we systematically examine classical machine learning models, temporal deep learning architecture, and tensor-decomposition–based methods on a gas-pipeline dataset using a fully temporally separated evaluation pipeline designed to mimic realistic deployment conditions. Our findings show that proper temporal handling and MNAR-aware preprocessing significantly alter the relative performance of popular anomaly-detection methods, providing practical guidance for designing reliable, leakage-resistant ICS intrusion-detection systems.

97 MATHEMATICS AND COMPUTING↗

Reliable machine prognostic health management in the presence of missing data

Prognostics and health management enables the prediction of future degradation and remaining useful life (RUL) for in-service systems based on historical and contemporary data, showing promise for many practical applications. One major challenge for prognostics is the common occurrence of missing values in time-series data, often caused by disruptions in sensor communication or hardware/software failures. Another major concern is that the sufficient prior knowledge of critical component degradation with a clear failure threshold is often not readily available in practice. These issues can significantly hinder the application of advanced signal and data analysis methods and consequently degrade the health management performance. In this article, we propose a novel data-driven framework that is capable of providing accurate and reliable predictions of degradation and RUL. In this approach, one-hot health state indicators are appended to the historical time series so that the model learns end-of-life automatically. A modified gate recurrent unit based variational autoencoder is employed in generative adversarial networks to model the temporal irregularity of the incomplete time series. Furthermore, experiments on multivariate time-series datasets collected from real-world aeroengines verify that significant performance improvement can be achieved using the proposed model for robust long-term prognostics.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Gap-filling eddy covariance methane fluxes: Comparison of machine learning model predictions and uncertainties at FLUXNET-CH4 wetlands

Time series of methane fluxes measured by eddy-covariance require gap-filling to estimate annual emissions. Gap-filling methane fluxes is challenging because of high variability and complex responses to multiple drivers. To date, there is no widely established gap-filling standard for methane, with regards both to the best model algorithms and predictors. In this study, we address the need for standardization by synthesizing results of gap-filling methods applied at 17 wetland sites spanning boreal to tropical regions including all major wetlands classes and two rice paddies. We introduce new procedures for: 1) creating realistic artificial gap scenarios, 2) training and evaluating gap-filling models without overstating performance, and 3) predicting half-hourly methane fluxes and annual emissions with robust uncertainty estimates. We tested a conventional method (marginal distribution sampling) and four machine learning algorithms - penalized linear regression, artificial neural networks, random forests, and boosted decision trees - and four predictor sets, including temporal, meteorological, ecosystem carbon and energy flux, and soil predictors. We find that the conventional method can achieve similar median performance to the machine learning models but is worse than the best machine learning models and relatively insensitive to predictor choices. Of the machine learning models, decision tree algorithms performed the best in cross-validation experiments, even with a baseline predictor set, and artificial neural networks showed comparable performance when using all predictors. Soil temperature was frequently the most important predictor whilst water table depth was important at sites with substantial water table fluctuations, highlighting the value of data on soil conditions. Raw gap-filling uncertainties from the machine learning models were underestimated and we propose a method to calibrate uncertainties to observations. Finally, we gap-fill and provide summary evaluation metrics for all 81 sites in the FLUXNET-CH4 community dataset and publicly release the python code for model development, evaluation, and uncertainty estimation.

42 ENGINEERING↗

High-dimensional data analytics in civil engineering: A review on matrix and tensor decomposition

Recent developments in sensing and monitoring techniques have led to the generation of high-dimensional data in the field of civil engineering. High-dimensional data analytics methods have thus been developed to interpret such complex data. Among the different high-dimensional data analytics techniques, matrix and tensor decomposition methods have acquired a notable interest in the civil engineering community over the past decade. Due to their unique ability to deal with highly redundant and correlated data, these methods are establishing themselves as promising and efficient tools to analyze high-dimensional data in the civil engineering arena. In this paper, high-dimensional data is referred to as a data set in which the number of features is comparable or larger than the number of observations. This review paper aims to summarize the applications of matrix and tensor decomposition methods in civil engineering over the last decade. The survey begins with a general overview of matrix and tensor decomposition followed by highlighting their significance in the field. Afterward, various applications of these high-dimensional data analytics methods in civil engineering are presented, while the advantages offered by these methods are discussed. Lastly, challenges and potential research avenues for employing matrix and tensor decomposition and future emerging trends for their novel use are highlighted.

42 ENGINEERING↗

Identifying Light-Duty Vehicle Travel from Large-Scale Multimodal Wearable GPS Data with Novelty Detection Algorithms

Identifying travel mode within travel survey data sets, especially light-duty vehicle (LDV) travel, is foundational, though nontrivial, to travel behavior analysis and fuel consumption estimation. Current travel mode detection approaches require well-sampled and balanced data sets with ground truth travel mode labels. They are rarely applied and validated on large-scale, real-world data sets, which may not satisfy the data requirements. This paper proposes an LDV travel mode detection model as a supplement to current travel mode detection methods, for the case when the training set is highly (and/or completely) unbalanced, to the extent that classical machine-learning approaches become difficult or impossible to deploy. The proposed model uses a novelty detection technique-one-class support vector machines (OCSVMs)-and a novel exhaustive feature extraction (EFE) technique on continuous time series data (i.e., Global Positioning System [GPS] speed profiles) for single-mode trip trajectories. Training and validation of the model are conducted on a large-scale, real-world data set. The proposed method accurately identifies LDV trips from a broad set of multimodal trips by leveraging a wealth of preexisting in-vehicle GPS travel data. Additional sensitivity analysis sheds light on the optimal training size, which will benefit applications limited by highly imbalanced data. The paper also discusses performance comparison with regular machine-learning approaches, the model's robustness, and the potential to extend the proposed model to multimodal prediction.

47 OTHER INSTRUMENTATION↗

Online PMU Missing Value Replacement Via Event-Participation Decomposition

We introduce a new method for online Phasor Measurement Unit (PMU) missing value replacement. Our approach allows us to decompose PMU event responses into a non-dynamic component (denoted the participation factor) that can be inferred directly from the past and a dynamic component that can be inferred directly from all other PMUs (denoted the event strength). When missing values occur, we can use these two components, which do not rely on the missing index, to estimate the correct value. The method is extremely fast and can easily be used for online applications. Furthermore, extensive testing on real power system event data reveals that our approach achieves state-of-the-art performance in terms of Mean Absolute Percent Errors (MAPEs) for PMU data dropped during event periods. Here, the method also yields an interpretable and simplified view of events for further analysis and applications. The method relies only on PMU data and does not take outside information such as network topology.

24 POWER TRANSMISSION AND DISTRIBUTION↗