Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data imputation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Long-term missing value imputation for time series data using deep neural networks

We present an approach that uses a deep learning model, in particular, a MultiLayer Perceptron, for estimating the missing values of a variable in multivariate time series data. We focus on filling a long continuous gap (e.g., multiple months of missing daily observations) rather than on individual randomly missing observations. Our proposed gap filling algorithm uses an automated method for determining the optimal MLP model architecture, thus allowing for optimal prediction performance for the given time series. We tested our approach by filling gaps of various lengths (three months to three years) in three environmental datasets with different time series characteristics, namely daily groundwater levels, daily soil moisture, and hourly Net Ecosystem Exchange. We compared the accuracy of the gap-filled values obtained with our approach to the widely used R-based time series gap filling methods ImputeTS and mtsdi. The results indicate that using an MLP for filling a large gap leads to better results, especially when the data behave nonlinearly. Thus, our approach enables the use of datasets that have a large gap in one variable, which is common in many long-term environmental monitoring observations.

97 MATHEMATICS AND COMPUTING↗

High-dimensional data analytics in civil engineering: A review on matrix and tensor decomposition

Recent developments in sensing and monitoring techniques have led to the generation of high-dimensional data in the field of civil engineering. High-dimensional data analytics methods have thus been developed to interpret such complex data. Among the different high-dimensional data analytics techniques, matrix and tensor decomposition methods have acquired a notable interest in the civil engineering community over the past decade. Due to their unique ability to deal with highly redundant and correlated data, these methods are establishing themselves as promising and efficient tools to analyze high-dimensional data in the civil engineering arena. In this paper, high-dimensional data is referred to as a data set in which the number of features is comparable or larger than the number of observations. This review paper aims to summarize the applications of matrix and tensor decomposition methods in civil engineering over the last decade. The survey begins with a general overview of matrix and tensor decomposition followed by highlighting their significance in the field. Afterward, various applications of these high-dimensional data analytics methods in civil engineering are presented, while the advantages offered by these methods are discussed. Lastly, challenges and potential research avenues for employing matrix and tensor decomposition and future emerging trends for their novel use are highlighted.

42 ENGINEERING↗

Effective Missing Value Imputation Methods for Building Monitoring Data

To understand behaviors of natural and man-made events, such as energy consumption of buildings, which accounts for 40% of energy uses in the US, we deploy automated monitoring devices to record periodic observations. However, such experimental and observation data often contains problems and irregularities that have to be cleaned up before analyses. Due to various conditions affecting sensor operations, the communication channels, recording steps, or the recording media, the recorded data might have missing values, errors, or anomalous values. An effective way to clean up these problems is to replace these missing values, errors and anomalous values with expected values, a process generally known as imputation. In this work, we survey commonly used missing value imputation techniques and compare their performance on a set of building monitoring data. To compare the different types of sensor measurements with widely varying characteristics, we use normalized root mean squared error (NRMSE) as the key metric for the effectiveness of the imputation methods. We additionally consider periodicity and run time when considering comparing methods. Through extensive testing, we find that for small gap sizes, up to 8 consecutive missing values, linear interpolation performs the best; for larger gaps stretching up to 48 consecutive missing values, K-nearest neighbors provides the most accurate imputations; for even larger gaps, more computational intensive methods, such as matrix factorization, achieve the smallest NRMSE. Additionally, we observe that these computationally intensive algorithms not only provide accurate imputations for large gaps, but are also more robust across all types of sensors.

Cho, B↗

Multi-agent voltage control in distribution systems using GAN-DRL-based approach

Active distribution grids can experience voltage fluctuations and violations due to the high penetration of variable distributed energy resources (DERs). These problems might occur because of the uncertain and variable generation natures of these resources, especially solar photovoltaic resources, during panel shadowing scenarios. Volt-VAR control (VVC) is an efficient method that controls the reactive power set-points of the inverters to regulate the voltage of distribution grids. Although several VVC approaches have been proposed recently, the performance of these approaches degrades significantly if behind-the-meter solar generation data are unobservable/missing. Therefore, it is necessary to impute missing/unobservable PV data accurately to be utilized in VVC approaches. Further, this paper proposes a model-free, data-driven, centrally trained, and decentrally executed multi-agent deep reinforcement learning-based VVC architecture to regulate the voltage of distribution networks. A generative adversarial network (GAN) is incorporated to impute the unobservable PV data accurately, which improves the performance of the proposed control architecture. The proposed multi-agent-soft-actor–critic algorithm (MASAC)-based VVC technique utilizes the actual PV dataset as well as the imputed dataset from the GAN framework to learn the optimal coordinated control policy for controlling the optimal reactive power set-points of PV inverters. The effectiveness of the proposed approach is analyzed on a modified IEEE 34-bus test case with added PV inverters. The results are compared and analyzed with a base case model with no VVC and VVC with a local droop control approach, genetic algorithm optimization, and a centralized soft actor–critic-based approach. Moreover, the performance of the proposed approach is compared with that of a multi-agent VVC framework without using the PV generation data and load information as the system state. The results illustrate that the proposed method with more state input improves the voltage profile and reduces the power loss of the network across various loading and PV generation scenarios.

14 SOLAR ENERGY↗

Linkage map construction using limited parental genotypic information

Abstract Genetic linkage maps based on single nucleotide polymorphisms (SNPs) represent an essential tool for a variety of genomic analyses. Today, next-generation sequencing (NGS) enables rapid genotyping of different mapping populations based on thousands of SNPs and the construction of highly saturated linkage maps. Nevertheless, missing data in the genotyping of the parental lines creates a bottleneck that determines the number of SNPs that can be used for the linkage map. As a proof of concept, a highly saturated genetic linkage map was constructed using the imputed genotypic data of a recombinant inbred line (RIL) population and the limited genotypic information of its parental lines. Two ABH genotype files were created from a pseudo-parental genotypic data set that includes all the SNPs present in the RIL population. In the first ABH file pseudo-parental 1 was considered parental A, while in the second pseudo-parental 1 was considered parental B. These two duplicate ABH genotype files were merged by chromosome and subjected to linkage map analysis. Since the ABH data were duplicated, two mirrored linkage groups were generated per chromosome. The correct linkage map was identified and selected based on the partial genotypic data of the parental lines. This strategy was effective for constructing a highly saturated linkage map of 33,421 SNPs based on the genotyping of 205 RILs and a limited number of 100 SNPs present in the parental lines. This strategy enables the use of all the NGS SNP data obtained from a low-coverage sequencing experiment in the mapping population.

59 BASIC BIOLOGICAL SCIENCES↗

Impact of duration and missing data on the long-term photovoltaic degradation rate estimation

Accurate quantification of photovoltaic (PV) system degradation rate (R D ) is essential for lifetime yield predictions. Although R D is a critical parameter, its estimation lacks a standardized methodology that can be applied on outdoor field data. The purpose of this paper is to investigate the impact of time period duration and missing data on R D by analyzing the performance of different techniques applied to synthetic PV system data at different linear R D patterns and known noise conditions. The analysis includes the application of different techniques to a 10-year synthetic dataset of a crystalline Silicon PV system, with emulated degradation levels and imputed missing data. Here, the analysis demonstrated that the accuracy of ordinary least squares (OLS), year-on-year (YOY), autoregressive integrated moving average (ARIMA) and robust principal component analysis (RPCA) techniques is affected by the evaluation duration with all techniques converging to lower R D deviations over the 10-year evaluation, apart from RPCA at high degradation levels. Moreover, the estimated R D is strongly affected by the amount of missing data. Filtering out the corrupted data yielded more accurate R D results for all techniques. It is proven that the application of a change-point detection stage is necessary and guidelines for accurate R D estimation are provided.

14 SOLAR ENERGY↗

Reliable machine prognostic health management in the presence of missing data

Prognostics and health management enables the prediction of future degradation and remaining useful life (RUL) for in-service systems based on historical and contemporary data, showing promise for many practical applications. One major challenge for prognostics is the common occurrence of missing values in time-series data, often caused by disruptions in sensor communication or hardware/software failures. Another major concern is that the sufficient prior knowledge of critical component degradation with a clear failure threshold is often not readily available in practice. These issues can significantly hinder the application of advanced signal and data analysis methods and consequently degrade the health management performance. In this article, we propose a novel data-driven framework that is capable of providing accurate and reliable predictions of degradation and RUL. In this approach, one-hot health state indicators are appended to the historical time series so that the model learns end-of-life automatically. A modified gate recurrent unit based variational autoencoder is employed in generative adversarial networks to model the temporal irregularity of the incomplete time series. Furthermore, experiments on multivariate time-series datasets collected from real-world aeroengines verify that significant performance improvement can be achieved using the proposed model for robust long-term prognostics.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Insights into Prismatic Loop Formation in Irradiated Fe–Cr Alloys from Hypothesis-Driven Active Learning and Causal Analysis

Neutron and electron irradiation experimental studies conducted on body-centered cubic Fe and Fe–Cr alloys have established two prismatic dislocation loop populations, which have Burgers vectors of either a/2$\langle$111$\rangle$ or a$\langle$100$\rangle$. Here, the loop formation depends on factors such as dose (D), dose rate (D rt ), temperature (T), chromium content (Cr%), and other alloying elements. Hence, it is important to understand how irradiation-induced dislocation loops evolve conditional upon the loop characteristics, such as loop density (DD), average loop size d̅, and irradiation parameters (D, D rt , T, and irradiation type), which is still an active area of research. To understand these complex structure–property relationships, machine learning (ML) is employed in a three-step approach. This includes imputing missing data with a k-nearest neighbor, generating functionalized features, and assessing feature importance with random forest classification and regression. Physics-based features are incorporated in a hypothesis-driven active learning scheme to overcome data unavailability challenges. Insights obtained from ML models (i) to categorize dislocation loop types, show the highest correlation with d̅; (ii) Log(DD), obtained through mathematical formulations involving D, Cr%, d̅, and T (e.g., Log(DD) ~ D + exp(-Cr%) + 1/d̅ and log(DD) ~ D + exp(-Cr%) + 1/T). Hypothesis-driven active learning is able to predict Log(DD) in which the experimental date is not known. Causal models verify cause–effect relationships for dislocation loop classification and irradiation factors in FeCr alloys.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Online PMU Missing Value Replacement Via Event-Participation Decomposition

We introduce a new method for online Phasor Measurement Unit (PMU) missing value replacement. Our approach allows us to decompose PMU event responses into a non-dynamic component (denoted the participation factor) that can be inferred directly from the past and a dynamic component that can be inferred directly from all other PMUs (denoted the event strength). When missing values occur, we can use these two components, which do not rely on the missing index, to estimate the correct value. The method is extremely fast and can easily be used for online applications. Furthermore, extensive testing on real power system event data reveals that our approach achieves state-of-the-art performance in terms of Mean Absolute Percent Errors (MAPEs) for PMU data dropped during event periods. Here, the method also yields an interpretable and simplified view of events for further analysis and applications. The method relies only on PMU data and does not take outside information such as network topology.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Satellite Embedding-Based Population Imputation for Areas with Missing Building Footprint Data: A Computer Vision-Based Approach

High-resolution population modeling is important for supporting effective decision-making across diverse sectors. LandScan Mosaic generates population estimates at the level of individual buildings and aggregates them to 3 arc-second grids, and this approach performs well in regions where building footprint data are comprehensive and reliable. However, large portions of the globe still suffer from incomplete, sparse, or entirely missing building stock datasets, creating a structural limitation for strictly building-based population models. To address this research gap, this study proposes a computer vision-based framework that employs Google Earth Engine satellite embeddings and UNet, which allows us to directly impute grid-level population estimates in building-data-deficient areas. Applied to Taiwan as a case study, the framework achieved strong predictive performance with R$^{2}$ of 0.89, RMSE of 18.70, and MAE of 8.41, outperforming traditional machine learning approaches. Notably, the proposed framework effectively addressed building false-positive errors inherent in Global Human Settlement Layer (GHSL) data, correctly identifying uninhabited areas that were erroneously classified as populated. The framework also offers significant advantages for global population mapping, particularly in terms of scalability and temporal consistency, thereby extending the coverage and accuracy of high-resolution population products in data-scarce regions worldwide. Urban planners, decision makers, and related stakeholders can obtain granular population distributions to support more accurate and targeted infrastructure investment, service delivery, resource allocation, and risk assessment decisions.

97 MATHEMATICS AND COMPUTING↗

Geostatistical interpolation of streambed hydrologic attributes with addition of left censored data and anisotropy

Spatial geostatistical interpolation of point measurements of streambed attributes in the hyporheic zone may be constrained by the streambed anisotropy, and data density and spatial distribution may significantly impact the results. Spatial clustering and low spatial data density can be caused by bedrock outcropping at the streambed limiting installation of in-stream piezometers. This study examines parameter error variability of the geostatistical interpolation using anisotropic interpolation methods and increasing the data density by adding left censored values (i.e., data below measurement limit) to locations where measurements were limited by exposed bedrock lining the streambed. The reduction in relative standard error of the interpolation was determined for the spatial distributions of streambed attributes including hydraulic conductivity, seepage flux, and mercury solute flux measured in two different years along a study reach in East Fork Poplar Creek, Tennessee, USA. Here, two methods to impute the left censored values were compared including the conventional half the detection limit substitution method, and the Stochastic Approximation of Expectation-Maximization (SAEM) algorithm, which both had comparable results. Imputing left censored data increased the data density to recommended ranges, reduced data clustering, increased the spatial dependence for some attributes, and reduced the standard error for each of the three attributes. For the reach considered herein, addition of the left censored values resulted in a larger error reduction than the consideration of anisotropy within the interpolation, which confirms the benefit of data addition to increase data density within data-limited river corridors.

58 GEOSCIENCES↗

Pharmacoepidemiology, Machine Learning and COVID-19: An intent-to-treat analysis of hydroxychloroquine, with or without azithromycin, and COVID-19 outcomes amongst hospitalized US Veterans

Hydroxychloroquine (HCQ) was proposed as an early therapy for coronavirus disease 2019 (COVID-19) after in vitro studies indicated possible benefit. Previous in vivo observational studies have presented conflicting results, though recent randomized clinical trials have reported no benefit from HCQ amongst hospitalized COVID-19 patients. In this work, we examined the effects of HCQ alone, and in combination with azithromycin, in a hospitalized COVID-19 positive, United States (US) Veteran population using a propensity score adjusted survival analysis with imputation of missing data. From March 1, 2020 through April 30, 2020, 64,055 US Veterans were tested for COVID-19 based on Veteran Affairs Healthcare Administration electronic health record data. Of the 7,193 positive cases, 2,809 were hospitalized, and 657 individuals were prescribed HCQ within the first 48-hours of hospitalization for the treatment of COVID-19. There was no apparent benefit associated with HCQ receipt, alone or in combination with azithromycin, and an increased risk of intubation when used in combination with azithromycin [Hazard Ratio (95% Confidence Interval): 1.55 (1.07, 2.24)]. In conclusion, we assessed the effectiveness of HCQ with or without azithromycin in treating patients hospitalized with COVID-19 using a national sample of the US Veteran population. Using rigorous study design and analytic methods to reduce confounding and bias, we found no evidence of a survival benefit from the administration of HCQ.

60 APPLIED LIFE SCIENCES↗

PVplr-stGNN 0.1.10

PV Performance Loss Rate Estimation using Spatio-temporal Graph Neural Networks PVplr-stGNN is a Python 3 package developed by the SDLE Research Center at Case Western Reserve University in Cleveland OH. This repository contains the full source PVplr-stGNN package. The package contains the PV-stGAE for missingness data detection and imputation and PV-DynGNN for PLR estimation.

Fan, Yangxin [Case Western Reserve Univ., Clevelan↗

Brownian bridge-based speed imputation technique for truck energy consumption and emissions estimation

The available truck Global Positioning System (GPS) data, typically collected with large time gaps, rely on imputation techniques to obtain second-by-second data that are required in models for estimating truck energy consumption and emissions. However, existing speed imputation techniques either require a large amount of high-resolution data for model training or rely on special movement assumptions. Here, to fill the gap and effectively apply the low-resolution truck GPS datasets, this paper proposes a simple imputation technique that adopts the Brownian bridge structure to impute missing speed data. The proposed technique introduces a feasible imputation region and a combined drift into the imputation procedure to capture vehicle acceleration constraint, travel distance constraint, and speed volatility. The calibrated model is applied to a set of low-resolution truck GPS data. The results demonstrate the robustness of the proposed technique in enhancing estimation accuracy when using low-resolution GPS data to estimate fuel consumption and emissions.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

FAIRification, Quality Assessment, and Missingness Pattern Discovery for Spatiotemporal Photovoltaic Data

The growth of the photovoltaic market has pushed the demand for power forecasting and performance evaluation for a huge population of PV power plants. Many of these power plants have spatiotemporal coherence that can be utilized for improving model accuracy. We have demonstrated in this paper the FAIRification of spatiotemporal PV time series data. Through the creation of a solar power plant ontology, we propose standards for the naming and structure of metadata used to describe the data from these power plants. Using the structure from this ontology, we have developed both R and Python packages for the automation of the FAIRification process. Going further, we have also developed an R package that automates the analysis of the quality of a data set through the designation of letter grades. To solve the issue of data missingness, we propose the use of St-GNN autoencoders to detect and impute missing values from a data set by utilizing data from power plants nearby.

14 SOLAR ENERGY↗

FAIRification, Quality Assessment, and Missingness Pattern Discovery for Spatiotemporal Photovoltaic Data

The ongoing growth of the photovoltaic market has pushed the demand for power forecasting and performance evaluation for a huge population of PV power plants. Through access to a large number of time series data sets from different power plants, we have found common issues that impede the modeling process. Namely, the time series data are hard to transfer between groups due to differences in variable nomenclature, and the quality of the data sets can vary. We address the issue of variable nomenclature by FAIRifying spatiotemporal PV time series data. Through the creation of a solar power plant ontology, we propose standards for the naming and structure of metadata used to describe the data from these power plants. Using the structure from this ontology, we have developed both R and Python packages for the automation of the FAIRification process. We have also developed an R package that automates the analysis of the quality of a data set through the designation of letter grades. With access to large time series data sets across many power plants, we can utilize spatiotemporal coherence between the sites in order to improve the quality of our data. To solve the issue of data missingness, we propose the use of Spatiotemporal-GNN autoencoders to detect and impute missing values from a data set by utilizing data from power plants nearby.

14 SOLAR ENERGY↗