Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “imputation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Long-term missing value imputation for time series data using deep neural networks

We present an approach that uses a deep learning model, in particular, a MultiLayer Perceptron, for estimating the missing values of a variable in multivariate time series data. We focus on filling a long continuous gap (e.g., multiple months of missing daily observations) rather than on individual randomly missing observations. Our proposed gap filling algorithm uses an automated method for determining the optimal MLP model architecture, thus allowing for optimal prediction performance for the given time series. We tested our approach by filling gaps of various lengths (three months to three years) in three environmental datasets with different time series characteristics, namely daily groundwater levels, daily soil moisture, and hourly Net Ecosystem Exchange. We compared the accuracy of the gap-filled values obtained with our approach to the widely used R-based time series gap filling methods ImputeTS and mtsdi. The results indicate that using an MLP for filling a large gap leads to better results, especially when the data behave nonlinearly. Thus, our approach enables the use of datasets that have a large gap in one variable, which is common in many long-term environmental monitoring observations.

97 MATHEMATICS AND COMPUTING↗

Leveraging large language models to address data scarcity in machine learning for graphene synthesis

Machine learning in experimental materials science faces significant challenges due to the scarcity of data, which are costly and time-consuming to generate, particularly when relying on in-house experiments. Literature data mining offers a potential solution but introduces issues like mixed data quality, inconsistent formats, and non-uniform reporting of synthesis parameters, resulting in partially missing and heterogeneous features across the dataset. Here, we propose data imputation and feature engineering methods that employ pre-trained large language models (LLMs) to enhance machine learning performance on scarce, heterogeneous datasets, demonstrated on graphene CVD synthesis data and the ML-HydPARK hydrogen storage dataset. GPT models perform data imputation via tailored prompting and semantic normalization of inconsistently reported features through embeddings, for example, to harmonize the complex nomenclature of CVD substrates. Beyond yielding more diverse and richer feature representations than traditional methods such as K-nearest neighbors (KNN) and Multivariate Imputation by Chained Equations (MICE), LLM-based data imputation is evaluated against dataset characteristics and prompting strategies. We vary the level of autonomy granted to the LLM, from generic prompting that leverages pre-trained knowledge for autonomous data generation to data-informed prompting that constrains outputs using target-specific information, and demonstrate which level of autonomy yields superior imputation performance across datasets and feature types. The proposed data engineering methods markedly improve downstream performance; for example, in graphene layer number classification using a support vector machine (SVM), binary accuracy increases from 39% to 65% and ternary accuracy from 52% to 72%. Fine-tuning experiments on both datasets show that combining our proposed LLM-based data imputation and feature encoding methods with numerical machine learning predictors outperforms standalone fine-tuned LLM predictors in data-scarce settings. The proposed strategies emphasize data enhancement techniques rather than refining learning architectures or regularizing loss functions, offering a broadly applicable framework for improving machine learning performance on scarce, inhomogeneous datasets.

Chemical vapor deposition↗

MultiPEM Toolbox: User Manual [Rev. 2]

This document explains use of the Multi-Phenomenology Explosion Monitoring (Multi PEM) Toolbox, a collection of R scripts for estimating the unknown device parameters of a new event with uncertainty quantification. The methodology and application used for illustration in this user manual are fully documented in a Los Alamos National Laboratory technical report hereafter designated “WPA” for reference. Additional details on the application are found in a recent journal article. Two assessment types are available: rapid and complete. Rapid assessments are conducted in two stages, as described in Section 2. In the first stage, calibration data are used to estimate forward and error model parameters (WPA, §5.1) and (if relevant) errors-in-variables yield values for calibration sources (WPA, §3, Equation (3)). In the second stage, new event data are used to estimate the unknown new event device parameters (WPA, §5.2) with uncertainty quantification. Two options for treating the inferred first stage parameters in second stage Bayesian analysis are available: fixing them at their maximum likelihood estimate (default), or multiple imputation. Multiple imputation involves utilizing several posterior samples (imputations) of the first stage parameters as fixed values in the second stage posterior sampling of the new event device parameters. Second stage sampling is conducted across imputations in parallel to improve computational efficiency. This method produces improved uncertainty quantification of the new event device parameters compared with the default treatment of the first stage parameters, at the expense of additional computation. Complete assessments are conducted in a single stage, as described in Section 3. Calibration and (if relevant) new event data are used simultaneously to estimate all forward model, error model, and (if relevant) new event device parameters with uncertainty quantification on the latter. As the name suggests, rapid assessments generally run substantially faster than complete assessments (even with multiple imputation), because the results of first stage analysis can be stored and incorporated into estimating a relatively low-dimensional space of new event device parameters whenever relevant new event data becomes available. On the other hand, complete assessments must be run on the full set of model and device parameters with calibration and new event data every time the latter becomes available.

97 MATHEMATICS AND COMPUTING↗

Evaluation of Time-Series Gap-Filling Methods for Solar Irradiance Applications

A complete solar resource data set is essential for any stage of a solar energy project - from feasibility studies to daily operations. But measured or modeled solar resource data are prone to data gaps and data quality issues. To mitigate these issues, a data imputation process should be implemented to obtain a complete and reliable temporal and spatial data series. This study focused on imputing temporal scales by applying random and artificial data gaps and then implementing eight imputation methods, including the Kalman filtering and smoothing and stine interpolations. These methods were implemented on 1-minute to half hourly irradiance data for 1 year using a few locations from the National Solar Radiation Database (NSRDB) and ground measurement data set. The results demonstrated that some of the simpler methods, such as the stine and linear interpolation methods, were the relatively best models based on the statistical metrics for imputing NSRDB and ground measurement data, respectively.

14 SOLAR ENERGY↗

Multi-agent voltage control in distribution systems using GAN-DRL-based approach

Active distribution grids can experience voltage fluctuations and violations due to the high penetration of variable distributed energy resources (DERs). These problems might occur because of the uncertain and variable generation natures of these resources, especially solar photovoltaic resources, during panel shadowing scenarios. Volt-VAR control (VVC) is an efficient method that controls the reactive power set-points of the inverters to regulate the voltage of distribution grids. Although several VVC approaches have been proposed recently, the performance of these approaches degrades significantly if behind-the-meter solar generation data are unobservable/missing. Therefore, it is necessary to impute missing/unobservable PV data accurately to be utilized in VVC approaches. Further, this paper proposes a model-free, data-driven, centrally trained, and decentrally executed multi-agent deep reinforcement learning-based VVC architecture to regulate the voltage of distribution networks. A generative adversarial network (GAN) is incorporated to impute the unobservable PV data accurately, which improves the performance of the proposed control architecture. The proposed multi-agent-soft-actor–critic algorithm (MASAC)-based VVC technique utilizes the actual PV dataset as well as the imputed dataset from the GAN framework to learn the optimal coordinated control policy for controlling the optimal reactive power set-points of PV inverters. The effectiveness of the proposed approach is analyzed on a modified IEEE 34-bus test case with added PV inverters. The results are compared and analyzed with a base case model with no VVC and VVC with a local droop control approach, genetic algorithm optimization, and a centralized soft actor–critic-based approach. Moreover, the performance of the proposed approach is compared with that of a multi-agent VVC framework without using the PV generation data and load information as the system state. The results illustrate that the proposed method with more state input improves the voltage profile and reduces the power loss of the network across various loading and PV generation scenarios.

14 SOLAR ENERGY↗

Revised monthly energy generation estimates for 1,500 hydroelectric power plants in the United States

Abstract The U.S. Energy Information Administration (EIA) conducts a regular survey (form EIA-923) to collect annual and monthly net generation for more than ten thousand U.S. power plants. Approximately 90% of the ~1,500 hydroelectric plants included in this data release are surveyed at annual resolution only and thus lack actual observations of monthly generation. For each of these plants, EIA imputes monthly generation values using the combined monthly generating pattern of other hydropower plants within the corresponding census division. The imputation method neglects local hydrology and reservoir operations, rendering the monthly data unsuitable for various research applications. Here we present an alternative approach to disaggregate each unobserved plant’s reported annual generation using proxies of monthly generation—namely historical monthly reservoir releases and average river discharge rates recorded downstream of each dam. Evaluation of the new dataset demonstrates substantial and robust improvement over the current imputation method, particularly if reservoir release data are available. The new dataset—named RectifHyd—provides an alternative to EIA-923 for U.S. scale, plant-level, monthly hydropower net generation (2001–2020). RectifHyd may be used to support power system studies or analyze within-year hydropower generation behavior at various spatial scales.

13 HYDRO ENERGY↗

Soybean ( Glycine max ) Haplotype Map (GmHapMap): a universal resource for soybean translational and functional genomics

Here, we describe a worldwide haplotype map for soybean (GmHapMap) constructed using whole-genome sequence data for 1007 Glycine max accessions and yielding 14.9 million variants as well as 4.3 M tag single-nucleotide polymorphisms (SNPs). When sampling random subsets of these accessions, the number of variants and tag SNPs plateaued beyond approximately 800 and 600 accessions, respectively. This suggests extensive coverage of diversity within the cultivated soybean. GmHapMap variants were imputed onto 21 618 previously genotyped accessions with up to 96% success for common alleles. A local association analysis was performed with the imputed data using markers located in a 1-Mb region known to contribute to seed oil content and enabled us to identify a candidate causal SNP residing in the NPC1 gene. We determined gene-centric haplotypes (407 867 GCHs) for the 55 589 genes and showed that such haplotypes can help to identify alleles that differ in the resulting phenotype. Finally, we predicted 18 031 putative loss-of-function (LOF) mutations in 10 662 genes and illustrated how such a resource can be used to explore gene function. The GmHapMap provides a unique worldwide resource for applied soybean genomics and breeding.

54 ENVIRONMENTAL SCIENCES↗

Scripts and raw/processed precipitation NRCS network data, Upper Colorado, 2008-2017

This data package consists of scripts and data that were used to generate all the results and figures for Mital et al., 2020 (https://doi.org/10.3389/frwa.2020.00020). The purpose of the study was to develop a new algorithm to impute (or gap-fill) missing daily precipitation data. Study area was the Upper Colorado Water Resources Region (UCWRR), over a time period of 2008-2017. In terms of data, the package consists of raw precipitation measurements from Natural Resources Conservation Service (NRCS) network. These raw measurements are compiled into a single csv file (NRCS_dates_2008_2017.csv). The corresponding metadata are compiled in Metadata_processed.csv. The package also consists of outputs generated during Baseline and Sequential imputation runs as documented in Mital et al., 2020. Scripts used to generate all the results and figures for Mital et al., 2020 are also uploaded. A readme file documenting the layout of the data archive is also uploaded, and a KML file shows the geographic range for the data.

54 ENVIRONMENTAL SCIENCES↗

Enhancing approximate modular Bayesian inference by emulating the conditional posterior

In modular Bayesian analyses, complex models are composed of distinct modules, each representing different aspects of the data or prior information. In this context, fully Bayesian approaches can sometimes lead to undesirable feedback between modules, compromising the integrity of the inference. The “cut-distribution” prevents unwanted influence between modules by “cutting” feedback. The direct sampling (DS) algorithm is standard practice for approximating the cut-distribution, but it can be computationally intensive, especially when the number of imputations required is large. An enhanced method is proposed, the Emulating the Conditional Posterior (ECP) algorithm, which leverages emulation to increase the number of imputations. Through numerical experiment it is demonstrated that the ECP algorithm outperforms the traditional DS approach in terms of accuracy and computational efficiency, particularly when resources are constrained. Here, it is also shown how the DS algorithm can be improved using ideas from design of experiments. Some practical recommendations are given for algorithm choice in modular Bayesian analyses.

97 MATHEMATICS AND COMPUTING↗

Geostatistical interpolation of streambed hydrologic attributes with addition of left censored data and anisotropy

Spatial geostatistical interpolation of point measurements of streambed attributes in the hyporheic zone may be constrained by the streambed anisotropy, and data density and spatial distribution may significantly impact the results. Spatial clustering and low spatial data density can be caused by bedrock outcropping at the streambed limiting installation of in-stream piezometers. This study examines parameter error variability of the geostatistical interpolation using anisotropic interpolation methods and increasing the data density by adding left censored values (i.e., data below measurement limit) to locations where measurements were limited by exposed bedrock lining the streambed. The reduction in relative standard error of the interpolation was determined for the spatial distributions of streambed attributes including hydraulic conductivity, seepage flux, and mercury solute flux measured in two different years along a study reach in East Fork Poplar Creek, Tennessee, USA. Here, two methods to impute the left censored values were compared including the conventional half the detection limit substitution method, and the Stochastic Approximation of Expectation-Maximization (SAEM) algorithm, which both had comparable results. Imputing left censored data increased the data density to recommended ranges, reduced data clustering, increased the spatial dependence for some attributes, and reduced the standard error for each of the three attributes. For the reach considered herein, addition of the left censored values resulted in a larger error reduction than the consideration of anisotropy within the interpolation, which confirms the benefit of data addition to increase data density within data-limited river corridors.

58 GEOSCIENCES↗

Low-rank Tensor Completion for PMU Data Recovery

This paper proposes a tensor completion method for the recovery of missing phasor measurement unit (PMU) measurements. Tensor completion as the general case of matrix completion has attracted increasing attention in recent years. The imputation accuracy for the existing matrix completion methods may be significantly reduced when there are consecutive data losses across multiple data channels. To tackle this issue, we explore the multi-way characteristics of PMU measurements by using a tensor model. We leverage the low-rank property of the PMU measurements and formulate the missing PMU data recovery problem as a low-rank tensor completion problem. An efficient algorithm based on alternating direction method of multipliers (ADMM) is developed to solve the tensor completion problem. The experiments using the real PMU dataset show that the proposed method exhibits better imputation accuracy compared with the conventional data recovery methods.

Ghasemkhani, Amir↗

Latent Neural ODE for Integrating Multi-Timescale Measurements in Smart Distribution Grids

Under a smart grid paradigm, there has been an increase in sensor installations to enhance situational awareness. The measurements from these sensors can be leveraged for real-time monitoring, control, and protection. However, these measurements are typically irregularly sampled. These measure-ments may also be intermittent due to communication bandwidth limitations. To tackle this problem, this paper proposes a novel latent neural ordinary differential equations (LODE) approach to aggregate the unevenly sampled multivariate time-series measurements. The proposed approach is flexible in performing both imputations and predictions while being computationally efficient. Simulation results on IEEE 37 bus test systems illustrate the efficiency of the proposed approach.

multi time-scale measurements↗

On the Impact of Income, Age, and Travel Distance on the Value of Time

The value of time (VOT) is a fundamental component used in transportation modeling, policy analysis, and economic appraisal. Decades of research and practice have empirically estimated the VOT across many factors (e.g., mode, purpose, time, comfort, etc.), yet little is known about its underlying form. Although it is well established that VOT can vary, it is still unclear whether patterns exist in this variation. The objective of this paper is not to merely estimate the VOT, but to model the VOT across multiple continuous and interacting variables. The purpose is to reveal its functional form with respect to mode, age, gender, purpose, income, and time of day to provide a generalizable understanding for future research and practice. Such an understanding can help develop simpler models and reduce the need for bespoke estimations for every conceivable variable perturbation. This research utilized a household travel survey containing 14,159 reported trips with imputed travel time and costs for the alternative mode choices. The average overall estimated VOT is 40.32 $/h, with results showing VOT varying log-linearly with income and trip distance, but following a Gaussian function (normal curve) with age. Overall, the results show that travel distance dominates VOT variation, which increases exponentially at a rate that is 3.61 times higher per mile of distance than per $10,000 of income, and that VOT by age peaks at age 54. This basic understanding of how the VOT varies sets the foundation for answering the subsequent question for why it might vary.

Engineering↗

LandScan Mosaic

The LandScan program at Oak Ridge National Laboratory (ORNL), in collaboration with the National Geospatial-Intelligence Agency (NGA), continues to deliver the most accurate and up to date global, high resolution gridded population data. Additionally, the latest advancements in the LandScan HD methodology led to reduced latency in development of rapid updates for geopolitical events. With momentum towards reporting more up to date population estimates, feedback from the user community expressed interest in reporting population estimates in ranges - whether to express a level of uncertainty or confirm to leadership and stakeholders the modeled data are estimates. Building upon the need to understand uncertainty or confidence in the modeled data and report ranges at the global scale, LandScan Mosaic was developed. LandScan Mosaic represents the next generation of high-resolution population modeling, building upon the established success of previous LandScan HD iterations. While LandScan HD employed a deterministic big data fusion approach, LandScan Mosaic enhances this methodology by integrating advanced machine learning techniques to impute missing, yet crucial, population model parameters. This advancement allows for probabilistic modeling of building occupancy and population distribution, incorporating uncertainty quantification through Monte Carlo sampling methods. By combining big data fusion with machine learning-driven imputation and stochastic modeling, LandScan Mosaic provides a more comprehensive and robust representation of population dynamics. LandScan Mosaic will be following the in the footsteps of its longstanding counterpart LandScan Global and releasing a global gridded population raster, at the 3-arcsecond resolution. This technical report documents the current stage of development of LandScan Mosaic, detailing the methodologies and data sources behind the modeling. Stakeholders are encouraged to use this document as an authoritative reference for insight into Mosaic’s data development processes. However, readers should note that LandScan Mosaic remains in a late-stage research and development phase, and methodologies and data presented here are subject to refinements ahead of the anticipated global release in Summer 2025. Feedback and inquiries from users and stakeholders are welcomed as we continue to refine and enhance this important population resource.

97 MATHEMATICS AND COMPUTING↗

Cleaning Images with Gaussian Process Regression

Many approaches to astronomical data reduction and analysis cannot tolerate missing data: corrupted pixels must first have their values imputed. This paper presents astrofix, a robust and flexible image imputation algorithm based on Gaussian process regression. Through an optimization process, astrofix chooses and applies a different interpolation kernel to each image, using a training set extracted automatically from that image. It naturally handles clusters of bad pixels and image edges and adapts to various instruments and image types. For bright pixels, the mean absolute error of astrofix is several times smaller than that of median replacement and interpolation by a Gaussian kernel. We demonstrate good performance on both imaging and spectroscopic data, including the SBIG 6303 0.4 m telescope and the FLOYDS spectrograph of Las Cumbres Observatory and the CHARIS integral-field spectrograph on the Subaru Telescope.

42 ENGINEERING↗

A consistent dataset for the net income distribution for 190 countries and aggregated to 32 geographical regions from 1958 to 2015

Abstract. Data on income distributions within and across countries are becoming increasingly important for informing analysis of income inequality and understanding the distributional consequences of climate change. While datasets on income distribution collected from household surveys are available for multiple countries, these datasets often do not represent the same concept of inequality (or income concept) and therefore make comparisons across countries, over time and across datasets difficult. Here, we present a consistent dataset of income distributions across 190 countries from 1958 to 2015 measured in terms of net income. We complement the observed values in this dataset with values imputed from a summary measure of the income distribution, specifically the Gini coefficient. For the imputation, we use a recently developed nonparametric principal-component-based approach that shows an excellent fit to data on income distributions compared to other approaches. We also present another version of this dataset aggregated from the country level to 32 geographical regions. Our dataset is developed for the purpose of calibrating models such as integrated human–Earth system models with detailed data on income distributions. This dataset will enable more robust analysis of income distribution at multiple scales. The latest version of our data are available on Zenodo: https://doi.org/10.5281/zenodo.7093997 (Narayan et al., 2022b).

97 MATHEMATICS AND COMPUTING↗