Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “imputation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Imputing historical statistics, soils information, and other land-use data to crop area

In foreign crop condition monitoring, satellite acquired imagery is routinely used. To facilitate interpretation of this imagery, it is advantageous to have estimates of the crop types and their extent for small area units, i.e., grid cells on a map represent, at 60 deg latitude, an area nominally 25 by 25 nautical miles in size. The feasibility of imputing historical crop statistics, soils information, and other ancillary data to crop area for a province in Argentina is studied.

Perry, C. R., Jr.↗

A Comparison of Time-Series Gap-Filling Methods to Impute Solar Radiation Data: Preprint

Complete solar resource data sets play a critical role at every stage of solar energy projects; however, measured or modeled solar resource data come with significant uncertainties and usually suffer from several issues, including, but not limited to, data gaps and data quality issues. To mitigate these issues, an appropriate data imputation method should be implemented to build a complete and reliable temporal (and spatial) database. Motivated by this, in this study, we extensively compare the performance of eight different gap-filling methods by creating random and artificial data gaps in (i) hourly irradiance data for 1 year using a few locations of the National Solar Radiation Database (NSRDB) and (ii) 1-minute ground measurement data sets from the Surface Radiation Budget Network (SURFRAD) and the National Renewable Energy Laboratory (NREL) stations.

clearness index↗

Long-term missing value imputation for time series data using deep neural networks

We present an approach that uses a deep learning model, in particular, a MultiLayer Perceptron, for estimating the missing values of a variable in multivariate time series data. We focus on filling a long continuous gap (e.g., multiple months of missing daily observations) rather than on individual randomly missing observations. Our proposed gap filling algorithm uses an automated method for determining the optimal MLP model architecture, thus allowing for optimal prediction performance for the given time series. We tested our approach by filling gaps of various lengths (three months to three years) in three environmental datasets with different time series characteristics, namely daily groundwater levels, daily soil moisture, and hourly Net Ecosystem Exchange. We compared the accuracy of the gap-filled values obtained with our approach to the widely used R-based time series gap filling methods ImputeTS and mtsdi. The results indicate that using an MLP for filling a large gap leads to better results, especially when the data behave nonlinearly. Thus, our approach enables the use of datasets that have a large gap in one variable, which is common in many long-term environmental monitoring observations.

97 MATHEMATICS AND COMPUTING↗

Leveraging large language models to address data scarcity in machine learning for graphene synthesis

Machine learning in experimental materials science faces significant challenges due to the scarcity of data, which are costly and time-consuming to generate, particularly when relying on in-house experiments. Literature data mining offers a potential solution but introduces issues like mixed data quality, inconsistent formats, and non-uniform reporting of synthesis parameters, resulting in partially missing and heterogeneous features across the dataset. Here, we propose data imputation and feature engineering methods that employ pre-trained large language models (LLMs) to enhance machine learning performance on scarce, heterogeneous datasets, demonstrated on graphene CVD synthesis data and the ML-HydPARK hydrogen storage dataset. GPT models perform data imputation via tailored prompting and semantic normalization of inconsistently reported features through embeddings, for example, to harmonize the complex nomenclature of CVD substrates. Beyond yielding more diverse and richer feature representations than traditional methods such as K-nearest neighbors (KNN) and Multivariate Imputation by Chained Equations (MICE), LLM-based data imputation is evaluated against dataset characteristics and prompting strategies. We vary the level of autonomy granted to the LLM, from generic prompting that leverages pre-trained knowledge for autonomous data generation to data-informed prompting that constrains outputs using target-specific information, and demonstrate which level of autonomy yields superior imputation performance across datasets and feature types. The proposed data engineering methods markedly improve downstream performance; for example, in graphene layer number classification using a support vector machine (SVM), binary accuracy increases from 39% to 65% and ternary accuracy from 52% to 72%. Fine-tuning experiments on both datasets show that combining our proposed LLM-based data imputation and feature encoding methods with numerical machine learning predictors outperforms standalone fine-tuned LLM predictors in data-scarce settings. The proposed strategies emphasize data enhancement techniques rather than refining learning architectures or regularizing loss functions, offering a broadly applicable framework for improving machine learning performance on scarce, inhomogeneous datasets.

Chemical vapor deposition↗

MultiPEM Toolbox: User Manual [Rev. 2]

This document explains use of the Multi-Phenomenology Explosion Monitoring (Multi PEM) Toolbox, a collection of R scripts for estimating the unknown device parameters of a new event with uncertainty quantification. The methodology and application used for illustration in this user manual are fully documented in a Los Alamos National Laboratory technical report hereafter designated “WPA” for reference. Additional details on the application are found in a recent journal article. Two assessment types are available: rapid and complete. Rapid assessments are conducted in two stages, as described in Section 2. In the first stage, calibration data are used to estimate forward and error model parameters (WPA, §5.1) and (if relevant) errors-in-variables yield values for calibration sources (WPA, §3, Equation (3)). In the second stage, new event data are used to estimate the unknown new event device parameters (WPA, §5.2) with uncertainty quantification. Two options for treating the inferred first stage parameters in second stage Bayesian analysis are available: fixing them at their maximum likelihood estimate (default), or multiple imputation. Multiple imputation involves utilizing several posterior samples (imputations) of the first stage parameters as fixed values in the second stage posterior sampling of the new event device parameters. Second stage sampling is conducted across imputations in parallel to improve computational efficiency. This method produces improved uncertainty quantification of the new event device parameters compared with the default treatment of the first stage parameters, at the expense of additional computation. Complete assessments are conducted in a single stage, as described in Section 3. Calibration and (if relevant) new event data are used simultaneously to estimate all forward model, error model, and (if relevant) new event device parameters with uncertainty quantification on the latter. As the name suggests, rapid assessments generally run substantially faster than complete assessments (even with multiple imputation), because the results of first stage analysis can be stored and incorporated into estimating a relatively low-dimensional space of new event device parameters whenever relevant new event data becomes available. On the other hand, complete assessments must be run on the full set of model and device parameters with calibration and new event data every time the latter becomes available.

97 MATHEMATICS AND COMPUTING↗

Evaluation of Time-Series Gap-Filling Methods for Solar Irradiance Applications

A complete solar resource data set is essential for any stage of a solar energy project - from feasibility studies to daily operations. But measured or modeled solar resource data are prone to data gaps and data quality issues. To mitigate these issues, a data imputation process should be implemented to obtain a complete and reliable temporal and spatial data series. This study focused on imputing temporal scales by applying random and artificial data gaps and then implementing eight imputation methods, including the Kalman filtering and smoothing and stine interpolations. These methods were implemented on 1-minute to half hourly irradiance data for 1 year using a few locations from the National Solar Radiation Database (NSRDB) and ground measurement data set. The results demonstrated that some of the simpler methods, such as the stine and linear interpolation methods, were the relatively best models based on the statistical metrics for imputing NSRDB and ground measurement data, respectively.

14 SOLAR ENERGY↗

Astronaut Preflight Cardiovascular Variables Associated with Vascular Compliance are Highly Correlated with Post-Flight Eye Outcome Measures in the Visual Impairment Intracranial Pressure (VIIP) Syndrome Following Long Duration Spaceflight

The detection of the first VIIP case occurred in 2005, and adequate eye outcome measures were available for 31 (67.4%) of the 46 long duration US crewmembers who had flown on the ISS since its first crewed mission in 2000. Therefore, this analysis is limited to a subgroup (22 males and 9 females). A "cardiovascular profile" for each astronaut was compiled by examining twelve individual parameters; eleven of these were preflight variables: systolic blood pressure, pulse pressure, body mass index, percentage body fat, LDL, HDL, triglycerides, use of anti‐lipid medication, fasting serum glucose, and maximal oxygen uptake in ml/kg. Each of these variables was averaged across three preflight annual physical exams. Astronaut age prior to the long duration mission, and inflight salt intake was also included in the analysis. The group of cardiovascular variables for each crew member was compared with seven VIIP eye outcome variables collected during the immediate post‐flight period: anterior-posterior axial length of the globe measured by ultrasound and optical biometry; optic nerve sheath diameter, optic nerve diameter, and optic nerve to sheath ratio‐ each measured by ultrasound and magnetic resonance imaging (MRI), intraocular pressure (IOP), change in manifest refraction, mean retinal nerve fiber layer (RNFL) on optical coherence tomography (OCT), and RNFL of the inferior and superior retinal quadrants. Since most of the VIIP eye outcome measures were added sequentially beginning in 2005, as knowledge of the syndrome improved, data were unavailable for 22.0% of the outcome measurements. To address the missing data, we employed multivariate multiple imputation techniques with predictive mean matching methods to accumulate 200 separate imputed datasets for analysis. We were able to impute data for the 22.0% of missing VIIP eye outcomes. We then applied Rubin's rules for collapsing the statistical results across our 200 multiply imputed data sets to assess the canonical correlation between the eye outcomes and the twelve astronaut cardiovascular variables available for all 31 subjects. Results: A highly significant canonical correlation was observed among the canonical solutions (p<.00001), with an average best canonical correlation of.97. The results suggest a strong association between astronauts' measures of cardiovascular health and the seven eye outcomes of the VIIP syndrome used in this analysis. Furthermore, the "joint test" revealed a significant difference in cardiovascular profile between male and female astronauts (Prob > F = 0.00001). Overall, female astronauts demonstrated a significantly healthier cardiovascular status. Individually, the female astronauts had significantly healthier profiles on seven of twelve cardiovascular variables than the men (p values ranging from <0.0001 to <0.05). Male astronauts did not demonstrate significantly healthier values on any of the twelve cardiovascular variables measured

Otto, Christian↗

Multi-agent voltage control in distribution systems using GAN-DRL-based approach

Active distribution grids can experience voltage fluctuations and violations due to the high penetration of variable distributed energy resources (DERs). These problems might occur because of the uncertain and variable generation natures of these resources, especially solar photovoltaic resources, during panel shadowing scenarios. Volt-VAR control (VVC) is an efficient method that controls the reactive power set-points of the inverters to regulate the voltage of distribution grids. Although several VVC approaches have been proposed recently, the performance of these approaches degrades significantly if behind-the-meter solar generation data are unobservable/missing. Therefore, it is necessary to impute missing/unobservable PV data accurately to be utilized in VVC approaches. Further, this paper proposes a model-free, data-driven, centrally trained, and decentrally executed multi-agent deep reinforcement learning-based VVC architecture to regulate the voltage of distribution networks. A generative adversarial network (GAN) is incorporated to impute the unobservable PV data accurately, which improves the performance of the proposed control architecture. The proposed multi-agent-soft-actor–critic algorithm (MASAC)-based VVC technique utilizes the actual PV dataset as well as the imputed dataset from the GAN framework to learn the optimal coordinated control policy for controlling the optimal reactive power set-points of PV inverters. The effectiveness of the proposed approach is analyzed on a modified IEEE 34-bus test case with added PV inverters. The results are compared and analyzed with a base case model with no VVC and VVC with a local droop control approach, genetic algorithm optimization, and a centralized soft actor–critic-based approach. Moreover, the performance of the proposed approach is compared with that of a multi-agent VVC framework without using the PV generation data and load information as the system state. The results illustrate that the proposed method with more state input improves the voltage profile and reduces the power loss of the network across various loading and PV generation scenarios.

14 SOLAR ENERGY↗

Revised monthly energy generation estimates for 1,500 hydroelectric power plants in the United States

Abstract The U.S. Energy Information Administration (EIA) conducts a regular survey (form EIA-923) to collect annual and monthly net generation for more than ten thousand U.S. power plants. Approximately 90% of the ~1,500 hydroelectric plants included in this data release are surveyed at annual resolution only and thus lack actual observations of monthly generation. For each of these plants, EIA imputes monthly generation values using the combined monthly generating pattern of other hydropower plants within the corresponding census division. The imputation method neglects local hydrology and reservoir operations, rendering the monthly data unsuitable for various research applications. Here we present an alternative approach to disaggregate each unobserved plant’s reported annual generation using proxies of monthly generation—namely historical monthly reservoir releases and average river discharge rates recorded downstream of each dam. Evaluation of the new dataset demonstrates substantial and robust improvement over the current imputation method, particularly if reservoir release data are available. The new dataset—named RectifHyd—provides an alternative to EIA-923 for U.S. scale, plant-level, monthly hydropower net generation (2001–2020). RectifHyd may be used to support power system studies or analyze within-year hydropower generation behavior at various spatial scales.

13 HYDRO ENERGY↗

Soybean ( Glycine max ) Haplotype Map (GmHapMap): a universal resource for soybean translational and functional genomics

Here, we describe a worldwide haplotype map for soybean (GmHapMap) constructed using whole-genome sequence data for 1007 Glycine max accessions and yielding 14.9 million variants as well as 4.3 M tag single-nucleotide polymorphisms (SNPs). When sampling random subsets of these accessions, the number of variants and tag SNPs plateaued beyond approximately 800 and 600 accessions, respectively. This suggests extensive coverage of diversity within the cultivated soybean. GmHapMap variants were imputed onto 21 618 previously genotyped accessions with up to 96% success for common alleles. A local association analysis was performed with the imputed data using markers located in a 1-Mb region known to contribute to seed oil content and enabled us to identify a candidate causal SNP residing in the NPC1 gene. We determined gene-centric haplotypes (407 867 GCHs) for the 55 589 genes and showed that such haplotypes can help to identify alleles that differ in the resulting phenotype. Finally, we predicted 18 031 putative loss-of-function (LOF) mutations in 10 662 genes and illustrated how such a resource can be used to explore gene function. The GmHapMap provides a unique worldwide resource for applied soybean genomics and breeding.

54 ENVIRONMENTAL SCIENCES↗

Scripts and raw/processed precipitation NRCS network data, Upper Colorado, 2008-2017

This data package consists of scripts and data that were used to generate all the results and figures for Mital et al., 2020 (https://doi.org/10.3389/frwa.2020.00020). The purpose of the study was to develop a new algorithm to impute (or gap-fill) missing daily precipitation data. Study area was the Upper Colorado Water Resources Region (UCWRR), over a time period of 2008-2017. In terms of data, the package consists of raw precipitation measurements from Natural Resources Conservation Service (NRCS) network. These raw measurements are compiled into a single csv file (NRCS_dates_2008_2017.csv). The corresponding metadata are compiled in Metadata_processed.csv. The package also consists of outputs generated during Baseline and Sequential imputation runs as documented in Mital et al., 2020. Scripts used to generate all the results and figures for Mital et al., 2020 are also uploaded. A readme file documenting the layout of the data archive is also uploaded, and a KML file shows the geographic range for the data.

54 ENVIRONMENTAL SCIENCES↗