Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data imputation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Design Choices in Anomaly Detection for Industrial Control Systems: Insights from Gas Pipeline Data

Industrial control systems (ICS) remain vulnerable to increasingly sophisticated cyberattacks, yet evaluating anomaly detection models in these environments is challenging due to temporal dependencies, missing-not-at-random patterns, and extremely imbalanced datasets. These factors make common practices—especially random data splits and naïve imputation—prone to severe temporal leakage, which can inflate reported performance and obscure real-world limitations. In this work, we systematically examine classical machine learning models, temporal deep learning architecture, and tensor-decomposition–based methods on a gas-pipeline dataset using a fully temporally separated evaluation pipeline designed to mimic realistic deployment conditions. Our findings show that proper temporal handling and MNAR-aware preprocessing significantly alter the relative performance of popular anomaly-detection methods, providing practical guidance for designing reliable, leakage-resistant ICS intrusion-detection systems.

97 MATHEMATICS AND COMPUTING↗

Visibility-enhanced model-free deep reinforcement learning algorithm for voltage control in realistic distribution systems using smart inverters

Increasing integration of distributed solar photovoltaic (PV) into distribution networks could result in adverse effects on grid operation. Traditional model-based control algorithms require accurate model information that is difficult to acquire and thus are challenging to implement in practice. Here, this paper proposes a surrogate model-enabled grid visibility scheme to empower deep reinforcement learning (DRL) approach for distribution network voltage regulation using PV inverters with minimal system knowledge. In contrast to existing DRL methods, this paper presents and corroborates the adverse impact of missing load information on DRL performance and, based on this finding, proposes a surrogate model methodology to impute load information utilizing observable data. Additionally, a multi-fidelity neural network is utilized to construct the DRL training environment, chosen for its efficient data utilization and enhanced robustness to data uncertainty. The feasibility and effectiveness of the proposed algorithm are assessed by considering DRL testing across varying degrees of observable load information and diverse training environments on a realistic power system.

14 SOLAR ENERGY↗

The Impact of Time-Aware Design Choices in ICS Anomaly Detection

Industrial control systems (ICS) remain vulnerable to increasingly sophisticated cyberattacks, yet evaluating anomaly detection models in these environments is challenging due to temporal dependencies, missing-not-at-random patterns, and extremely imbalanced datasets. These factors make common practices—especially random data splits and na¨ıve imputation— prone to severe temporal leakage, which can inflate reported performance and obscure real-world limitations. In this work, we systematically examine classical machine learning models, temporal deep learning architecture, and tensordecomposition– based methods on a gas-pipeline dataset using a fully temporally separated evaluation pipeline designed to mimic realistic deployment conditions. Our findings show that proper temporal handling and MNAR-aware preprocessing significantly alter the relative performance of popular anomaly-detection methods, providing practical guidance for designing reliable, leakage-resistant ICS intrusion-detection systems.

97 MATHEMATICS AND COMPUTING↗

AUTONOMIE VID

Autonomie Vehicle Information Database (VID) offers a comprehensive list of vehicle specifications since 1990. The database details more than 65,000 vehicles with hundreds of attributes. The database is the result of the development of a general automated data collection framework as well as the development of building blocks for processing, cleaning, integrating and analyzing complex data. The data has undergone several layers of outlier detections processes, machine learning based imputations methods have been used to deal with missing data problems, and new fields have been created according to the rules of feature engineering. Thanks to this streamlined data pipelines, the resulting processed aggregated data should deliver a unique level of information to the user in which the content can be efficiently maintained and updated.

Moswd, Ayman↗

Revised monthly energy generation estimates for 1,500 hydroelectric power plants in the United States

Abstract The U.S. Energy Information Administration (EIA) conducts a regular survey (form EIA-923) to collect annual and monthly net generation for more than ten thousand U.S. power plants. Approximately 90% of the ~1,500 hydroelectric plants included in this data release are surveyed at annual resolution only and thus lack actual observations of monthly generation. For each of these plants, EIA imputes monthly generation values using the combined monthly generating pattern of other hydropower plants within the corresponding census division. The imputation method neglects local hydrology and reservoir operations, rendering the monthly data unsuitable for various research applications. Here we present an alternative approach to disaggregate each unobserved plant’s reported annual generation using proxies of monthly generation—namely historical monthly reservoir releases and average river discharge rates recorded downstream of each dam. Evaluation of the new dataset demonstrates substantial and robust improvement over the current imputation method, particularly if reservoir release data are available. The new dataset—named RectifHyd—provides an alternative to EIA-923 for U.S. scale, plant-level, monthly hydropower net generation (2001–2020). RectifHyd may be used to support power system studies or analyze within-year hydropower generation behavior at various spatial scales.

13 HYDRO ENERGY↗

Identifying Light-Duty Vehicle Travel from Large-Scale Multimodal Wearable GPS Data with Novelty Detection Algorithms

Identifying travel mode within travel survey data sets, especially light-duty vehicle (LDV) travel, is foundational, though nontrivial, to travel behavior analysis and fuel consumption estimation. Current travel mode detection approaches require well-sampled and balanced data sets with ground truth travel mode labels. They are rarely applied and validated on large-scale, real-world data sets, which may not satisfy the data requirements. This paper proposes an LDV travel mode detection model as a supplement to current travel mode detection methods, for the case when the training set is highly (and/or completely) unbalanced, to the extent that classical machine-learning approaches become difficult or impossible to deploy. The proposed model uses a novelty detection technique-one-class support vector machines (OCSVMs)-and a novel exhaustive feature extraction (EFE) technique on continuous time series data (i.e., Global Positioning System [GPS] speed profiles) for single-mode trip trajectories. Training and validation of the model are conducted on a large-scale, real-world data set. The proposed method accurately identifies LDV trips from a broad set of multimodal trips by leveraging a wealth of preexisting in-vehicle GPS travel data. Additional sensitivity analysis sheds light on the optimal training size, which will benefit applications limited by highly imbalanced data. The paper also discusses performance comparison with regular machine-learning approaches, the model's robustness, and the potential to extend the proposed model to multimodal prediction.

47 OTHER INSTRUMENTATION↗

Scripts and raw/processed precipitation NRCS network data, Upper Colorado, 2008-2017

This data package consists of scripts and data that were used to generate all the results and figures for Mital et al., 2020 (https://doi.org/10.3389/frwa.2020.00020). The purpose of the study was to develop a new algorithm to impute (or gap-fill) missing daily precipitation data. Study area was the Upper Colorado Water Resources Region (UCWRR), over a time period of 2008-2017. In terms of data, the package consists of raw precipitation measurements from Natural Resources Conservation Service (NRCS) network. These raw measurements are compiled into a single csv file (NRCS_dates_2008_2017.csv). The corresponding metadata are compiled in Metadata_processed.csv. The package also consists of outputs generated during Baseline and Sequential imputation runs as documented in Mital et al., 2020. Scripts used to generate all the results and figures for Mital et al., 2020 are also uploaded. A readme file documenting the layout of the data archive is also uploaded, and a KML file shows the geographic range for the data.

54 ENVIRONMENTAL SCIENCES↗

A review of imputation strategies for isobaric labeling-based shotgun proteomics

The throughput efficiency and increased depth of coverage provided by isobaric-labeled proteomics measurements have led to increased usage of these techniques. However, the structure of missing data is uniquely different than unlabeled studies. In this review, we compare the efficacy of nine imputation methods on a CPTAC proteomics iTRAQ dataset. Imputation methods were evaluated with regard to accuracy, variability, statistical hypothesis test inference and run time over datasets consisting of varying number of iTRAQ plexes and percentages of missing data. In general, expectation maximization and random forest imputation methods yielded the best performances, and constant-based methods performed poorly consistently across all dataset sizes and percentages of missing values. For datasets with small sample sizes and higher percentages of missing data, results indicate that statistical inference with no imputation may be preferable. Based on the findings in this review, there are core imputation methods that perform higher for isobaric-labeled proteomics data, but great care and consideration as to whether imputation should be used should be given for datasets comprised of a small number of samples, as well as to factors such as computational time and reproducibility of imputation values.

Bramer, Lisa M.↗

Observed and Imputed Volumetric Soil Water Content Timeseries for the New Mexico Elevation Gradient

Reliable soil water content (SWC) data are essential for understanding dryland ecosystem dynamics, but high-frequency SWC sensors often fail, creating gaps in critical datasets. To address this, we developed a Bayesian mixture model that imputes missing SWC using both linear interpolation and an ecosystem water balance model (SOILWAT2), tested across six AmeriFlux eddy covariance tower sites in the New Mexico Elevation Gradient, demonstrating its effectiveness in reconstructing SWC patterns while providing insights into the factors driving SWC variability. Daily volumetric soil water content (SWC) data are provided as csv-formatted spreadsheets for the six AmeriFlux sites (US-Seg, US-Ses, US-Wjs, US-Mpi, US-Vcp, and US-Vcs). For each site there is an observed SWC file (site_SWC_gapfill.csv) and a file that contains imputed SWC (imputed_SWC_site.csv). The observed SWC files contain temperature corrected sensor values, tower precipitation data, as well as outputs from SOILWAT2 simulations that were used to impute SWC. The imputed files contain the original observed SWC values and the imputed missing SWC values. When SWC was missing from the original data, the missing value was imputed based on the Bayesian imputation mixture model. The posterior mean of all imputed values is reported as "mean_X". When the observed SWC was NOT missing, mean_X = observed SWC value (original data). The standard deviation, 2.5th percentile and the 97.5th percentile for the imputed values are also reported in the imputed files. There are readme text files for each file type explaining the contents of each column.

54 ENVIRONMENTAL SCIENCES↗

A sorghum practical haplotype graph facilitates genome‐wide imputation and cost‐effective genomic prediction

Abstract Successful management and utilization of increasingly large genomic datasets is essential for breeding programs to accelerate cultivar development. To help with this, we developed a Sorghum bicolor Practical Haplotype Graph (PHG) pangenome database that stores haplotypes and variant information. We developed two PHGs in sorghum that were used to identify genome‐wide variants for 24 founders of the Chibas sorghum breeding program from 0.01x sequence coverage. The PHG called single nucleotide polymorphisms (SNPs) with 5.9% error at 0.01x coverage—only 3% higher than PHG error when calling SNPs from 8x coverage sequence. Additionally, 207 progenies from the Chibas genomic selection (GS) training population were sequenced and processed through the PHG. Missing genotypes were imputed from PHG parental haplotypes and used for genomic prediction. Mean prediction accuracies with PHG SNP calls range from .57–.73 and are similar to prediction accuracies obtained with genotyping‐by‐sequencing or targeted amplicon sequencing (rhAmpSeq) markers. This study demonstrates the use of a sorghum PHG to impute SNPs from low‐coverage sequence data and shows that the PHG can unify genotype calls across multiple sequencing platforms. By reducing input sequence requirements, the PHG can decrease the cost of genotyping, make GS more feasible, and facilitate larger breeding populations. Our results demonstrate that the PHG is a useful research and breeding tool that maintains variant information from a diverse group of taxa, stores sequence data in a condensed but readily accessible format, unifies genotypes across genotyping platforms, and provides a cost‐effective option for genomic selection.

Jensen, Sarah E.↗

Generation and evaluation of synthetic patient data

Background: Machine learning (ML) has made a significant impact in medicine and cancer research; however, its impact in these areas has been undeniably slower and more limited than in other application domains. A major reason for this has been the lack of availability of patient data to the broader ML research community, in large part due to patient privacy protection concerns. High-quality, realistic, synthetic datasets can be leveraged to accelerate methodological developments in medicine. By and large, medical data is high dimensional and often categorical. These characteristics pose multiple modeling challenges. Methods: In this paper, we evaluate three classes of synthetic data generation approaches; probabilistic models, classification-based imputation models, and generative adversarial neural networks. Metrics for evaluating the quality of the generated synthetic datasets are presented and discussed. Results: While the results and discussions are broadly applicable to medical data, for demonstration purposes we generate synthetic datasets for cancer based on the publicly available cancer registry data from the Surveillance Epidemiology and End Results (SEER) program. Specifically, our cohort consists of breast, respiratory, and non-solid cancer cases diagnosed between 2010 and 2015, which includes over 360,000 individual cases. Conclusions: We discuss the trade-offs of the different methods and metrics, providing guidance on considerations for the generation and usage of medical synthetic data.

59 BASIC BIOLOGICAL SCIENCES↗

Anomaly Detection and Approximate Similarity Searches of Transients in Real-time Data Streams

Abstract We present Lightcurve Anomaly Identification and Similarity Search ( LAISS ), an automated pipeline to detect anomalous astrophysical transients in real-time data streams. We deploy our anomaly detection model on the nightly Zwicky Transient Facility (ZTF) Alert Stream via the ANTARES broker, identifying a manageable ∼1–5 candidates per night for expert vetting and coordinating follow-up observations. Our method leverages statistical light-curve and contextual host galaxy features within a random forest classifier, tagging transients of rare classes ( spectroscopic anomalies), of uncommon host galaxy environments ( contextual anomalies), and of peculiar or interaction-powered phenomena ( behavioral anomalies). Moreover, we demonstrate the power of a low-latency (∼ms) approximate similarity search method to find transient analogs with similar light-curve evolution and host galaxy environments. We use analogs for data-driven discovery, characterization, (re)classification, and imputation in retrospective and real-time searches. To date, we have identified ∼50 previously known and previously missed rare transients from real-time and retrospective searches, including but not limited to superluminous supernovae (SLSNe), tidal disruption events, SNe IIn, SNe IIb, SNe I-CSM, SNe Ia-91bg-like, SNe Ib, SNe Ic, SNe Ic-BL, and M31 novae. Lastly, we report the discovery of 325 total transients, all observed between 2018 and 2021 and absent from public catalogs (∼1% of all ZTF Astronomical Transient reports to the Transient Name Server through 2021). These methods enable a systematic approach to finding the “needle in the haystack” in large-volume data streams. Because of its integration with the ANTARES broker, LAISS is built to detect exciting transients in Rubin data.

79 ASTRONOMY AND ASTROPHYSICS↗

Low-rank Tensor Completion for PMU Data Recovery

This paper proposes a tensor completion method for the recovery of missing phasor measurement unit (PMU) measurements. Tensor completion as the general case of matrix completion has attracted increasing attention in recent years. The imputation accuracy for the existing matrix completion methods may be significantly reduced when there are consecutive data losses across multiple data channels. To tackle this issue, we explore the multi-way characteristics of PMU measurements by using a tensor model. We leverage the low-rank property of the PMU measurements and formulate the missing PMU data recovery problem as a low-rank tensor completion problem. An efficient algorithm based on alternating direction method of multipliers (ADMM) is developed to solve the tensor completion problem. The experiments using the real PMU dataset show that the proposed method exhibits better imputation accuracy compared with the conventional data recovery methods.

Ghasemkhani, Amir↗

LandScan Mosaic

The LandScan program at Oak Ridge National Laboratory (ORNL), in collaboration with the National Geospatial-Intelligence Agency (NGA), continues to deliver the most accurate and up to date global, high resolution gridded population data. Additionally, the latest advancements in the LandScan HD methodology led to reduced latency in development of rapid updates for geopolitical events. With momentum towards reporting more up to date population estimates, feedback from the user community expressed interest in reporting population estimates in ranges - whether to express a level of uncertainty or confirm to leadership and stakeholders the modeled data are estimates. Building upon the need to understand uncertainty or confidence in the modeled data and report ranges at the global scale, LandScan Mosaic was developed. LandScan Mosaic represents the next generation of high-resolution population modeling, building upon the established success of previous LandScan HD iterations. While LandScan HD employed a deterministic big data fusion approach, LandScan Mosaic enhances this methodology by integrating advanced machine learning techniques to impute missing, yet crucial, population model parameters. This advancement allows for probabilistic modeling of building occupancy and population distribution, incorporating uncertainty quantification through Monte Carlo sampling methods. By combining big data fusion with machine learning-driven imputation and stochastic modeling, LandScan Mosaic provides a more comprehensive and robust representation of population dynamics. LandScan Mosaic will be following the in the footsteps of its longstanding counterpart LandScan Global and releasing a global gridded population raster, at the 3-arcsecond resolution. This technical report documents the current stage of development of LandScan Mosaic, detailing the methodologies and data sources behind the modeling. Stakeholders are encouraged to use this document as an authoritative reference for insight into Mosaic’s data development processes. However, readers should note that LandScan Mosaic remains in a late-stage research and development phase, and methodologies and data presented here are subject to refinements ahead of the anticipated global release in Summer 2025. Feedback and inquiries from users and stakeholders are welcomed as we continue to refine and enhance this important population resource.

97 MATHEMATICS AND COMPUTING↗

Cleaning Images with Gaussian Process Regression

Many approaches to astronomical data reduction and analysis cannot tolerate missing data: corrupted pixels must first have their values imputed. This paper presents astrofix, a robust and flexible image imputation algorithm based on Gaussian process regression. Through an optimization process, astrofix chooses and applies a different interpolation kernel to each image, using a training set extracted automatically from that image. It naturally handles clusters of bad pixels and image edges and adapts to various instruments and image types. For bright pixels, the mean absolute error of astrofix is several times smaller than that of median replacement and interpolation by a Gaussian kernel. We demonstrate good performance on both imaging and spectroscopic data, including the SBIG 6303 0.4 m telescope and the FLOYDS spectrograph of Las Cumbres Observatory and the CHARIS integral-field spectrograph on the Subaru Telescope.

42 ENGINEERING↗

Challenging problems of quality assurance and quality control (QA/QC) of meteorological time series data

Abstract Representativeness and quality of collected meteorological data impact accuracy and precision of climate, hydrological, and biogeochemical analyses and predictions. We developed a comprehensive Quality Assurance (QA) and Quality Control (QC) statistical framework, consisting of three major phases: Phase I—Preliminary data exploration, i.e., processing of raw datasets, with the challenging problems of time formatting and combining datasets of different lengths and different time intervals; Phase II—QA of the datasets, including detecting and flagging of duplicates, outliers, and extreme data; and Phase III—the development of time series of a desired frequency, imputation of missing values, visualization and a final statistical summary. The paper includes two use cases based on the time series data collected at the Billy Barr meteorological station (East River Watershed, Colorado), and the Barro Colorado Island (BCI, Panama) meteorological station. The developed statistical framework is suitable for both real-time and post-data-collection QA/QC analysis of meteorological datasets.

54 ENVIRONMENTAL SCIENCES↗

A novel data gaps filling method for solar PV output forecasting

This study proposes a modified gaps filling method, expanding the column mean imputation method and evaluated using randomly generated missing values comprising 5%, 10%, 15%, and 20% of the original data on power output. The XGBoost algorithm was implemented as a forecasting model using the original and processed datasets and two sources of solar radiation data, namely, Shortwave Radiation (SWR) from Advanced Himawari Imager 8 (AHI-8) and Surface Solar Radiation Downward (SSRD) from ERA5 global reanalysis data. Further, the accuracy of the two sets of forecasted power output was evaluated using Root Mean Square Error (RMSE) and Mean Absolute Error (MAE). Results show that by applying the proposed gap filling method and using SWR in forecasting solar photovoltaic (PV) output, the improvement in the RMSE and MAE values range from 12.52% to 24.30% and from 21.10% to 31.31%, respectively. Meanwhile, using SSRD, the improvement in the RMSE values range from 14.01% to 28.54% and MAE values from 22.39% to 35.53%. To further evaluate the accuracy of the proposed gap-filling method, the proposed method could be validated using different datasets and other forecasting methods. Future studies could also consider applying the said method to datasets with data gaps higher than 20%.

Energy & Fuels↗

A genomic data archive from the Network for Pancreatic Organ donors with Diabetes

The Network for Pancreatic Organ donors with Diabetes (nPOD) is the largest biorepository of human pancreata and associated immune organs from donors with type 1 diabetes (T1D), maturity-onset diabetes of the young (MODY), cystic fibrosis-related diabetes (CFRD), type 2 diabetes (T2D), gestational diabetes, islet autoantibody positivity (AAb+), and without diabetes. nPOD recovers, processes, analyzes, and distributes high-quality biospecimens, collected using optimized standard operating procedures, and associated de-identified data/metadata to researchers around the world. Herein describes the release of high-parameter genotyping data from this collection. 372 donors were genotyped using a custom precision medicine single nucleotide polymorphism (SNP) microarray. Data were technically validated using published algorithms to evaluate donor relatedness, ancestry, imputed HLA, and T1D genetic risk score. Additionally, 207 donors were assessed for rare known and novel coding region variants via whole exome sequencing (WES). These data are publicly-available to enable genotype-specific sample requests and the study of novel genotype:phenotype associations, aiding in the mission of nPOD to enhance understanding of diabetes pathogenesis to promote the development of novel therapies.

59 BASIC BIOLOGICAL SCIENCES↗