Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Imputation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

High-resolution mapping reveals hotspots and sex-biased recombination in Populus trichocarpa

Abstract Fine-scale meiotic recombination is fundamental to the outcome of natural and artificial selection. Here, dense genetic mapping and haplotype reconstruction were used to estimate recombination for a full factorial Populus trichocarpa cross of 7 males and 7 females. Genomes of the resulting 49 full-sib families (N = 829 offspring) were resequenced, and high-fidelity biallelic SNP/INDELs and pedigree information were used to ascertain allelic phase and impute progeny genotypes to recover gametic haplotypes. The 14 parental genetic maps contained 1,820 SNP/INDELs on average that covered 376.7 Mb of physical length across 19 chromosomes. Comparison of parental and progeny haplotypes allowed fine-scale demarcation of cross-over regions, where 38,846 cross-over events in 1,658 gametes were observed. Cross-over events were positively associated with gene density and negatively associated with GC content and long-terminal repeats. One of the most striking findings was higher rates of cross-overs in males in 8 out of 19 chromosomes. Regions with elevated male cross-over rates had lower gene density and GC content than windows showing no sex bias. High-resolution analysis identified 67 candidate cross-over hotspots spread throughout the genome. DNA sequence motifs enriched in these regions showed striking similarity to those of maize, Arabidopsis, and wheat. These findings, and recombination estimates, will be useful for ongoing efforts to accelerate domestication of this and other biomass feedstocks, as well as future studies investigating broader questions related to evolutionary history, perennial development, phenology, wood formation, vegetative propagation, and dioecy that cannot be studied using annual plant model systems.

59 BASIC BIOLOGICAL SCIENCES↗

Identifying strengths and weaknesses of methods for computational network inference from single-cell RNA-seq data

Single-cell RNA-sequencing (scRNA-seq) offers unparalleled insight into the transcriptional programs of different cellular states by measuring the transcriptome of thousands of individual cells. An emerging problem in the analysis of scRNA-seq is the inference of transcriptional gene regulatory networks and a number of methods with different learning frameworks have been developed to address this problem. Here, we present an expanded benchmarking study of eleven recent network inference methods on seven published scRNA-seq datasets in human, mouse, and yeast considering different types of gold standard networks and evaluation metrics. We evaluate methods based on their computing requirements as well as on their ability to recover the network structure. We find that, while most methods have a modest recovery of experimentally derived interactions based on global metrics such as Area Under the Precision Recall curve, methods are able to capture targets of regulators that are relevant to the system under study. Among the top performing methods that use only expression were SCENIC, PIDC, MERLIN or Correlation. Addition of prior biological knowledge and the estimation of transcription factor activities resulted in the best overall performance with the Inferelator and MERLIN methods that use prior knowledge outperforming methods that use expression alone. We found that imputation for network inference did not improve network inference accuracy and could be detrimental. Comparisons of inferred networks for comparable bulk conditions showed that the networks inferred from scRNA-seq datasets are often better or at par with the networks inferred from bulk datasets. Our analysis should be beneficial in selecting methods for network inference. At the same time, this highlights the need for improved methods and better gold standards for regulatory network inference from scRNAseq datasets.

59 BASIC BIOLOGICAL SCIENCES↗

Integrating Intermediate Traits in Phylogenetic Genotype-to-Phenotype Studies

A major goal of research in evolution and genetics is linking genotype to phenotype. This work could be direct, such as determining the genetic basis of a phenotype by leveraging genetic variation or divergence in a developmental, physiological, or behavioral trait. The work could also involve studying the evolutionary phenomena (e.g., reproductive isolation, adaptation, sexual dimorphism, behavior) that reveal an indirect link between genotype and a trait of interest. When the phenotype diverges across evolutionarily distinct lineages, this genotype-to-phenotype problem can be addressed using phylogenetic genotype-to-phenotype (PhyloG2P) mapping, which uses genetic signatures and convergent phenotypes on a phylogeny to infer the genetic bases of traits. The PhyloG2P approach has proven powerful in revealing key genetic changes associated with diverse traits, including the mammalian transition to marine environments and transitions between major mechanisms of photosynthesis. However, there are several intermediate traits layered in between genotype and the phenotype of interest, including but not limited to transcriptional profiles, chromatin states, protein abundances, structures, modifications, metabolites, and physiological parameters. Each intermediate trait is interesting and informative in its own right, but synthesis across data types has great promise for providing a deep, integrated, and predictive understanding of how genotypes drive phenotypic differences and convergence. We argue that an expanded PhyloG2P framework (the PhyloG2P matrix) that explicitly considers intermediate traits, and imputes those that are prohibitive to obtain, will allow a better mechanistic understanding of any trait of interest. Furthermore, this approach provides a proxy for functional validation and mechanistic understanding in organisms where laboratory manipulation is impractical.

59 BASIC BIOLOGICAL SCIENCES↗

Redshifts of radio sources in the Million Quasars Catalogue from machine learning

ABSTRACT With the aim of using machine learning techniques to obtain photometric redshifts based upon a source’s radio spectrum alone, we have extracted the radio sources from the Million Quasars Catalogue. Of these, 44 119 have a spectroscopic redshift, required for model validation, and for which photometry could be obtained. Using the radio spectral properties as features, we fail to find a model which can reliably predict the redshifts, although there is the suggestion that the models improve with the size of the training sample. Using the near-infrared–optical–ultraviolet bands magnitudes, we obtain reliable predictions based on the 12 503 radio sources which have all of the required photometry. From the 80:20 training–validation split, this gives only 2501 validation sources, although training the sample upon our previous SDSS model gives comparable results for all 12 503 sources. This makes us confident that SkyMapper, which will survey southern sky in the u, v, g, r, i, z bands, can be used to predict the redshifts of radio sources detected with the Square Kilometre Array. By using machine learning to impute the magnitudes missing from much of the sample, we can predict the redshifts for 32 698 sources, an increase from 28 to 74 per cent of the sample, at the cost of increasing the outlier fraction by a factor of 1.4. While the ‘optical’ band data prove successful, at this stage we cannot rule out the possibility of a radio photometric redshift, given sufficient data which may be necessary to overcome the relatively featureless radio spectra.

79 ASTRONOMY AND ASTROPHYSICS↗

Integration of ultra-low coverage whole-genome sequences for reconstructing the evolutionary history of Galapagos giant tortoises

Genomic data from contemporary and historical samples often need to be coupled for evolutionary reconstructions of multitaxon complexes. However, the genetic data recovered from historical samples may result only in ultra-low coverage whole-genome sequences (ulcWGS; <0.15× depth), leading to inaccurate evolutionary inferences given a preponderance of missing data. Using the Galapagos giant tortoise radiation as a study system (Chelonoidis spp., composed of 13 extant and four extinct lineages), we assembled a novel methodological pipeline that removes potential noise introduced by the missing data and enhances the evolutionary signal from ulcWGS samples. We leveraged existing tools for phylogenomic placement (EPA-ng), population genomic structure (smartsnp) and admixture (Admixfrog, NGSadmix) to demonstrate that the evolutionary history of samples can be uncovered with sequencing depths as low as 0.008–0.139×. Importantly, these approaches do not use genotype imputation of the ulcWGS samples, which would require extensive reference datasets. Our application to two cases of extinct lineages of Galapagos giant tortoises, with and without references from the same lineage, demonstrates the general value of the approach. We confirm where the extinct lineages from San Cristóbal and Santa Fe islands fit into the Galapagos giant tortoise radiation, and that these lineages were evolutionarily distinct entities.

ancient DNA↗

The Impact of Time-Aware Design Choices in ICS Anomaly Detection

Industrial control systems (ICS) remain vulnerable to increasingly sophisticated cyberattacks, yet evaluating anomaly detection models in these environments is challenging due to temporal dependencies, missing-not-at-random patterns, and extremely imbalanced datasets. These factors make common practices—especially random data splits and na¨ıve imputation— prone to severe temporal leakage, which can inflate reported performance and obscure real-world limitations. In this work, we systematically examine classical machine learning models, temporal deep learning architecture, and tensordecomposition– based methods on a gas-pipeline dataset using a fully temporally separated evaluation pipeline designed to mimic realistic deployment conditions. Our findings show that proper temporal handling and MNAR-aware preprocessing significantly alter the relative performance of popular anomaly-detection methods, providing practical guidance for designing reliable, leakage-resistant ICS intrusion-detection systems.

97 MATHEMATICS AND COMPUTING↗

Detection of Stealthy False Data Injection Attacks in Unobservable Distribution Networks

In this paper, a composite scheme is proposed for detecting stealthy data manipulation attacks on distribution system which is unobservable with standard least squares based state estimators. This technique has three stages where the process of data imputation, voltage phasor estimation and the bad data detection are carried out in a systematic manner. The proposed approach is then integrated with moving target defense strategies which perturbs the network parameters to reveal stealthy false data injection attacks. The proposed approach is tested is validated on a three-phase, unbalanced 37-node distribution system and its results are presented. It is shown that the proposed approach has the ability to accurately detect the presence of FDI attacks using limited measurements (i.e., the test system is unobservable).

Rajasekaran, James K.↗

FAIRification, Quality Assessment, and Missingness Pattern Discovery for Spatiotemporal Photovoltaic Data

The growth of the photovoltaic market has pushed the demand for power forecasting and performance evaluation for a huge population of PV power plants. Many of these power plants have spatiotemporal coherence that can be utilized for improving model accuracy. We have demonstrated in this paper the FAIRification of spatiotemporal PV time series data. Through the creation of a solar power plant ontology, we propose standards for the naming and structure of metadata used to describe the data from these power plants. Using the structure from this ontology, we have developed both R and Python packages for the automation of the FAIRification process. Going further, we have also developed an R package that automates the analysis of the quality of a data set through the designation of letter grades. To solve the issue of data missingness, we propose the use of St-GNN autoencoders to detect and impute missing values from a data set by utilizing data from power plants nearby.

14 SOLAR ENERGY↗

Bayesian Framework for Multi-Timescale State Estimation in Low-Observable Distribution Systems

To support the smart grid paradigm, there has been a significant increase in sensor deployments and metering infrastructure in distribution systems. However, the measurements provided by these sensors and metering devices are typically sampled at different rates and could suffer from losses during the aggregation process. It is crucial to effectively reconcile the time-series measurements for a reliable state estimation. While weighted least squares has been the traditional approach for state estimation, sparsity-based approaches like matrix completion have become popular due to their superior performance in low-observability conditions. This paper proposes a Bayesian framework for both multi-timescale data aggregation and matrix completion based state estimation. Specifically, the multiscale time-series data aggregated from heterogenous sources are reconciled using a multitask Gaussian process that exploits the spatio-temporal correlations. Here, the resulting consistent timeseries alongwith the confidence bound on the imputations are fed into a Bayesian matrix completion method augmented with linearized power-flow constraints to accurately estimate the states in low-observability conditions. Results on three phase unbalanced IEEE 37 and IEEE 123 bus test systems reveal the superior performance of the proposed Bayesian framework. The computational complexity for the proposed Bayesian framework is also quantified.

42 ENGINEERING↗

PVplr-stGNN 0.1.10

PV Performance Loss Rate Estimation using Spatio-temporal Graph Neural Networks PVplr-stGNN is a Python 3 package developed by the SDLE Research Center at Case Western Reserve University in Cleveland OH. This repository contains the full source PVplr-stGNN package. The package contains the PV-stGAE for missingness data detection and imputation and PV-DynGNN for PLR estimation.

Fan, Yangxin [Case Western Reserve Univ., Clevelan↗

Phasor-Measurement-Unit-Based Data Analytics Using Digital Twin and PhasorAnalytics Software

A major objective of this project was to apply GE’s commercial machine learning and data analytics toolsets to large-scale, real-world, anonymized Phasor Measurement Unit (PMU) datasets in order to extract signatures, correlated and/or causal factors, and precursor patterns associated with significant power system phenomena. The project had a particular emphasis on extraction of insights relevant to asset health monitoring, real-time load modeling and cybersecurity monitoring. Additionally, the team was directed to undertake a comprehensive data quality analysis for the provided datasets and encouraged to estimate the ‘machine-learning readiness’ of the datasets by documenting any major obstacles to the application of commercial machine learning algorithms. To accomplish the aforementioned objectives, the project team’s work centered around the identification of key event signatures and application of the identified event signatures for event detection and event classification. The industry-validated, semi-supervised machine learning strategy employed for event signature identification involved several major tasks, including data-preprocessing, generation of an overabundance of features, normal data identification, normality modeling, and event signature identification through a methodical, quantitative ranking of features in order of relevance to each studied event type. Throughout the project, data quality issues and mitigation techniques were investigated. In this report, insights are provided regarding the readiness of the provided synchrophasor datasets for application of machine learning and data analytics. The methodologies employed for this technical strategy are summarized in this report. With regards to data preprocessing and feature generation, the provided Training and Test Datasets were ingested into GE’s big data environment. Subsequently, the team applied bad data cleansing and data imputation scripts, event detection scripts, and application programming interfaces (APIs) to the datasets for convenient data access. The project team completed development and validation of dozens of physics-based, statistics-based and transformation-based feature functions used for the extraction of over 60 synchrophasor features. Using a new parallel feature generation technology developed on this project, over 60 features have been rapidly generated for the full two years’ worth of Training and Test Dataset data associated with both the Eastern and Western interconnects. Even accommodating for temporal down-sampling inherent to the feature extraction procedure, this parallel feature generation activity resulted in a massive feature set with a storage requirement approximately equal to that of the raw training dataset itself. With regards to normal data identification and normality modeling, a normality model was built using the feature data extracted from the Training Dataset and iteratively refined subsequent to incremental adjustments and expansions of the Training Dataset feature data. With respect to event characterization and signature identification, an event signature identification pipeline was developed and used in conjunction with the normality model to identify over 15 event signatures for key event categories within the Training Dataset. The identified event signatures were used to characterize hundreds of key events in terms of relative severity, duration, and location of the event. An investigation was undertaken to identify correlated and causal factors involved in transformer events. A separate investigation into temporal trends in ring-down analysis results was undertaken to determine possible associations between system dynamics and various other factors such as loading, season or year. To validate the identified event signatures, additional work was undertaken to develop signature-based anomaly detection and classification tools suitable for convenient application to the synchrophasor datasets. The anomaly detection and classification tools, suitable for online application, were then applied to the entirety of the Eastern Interconnect Training and Test Datasets. Performance of the event detection and classification tools was evaluated upon receipt of the Test Dataset event logs (i.e., the labels for events contained in the Test Dataset), and promising results were obtained despite several challenges (documented herein) associated with application of supervised or semi-supervised machine learning methods to large-scale, anonymized datasets. Finally, the detection and classification tools were used to detect, classify, and characterize thousands of new events not included in the original event logs provided by the DOE within both the Training and Test Datasets.

24 POWER TRANSMISSION AND DISTRIBUTION↗

FAIRification, Quality Assessment, and Missingness Pattern Discovery for Spatiotemporal Photovoltaic Data

The ongoing growth of the photovoltaic market has pushed the demand for power forecasting and performance evaluation for a huge population of PV power plants. Through access to a large number of time series data sets from different power plants, we have found common issues that impede the modeling process. Namely, the time series data are hard to transfer between groups due to differences in variable nomenclature, and the quality of the data sets can vary. We address the issue of variable nomenclature by FAIRifying spatiotemporal PV time series data. Through the creation of a solar power plant ontology, we propose standards for the naming and structure of metadata used to describe the data from these power plants. Using the structure from this ontology, we have developed both R and Python packages for the automation of the FAIRification process. We have also developed an R package that automates the analysis of the quality of a data set through the designation of letter grades. With access to large time series data sets across many power plants, we can utilize spatiotemporal coherence between the sites in order to improve the quality of our data. To solve the issue of data missingness, we propose the use of Spatiotemporal-GNN autoencoders to detect and impute missing values from a data set by utilizing data from power plants nearby.

14 SOLAR ENERGY↗

Jobs, jobs, jobs: what’s an analyst to do?

Analysts and economists often face the task of using employment metrics to characterize industries of interest. Some key challenges can be understanding where to find employment metrics, the differences in various employment metrics, and when each metric should be used. This article analyzes a variety of publicly available employment data for the United States and compares these data. A detailed description of the intricacies of each data source is provided, which covers factors such as regionality, industry breakout, periodicity, and the types of jobs included. This article provides several case study examples, using the oil and gas extraction, coal mining, and chemical manufacturing sectors to portray challenges data users may face when developing employment estimates that suit their needs. Data users should be aware of a variety of data sources to understand alternative analysis options when data limitations are present and to determine which data source best meets their needs. Instances may occur in which information from one dataset may be used to help impute missing values.

99 GENERAL AND MISCELLANEOUS↗

Who Has the Upper HAND in a Hurricane? Flood Risk, Weather Shocks, and Property Prices

Flood risk is increasing in the United States, but how property markets respond to this hazard is unclear. Focusing on Houston, Texas, we analyze changes in flood risk capitalization in property prices after three major Gulf Coast hurricanes. Here, we compare two measures of flood risk–floodplain designations from the Federal Emergency Management Agency (FEMA), and a hydrologically-imputed, continuous measure, Height Above the Nearest Drainage (HAND). HAND affects property prices after a storm in intuitive ways, while we observe counterintuitive post-storm price premiums within FEMA floodplains. Plausibly exogenous measures like HAND may be useful tools for characterizing flood risk in economic analyses.

Plough, Julian A. [University of North Carolina, C↗

Physics-Informed Deep Learning for Reconstruction of Spatial Missing Climate Information in the Antarctic

Understanding the influence of the Antarctic on the global climate is crucial for the prediction of global warming. However, due to very few observation sites, it is difficult to reconstruct the rational spatial pattern by filling in the missing values from the limited site observations. To tackle this challenge, regional spatial gap-filling methods, such as Kriging and inverse distance weighted (IDW), are regularly used in geoscience. Nevertheless, the reconstructing credibility of these methods is undesirable when the spatial structure has massive missing pieces. Inspired by image inpainting, we propose a novel deep learning method that demonstrates a good effect by embedding the physics-aware initialization of deep learning methods for rapid learning and capturing the spatial dependence for the high-fidelity imputation of missing areas. We create the benchmark dataset that artificially masks the Antarctic region with ratios of 30%, 50% and 70%. The reconstructing monthly mean surface temperature using the deep learning image inpainting method RFR (Recurrent Feature Reasoning) exhibits an average of 63% and 71% improvement of accuracy over Kriging and IDW under different missing rates. With regard to wind speed, there are still 36% and 50% improvements. In particular, the achieved improvement is even better for the larger missing ratio, such as under the 70% missing rate, where the accuracy of RFR is 68% and 74% higher than Kriging and IDW for temperature and also 38% and 46% higher for wind speed. In addition, the PI-RFR (Physics-Informed Recurrent Feature Reasoning) method we proposed is initialized using the spatial pattern data simulated by the numerical climate model instead of the unified average. Compared with RFR, PI-RFR has an average accuracy improvement of 10% for temperature and 9% for wind speed. When applied to reconstruct the spatial pattern based on the Antarctic site observations, where the missing rate is over 90%, the proposed method exhibits more spatial characteristics than Kriging and IDW.

54 ENVIRONMENTAL SCIENCES↗

Functional Data Analysis for Extracting the Intrinsic Dimensionality of Spectra: Application to Chemical Homogeneity in the Open Cluster M67

High-resolution spectroscopic surveys of the Milky Way have entered the Big Data regime and have opened avenues for solving outstanding questions in Galactic archeology. However, exploiting their full potential is limited by complex systematics, whose characterization has not received much attention in modern spectroscopic analyses. In this work, we present a novel method to disentangle the component of spectral data space intrinsic to the stars from that due to systematics. Using functional principal component analysis on a sample of 18,933 giant spectra from APOGEE, we find that the intrinsic structure above the level of observational uncertainties requires ≈10 functional principal components (FPCs). Our FPCs can reduce the dimensionality of spectra, remove systematics, and impute masked wavelengths, thereby enabling accurate studies of stellar populations. To demonstrate the applicability of our FPCs, we use them to infer stellar parameters and abundances of 28 giants in the open cluster M67. We employ Sequential Neural Likelihood, a simulation-based Bayesian inference method that learns likelihood functions using neural density estimators, to incorporate non-Gaussian effects in spectral likelihoods. By hierarchically combining the inferred abundances, we limit the spread of the following elements in M67: Fe ≲ 0.02 dex; C ≲ 0.03 dex; O, Mg, Si, Ni ≲ 0.04 dex; Ca ≲ 0.05 dex; N, Al ≲ 0.07 dex (at 68% confidence). Our constraints suggest a lack of self-pollution by core-collapse supernovae in M67, which has promising implications for the future of chemical tagging to understand the star formation history and dynamical evolution of the Milky Way.

79 ASTRONOMY AND ASTROPHYSICS↗

Anomaly Detection and Approximate Similarity Searches of Transients in Real-time Data Streams

Abstract We present Lightcurve Anomaly Identification and Similarity Search ( LAISS ), an automated pipeline to detect anomalous astrophysical transients in real-time data streams. We deploy our anomaly detection model on the nightly Zwicky Transient Facility (ZTF) Alert Stream via the ANTARES broker, identifying a manageable ∼1–5 candidates per night for expert vetting and coordinating follow-up observations. Our method leverages statistical light-curve and contextual host galaxy features within a random forest classifier, tagging transients of rare classes ( spectroscopic anomalies), of uncommon host galaxy environments ( contextual anomalies), and of peculiar or interaction-powered phenomena ( behavioral anomalies). Moreover, we demonstrate the power of a low-latency (∼ms) approximate similarity search method to find transient analogs with similar light-curve evolution and host galaxy environments. We use analogs for data-driven discovery, characterization, (re)classification, and imputation in retrospective and real-time searches. To date, we have identified ∼50 previously known and previously missed rare transients from real-time and retrospective searches, including but not limited to superluminous supernovae (SLSNe), tidal disruption events, SNe IIn, SNe IIb, SNe I-CSM, SNe Ia-91bg-like, SNe Ib, SNe Ic, SNe Ic-BL, and M31 novae. Lastly, we report the discovery of 325 total transients, all observed between 2018 and 2021 and absent from public catalogs (∼1% of all ZTF Astronomical Transient reports to the Transient Name Server through 2021). These methods enable a systematic approach to finding the “needle in the haystack” in large-volume data streams. Because of its integration with the ANTARES broker, LAISS is built to detect exciting transients in Rubin data.

79 ASTRONOMY AND ASTROPHYSICS↗

A graph neural network (GNN) approach to basin-scale river network learning: the role of physics-based connectivity and data fusion

Abstract. Rivers and river habitats around the world are under sustained pressure from human activities and the changing global environment. Our ability to quantify and manage the river states in a timely manner is critical for protecting the public safety and natural resources. In recent years, vector-based river network models have enabled modeling of large river basins at increasingly fine resolutions, but are computationally demanding. This work presents a multistage, physics-guided, graph neural network (GNN) approach for basin-scale river network learning and streamflow forecasting. During training, we train a GNN model to approximate outputs of a high-resolution vector-based river network model; we then fine-tune the pretrained GNN model with streamflow observations. We further apply a graph-based, data-fusion step to correct prediction biases. The GNN-based framework is first demonstrated over a snow-dominated watershed in the western United States. A series of experiments are performed to test different training and imputation strategies. Results show that the trained GNN model can effectively serve as a surrogate of the process-based model with high accuracy, with median Kling–Gupta efficiency (KGE) greater than 0.97. Application of the graph-based data fusion further reduces mismatch between the GNN model and observations, with as much as 50 % KGE improvement over some cross-validation gages. To improve scalability, a graph-coarsening procedure is introduced and is demonstrated over a much larger basin. Results show that graph coarsening achieves comparable prediction skills at only a fraction of training cost, thus providing important insights into the degree of physical realism needed for developing large-scale GNN-based river network models.

54 ENVIRONMENTAL SCIENCES↗