Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “imputation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Insights into Prismatic Loop Formation in Irradiated Fe–Cr Alloys from Hypothesis-Driven Active Learning and Causal Analysis

Neutron and electron irradiation experimental studies conducted on body-centered cubic Fe and Fe–Cr alloys have established two prismatic dislocation loop populations, which have Burgers vectors of either a/2$\langle$111$\rangle$ or a$\langle$100$\rangle$. Here, the loop formation depends on factors such as dose (D), dose rate (D rt ), temperature (T), chromium content (Cr%), and other alloying elements. Hence, it is important to understand how irradiation-induced dislocation loops evolve conditional upon the loop characteristics, such as loop density (DD), average loop size d̅, and irradiation parameters (D, D rt , T, and irradiation type), which is still an active area of research. To understand these complex structure–property relationships, machine learning (ML) is employed in a three-step approach. This includes imputing missing data with a k-nearest neighbor, generating functionalized features, and assessing feature importance with random forest classification and regression. Physics-based features are incorporated in a hypothesis-driven active learning scheme to overcome data unavailability challenges. Insights obtained from ML models (i) to categorize dislocation loop types, show the highest correlation with d̅; (ii) Log(DD), obtained through mathematical formulations involving D, Cr%, d̅, and T (e.g., Log(DD) ~ D + exp(-Cr%) + 1/d̅ and log(DD) ~ D + exp(-Cr%) + 1/T). Hypothesis-driven active learning is able to predict Log(DD) in which the experimental date is not known. Causal models verify cause–effect relationships for dislocation loop classification and irradiation factors in FeCr alloys.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Automated Gold Nanorod Spectral Morphology Analysis Pipeline

The development of a colloidal synthesis procedure to produce nanomaterials with high shape and size purity is often a time-consuming, iterative process. This is often due to quantitative uncertainties in the required reaction conditions and the time, resources, and expertise intensive characterization methods required for quantitative determination of nanomaterial size and shape. Absorption spectroscopy is often the easiest method for colloidal nanomaterial characterization. However, due to the lack of a reliable method to extract nanoparticle shapes from absorption spectroscopy, it is generally treated as a more qualitative measure for metal nanoparticles. This work demonstrates a gold nanorod (AuNR) spectral morphology analysis tool, called AuNR-SMA, which is a fast and accurate method to extract quantitative structural information from colloidal AuNR absorption spectra. To demonstrate the practical utility of this model, we apply it to three distinct applications. First, we demonstrate this model's utility as an automated analysis tool in a high-throughput AuNR synthesis procedure by generating quantitative size information from optical spectra. Second, we use the predictions generated by this model to train a machine learning model to predict the resulting AuNR size distributions under specified reaction conditions. Third, we apply this model to spectra extracted from the literature where no size distributions are reported and impute unreported quantitative information on AuNR synthesis. This approach can potentially be extended to any other nanocrystal system where absorption spectra are size dependent, and accurate numerical simulation of absorption spectra is possible. In addition, this pipeline could be integrated into automated synthesis apparatuses to provide interpretable data from simple measurements, help explore the synthesis science of nanoparticles in a rational manner, or facilitate closed-loop workflows.

36 MATERIALS SCIENCE↗

A Coupled Deep Learning Model for Estimating Surface NO 2 Levels from Remote Sensing Data: 15-Year Study Over the Contiguous United States

This study proposes a novel two-step deep learning (DL) model for estimating surface NO 2 concentrations using satellite data over the contiguous United States (CONUS) from 2005 to 2019. The first phase of the model uses partial convolutional neural network (PCNN), an advanced DL model that accurately imputes gaps between surface NO 2 stations and creates 5,478 daily-mean NO 2 grids (PCNN-NO 2 ) of the 2005-2019 period over the study area. We then feed the PCNN-NO 2 , along with other predictor variables, into a deep neural network (DNN) to estimate surface NO 2 levels, achieving exceptional performance with a correlation coefficient of 0.975 to 0.978, a mean absolute bias of 0.99 ppb to 1.38 ppb, and a root mean square error of 1.47 ppb to 1.97 ppb. Spatial cross-validation results also indicate strong spatial performance of PCNN-DNN surface NO 2 estimates. In addition to its accurate estimates, the PCNN-DNN model consistently generates estimated NO 2 grids without any missing values, improving the quality of various applications such as emission reduction strategies and public health studies. Between 2005 and 2019, the 5,478 daily estimated NO 2 grids over the CONUS reveal significant reductions in NO 2 levels in fourteen major urban environments: Washington D.C. (-43%), New York (-45%), Los Angeles (-38%), Chicago (-25%), Boston (-43%), Houston (-34%), Dallas (-40%), Philadelphia (-41%), Phoenix (-38%), Detroit (-20%), Denver (-23%), Atlanta (-0.7%), Cincinnati (-38%), and Pittsburgh (-56%). Furthermore, the study shows that the denser urban regions that in-situ stations are installed in, the higher the difference between in-situ observations and regional-mean NO 2 levels.

54 ENVIRONMENTAL SCIENCES↗

Predicting Large‐Scale Systematic Missing Pipe Attributes in Water Distribution Networks

Water distribution network (WDN) models are an essential tool used by water utilities for hydraulic analysis. Unfortunately, missing data and insufficient resources often make creating and maintaining these models unfeasible. Existing methods to address missing pipe properties, like sequential imputation for missing values and reconstruction using graph metrics, are designed to accommodate random patterns of missing information and require a significant percentage of the system's attributes to be known. However, these data completeness assumptions do not always align with real‐world scenarios where large sections of the WDN model have missing data. To address this challenge, this study proposes a data‐driven approach for estimating pipe diameter when considering different spatial patterns and degrees of data completeness (i.e., 0%–90%). Using data from 16 WDNs in Kentucky, this study compares the use of machine learning (ML) using topological and geospatial features against an existing deterministic approach. Results demonstrate that WDN models with pipe diameters predicted by the proposed ML method had comparable hydraulic performance to the ground truth models. Moreover, results showed that ML method performance varies between WDNs of differing topological classification. Insights from this study help advance the ability to leverage partial data to create and maintain WDN models amid uncertainty and inadequate resources.

Poff, Jason W. [Oregon State Univ., Corvallis, OR ↗

A genomic data archive from the Network for Pancreatic Organ donors with Diabetes

The Network for Pancreatic Organ donors with Diabetes (nPOD) is the largest biorepository of human pancreata and associated immune organs from donors with type 1 diabetes (T1D), maturity-onset diabetes of the young (MODY), cystic fibrosis-related diabetes (CFRD), type 2 diabetes (T2D), gestational diabetes, islet autoantibody positivity (AAb+), and without diabetes. nPOD recovers, processes, analyzes, and distributes high-quality biospecimens, collected using optimized standard operating procedures, and associated de-identified data/metadata to researchers around the world. Herein describes the release of high-parameter genotyping data from this collection. 372 donors were genotyped using a custom precision medicine single nucleotide polymorphism (SNP) microarray. Data were technically validated using published algorithms to evaluate donor relatedness, ancestry, imputed HLA, and T1D genetic risk score. Additionally, 207 donors were assessed for rare known and novel coding region variants via whole exome sequencing (WES). These data are publicly-available to enable genotype-specific sample requests and the study of novel genotype:phenotype associations, aiding in the mission of nPOD to enhance understanding of diabetes pathogenesis to promote the development of novel therapies.

59 BASIC BIOLOGICAL SCIENCES↗

A novel data gaps filling method for solar PV output forecasting

This study proposes a modified gaps filling method, expanding the column mean imputation method and evaluated using randomly generated missing values comprising 5%, 10%, 15%, and 20% of the original data on power output. The XGBoost algorithm was implemented as a forecasting model using the original and processed datasets and two sources of solar radiation data, namely, Shortwave Radiation (SWR) from Advanced Himawari Imager 8 (AHI-8) and Surface Solar Radiation Downward (SSRD) from ERA5 global reanalysis data. Further, the accuracy of the two sets of forecasted power output was evaluated using Root Mean Square Error (RMSE) and Mean Absolute Error (MAE). Results show that by applying the proposed gap filling method and using SWR in forecasting solar photovoltaic (PV) output, the improvement in the RMSE and MAE values range from 12.52% to 24.30% and from 21.10% to 31.31%, respectively. Meanwhile, using SSRD, the improvement in the RMSE values range from 14.01% to 28.54% and MAE values from 22.39% to 35.53%. To further evaluate the accuracy of the proposed gap-filling method, the proposed method could be validated using different datasets and other forecasting methods. Future studies could also consider applying the said method to datasets with data gaps higher than 20%.

Energy & Fuels↗

Constructing a Simulation Surrogate with Partially Observed Output

Gaussian process surrogates are a popular alternative to directly using computationally expensive simulation models. When the simulation output consists of many responses, dimension-reduction techniques are often employed to construct these surrogates. However, surrogate methods with dimension reduction generally rely on complete output training data. This article proposes a new Gaussian process surrogate method that permits the use of partially observed output while remaining computationally efficient. The new method involves the imputation of missing values and the adjustment of the covariance matrix used for Gaussian process inference. The resulting surrogate represents the available responses, disregards the missing responses, and provides meaningful uncertainty quantification. In conclusion, the proposed approach is shown to offer sharper inference than alternatives in a simulation study and a case study where an energy density functional model that frequently returns incomplete output is calibrated.

42 ENGINEERING↗

Generating synthetic occupants for use in building performance simulation

Occupant behaviour simulation frameworks can employ synthetic populations to characterize occupancy and behavioural patterns in buildings based on observed demographic data at a certain geographical location. For buildings, very few synthetic occupant populations have been generated. This paper uses a Bayesian Networks (BN) structural learning approach to synthesize populations of occupants in a multi-family housing case study. Two additional cases of office occupants and senior housing residents are considered as a cross-case comparison. Furthermore, we draw upon the extended version of drivers-needs-actions-systems (DNAS) framework to guide the selection of variables and data imputation. Our results show that the BN approach is powerful in learning the structure of data sets. The synthetic data sets successfully match the joint distributions of the underlying combined data sets. Experiments on the multi-family housing particularly show better performance than the office and senior housing cases.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Pharmacoepidemiology, Machine Learning and COVID-19: An intent-to-treat analysis of hydroxychloroquine, with or without azithromycin, and COVID-19 outcomes amongst hospitalized US Veterans

Hydroxychloroquine (HCQ) was proposed as an early therapy for coronavirus disease 2019 (COVID-19) after in vitro studies indicated possible benefit. Previous in vivo observational studies have presented conflicting results, though recent randomized clinical trials have reported no benefit from HCQ amongst hospitalized COVID-19 patients. In this work, we examined the effects of HCQ alone, and in combination with azithromycin, in a hospitalized COVID-19 positive, United States (US) Veteran population using a propensity score adjusted survival analysis with imputation of missing data. From March 1, 2020 through April 30, 2020, 64,055 US Veterans were tested for COVID-19 based on Veteran Affairs Healthcare Administration electronic health record data. Of the 7,193 positive cases, 2,809 were hospitalized, and 657 individuals were prescribed HCQ within the first 48-hours of hospitalization for the treatment of COVID-19. There was no apparent benefit associated with HCQ receipt, alone or in combination with azithromycin, and an increased risk of intubation when used in combination with azithromycin [Hazard Ratio (95% Confidence Interval): 1.55 (1.07, 2.24)]. In conclusion, we assessed the effectiveness of HCQ with or without azithromycin in treating patients hospitalized with COVID-19 using a national sample of the US Veteran population. Using rigorous study design and analytic methods to reduce confounding and bias, we found no evidence of a survival benefit from the administration of HCQ.

60 APPLIED LIFE SCIENCES↗

High-resolution mapping reveals hotspots and sex-biased recombination in Populus trichocarpa

Abstract Fine-scale meiotic recombination is fundamental to the outcome of natural and artificial selection. Here, dense genetic mapping and haplotype reconstruction were used to estimate recombination for a full factorial Populus trichocarpa cross of 7 males and 7 females. Genomes of the resulting 49 full-sib families (N = 829 offspring) were resequenced, and high-fidelity biallelic SNP/INDELs and pedigree information were used to ascertain allelic phase and impute progeny genotypes to recover gametic haplotypes. The 14 parental genetic maps contained 1,820 SNP/INDELs on average that covered 376.7 Mb of physical length across 19 chromosomes. Comparison of parental and progeny haplotypes allowed fine-scale demarcation of cross-over regions, where 38,846 cross-over events in 1,658 gametes were observed. Cross-over events were positively associated with gene density and negatively associated with GC content and long-terminal repeats. One of the most striking findings was higher rates of cross-overs in males in 8 out of 19 chromosomes. Regions with elevated male cross-over rates had lower gene density and GC content than windows showing no sex bias. High-resolution analysis identified 67 candidate cross-over hotspots spread throughout the genome. DNA sequence motifs enriched in these regions showed striking similarity to those of maize, Arabidopsis, and wheat. These findings, and recombination estimates, will be useful for ongoing efforts to accelerate domestication of this and other biomass feedstocks, as well as future studies investigating broader questions related to evolutionary history, perennial development, phenology, wood formation, vegetative propagation, and dioecy that cannot be studied using annual plant model systems.

59 BASIC BIOLOGICAL SCIENCES↗

Identifying strengths and weaknesses of methods for computational network inference from single-cell RNA-seq data

Single-cell RNA-sequencing (scRNA-seq) offers unparalleled insight into the transcriptional programs of different cellular states by measuring the transcriptome of thousands of individual cells. An emerging problem in the analysis of scRNA-seq is the inference of transcriptional gene regulatory networks and a number of methods with different learning frameworks have been developed to address this problem. Here, we present an expanded benchmarking study of eleven recent network inference methods on seven published scRNA-seq datasets in human, mouse, and yeast considering different types of gold standard networks and evaluation metrics. We evaluate methods based on their computing requirements as well as on their ability to recover the network structure. We find that, while most methods have a modest recovery of experimentally derived interactions based on global metrics such as Area Under the Precision Recall curve, methods are able to capture targets of regulators that are relevant to the system under study. Among the top performing methods that use only expression were SCENIC, PIDC, MERLIN or Correlation. Addition of prior biological knowledge and the estimation of transcription factor activities resulted in the best overall performance with the Inferelator and MERLIN methods that use prior knowledge outperforming methods that use expression alone. We found that imputation for network inference did not improve network inference accuracy and could be detrimental. Comparisons of inferred networks for comparable bulk conditions showed that the networks inferred from scRNA-seq datasets are often better or at par with the networks inferred from bulk datasets. Our analysis should be beneficial in selecting methods for network inference. At the same time, this highlights the need for improved methods and better gold standards for regulatory network inference from scRNAseq datasets.

59 BASIC BIOLOGICAL SCIENCES↗

Integrating Intermediate Traits in Phylogenetic Genotype-to-Phenotype Studies

A major goal of research in evolution and genetics is linking genotype to phenotype. This work could be direct, such as determining the genetic basis of a phenotype by leveraging genetic variation or divergence in a developmental, physiological, or behavioral trait. The work could also involve studying the evolutionary phenomena (e.g., reproductive isolation, adaptation, sexual dimorphism, behavior) that reveal an indirect link between genotype and a trait of interest. When the phenotype diverges across evolutionarily distinct lineages, this genotype-to-phenotype problem can be addressed using phylogenetic genotype-to-phenotype (PhyloG2P) mapping, which uses genetic signatures and convergent phenotypes on a phylogeny to infer the genetic bases of traits. The PhyloG2P approach has proven powerful in revealing key genetic changes associated with diverse traits, including the mammalian transition to marine environments and transitions between major mechanisms of photosynthesis. However, there are several intermediate traits layered in between genotype and the phenotype of interest, including but not limited to transcriptional profiles, chromatin states, protein abundances, structures, modifications, metabolites, and physiological parameters. Each intermediate trait is interesting and informative in its own right, but synthesis across data types has great promise for providing a deep, integrated, and predictive understanding of how genotypes drive phenotypic differences and convergence. We argue that an expanded PhyloG2P framework (the PhyloG2P matrix) that explicitly considers intermediate traits, and imputes those that are prohibitive to obtain, will allow a better mechanistic understanding of any trait of interest. Furthermore, this approach provides a proxy for functional validation and mechanistic understanding in organisms where laboratory manipulation is impractical.

59 BASIC BIOLOGICAL SCIENCES↗

Redshifts of radio sources in the Million Quasars Catalogue from machine learning

ABSTRACT With the aim of using machine learning techniques to obtain photometric redshifts based upon a source’s radio spectrum alone, we have extracted the radio sources from the Million Quasars Catalogue. Of these, 44 119 have a spectroscopic redshift, required for model validation, and for which photometry could be obtained. Using the radio spectral properties as features, we fail to find a model which can reliably predict the redshifts, although there is the suggestion that the models improve with the size of the training sample. Using the near-infrared–optical–ultraviolet bands magnitudes, we obtain reliable predictions based on the 12 503 radio sources which have all of the required photometry. From the 80:20 training–validation split, this gives only 2501 validation sources, although training the sample upon our previous SDSS model gives comparable results for all 12 503 sources. This makes us confident that SkyMapper, which will survey southern sky in the u, v, g, r, i, z bands, can be used to predict the redshifts of radio sources detected with the Square Kilometre Array. By using machine learning to impute the magnitudes missing from much of the sample, we can predict the redshifts for 32 698 sources, an increase from 28 to 74 per cent of the sample, at the cost of increasing the outlier fraction by a factor of 1.4. While the ‘optical’ band data prove successful, at this stage we cannot rule out the possibility of a radio photometric redshift, given sufficient data which may be necessary to overcome the relatively featureless radio spectra.

79 ASTRONOMY AND ASTROPHYSICS↗

Integration of ultra-low coverage whole-genome sequences for reconstructing the evolutionary history of Galapagos giant tortoises

Genomic data from contemporary and historical samples often need to be coupled for evolutionary reconstructions of multitaxon complexes. However, the genetic data recovered from historical samples may result only in ultra-low coverage whole-genome sequences (ulcWGS; <0.15× depth), leading to inaccurate evolutionary inferences given a preponderance of missing data. Using the Galapagos giant tortoise radiation as a study system (Chelonoidis spp., composed of 13 extant and four extinct lineages), we assembled a novel methodological pipeline that removes potential noise introduced by the missing data and enhances the evolutionary signal from ulcWGS samples. We leveraged existing tools for phylogenomic placement (EPA-ng), population genomic structure (smartsnp) and admixture (Admixfrog, NGSadmix) to demonstrate that the evolutionary history of samples can be uncovered with sequencing depths as low as 0.008–0.139×. Importantly, these approaches do not use genotype imputation of the ulcWGS samples, which would require extensive reference datasets. Our application to two cases of extinct lineages of Galapagos giant tortoises, with and without references from the same lineage, demonstrates the general value of the approach. We confirm where the extinct lineages from San Cristóbal and Santa Fe islands fit into the Galapagos giant tortoise radiation, and that these lineages were evolutionarily distinct entities.

ancient DNA↗

The Impact of Time-Aware Design Choices in ICS Anomaly Detection

Industrial control systems (ICS) remain vulnerable to increasingly sophisticated cyberattacks, yet evaluating anomaly detection models in these environments is challenging due to temporal dependencies, missing-not-at-random patterns, and extremely imbalanced datasets. These factors make common practices—especially random data splits and na¨ıve imputation— prone to severe temporal leakage, which can inflate reported performance and obscure real-world limitations. In this work, we systematically examine classical machine learning models, temporal deep learning architecture, and tensordecomposition– based methods on a gas-pipeline dataset using a fully temporally separated evaluation pipeline designed to mimic realistic deployment conditions. Our findings show that proper temporal handling and MNAR-aware preprocessing significantly alter the relative performance of popular anomaly-detection methods, providing practical guidance for designing reliable, leakage-resistant ICS intrusion-detection systems.

97 MATHEMATICS AND COMPUTING↗

Detection of Stealthy False Data Injection Attacks in Unobservable Distribution Networks

In this paper, a composite scheme is proposed for detecting stealthy data manipulation attacks on distribution system which is unobservable with standard least squares based state estimators. This technique has three stages where the process of data imputation, voltage phasor estimation and the bad data detection are carried out in a systematic manner. The proposed approach is then integrated with moving target defense strategies which perturbs the network parameters to reveal stealthy false data injection attacks. The proposed approach is tested is validated on a three-phase, unbalanced 37-node distribution system and its results are presented. It is shown that the proposed approach has the ability to accurately detect the presence of FDI attacks using limited measurements (i.e., the test system is unobservable).

Rajasekaran, James K.↗

Model-Free State Estimation Using Low-Rank Canonical Polyadic Decomposition

As electric grids experience high penetration levels of renewable generation, fundamental changes are required to address real-time situational awareness. Here, we utilize unique traits of tensors to devise a model-free situational awareness and energy forecasting framework for distribution networks. This work formulates the state of the network at multiple time instants as a three-way tensor; hence, recovering full state information of the network is tantamount to estimating all the values of the tensor. Given measurements received from µphasor measurement units and/or smart meters, the recovery of unobserved quantities is carried out using the low-rank canonical polyadic decomposition of the state tensor—that is, the state estimation task is posed as a tensor imputation problem utilizing observed patterns in measured (sampled) quantities. Two structured sampling schemes are considered, namely, asynchronous slab and fiber sampling. For both schemes, we present sufficient conditions on the number of sampled slabs and fibers that guarantee identifiability of the factors of the state tensor. Numerical results demonstrate the ability of the proposed framework to achieve high estimation accuracy in multiple sampling scenarios.

42 ENGINEERING↗

FAIRification, Quality Assessment, and Missingness Pattern Discovery for Spatiotemporal Photovoltaic Data

The growth of the photovoltaic market has pushed the demand for power forecasting and performance evaluation for a huge population of PV power plants. Many of these power plants have spatiotemporal coherence that can be utilized for improving model accuracy. We have demonstrated in this paper the FAIRification of spatiotemporal PV time series data. Through the creation of a solar power plant ontology, we propose standards for the naming and structure of metadata used to describe the data from these power plants. Using the structure from this ontology, we have developed both R and Python packages for the automation of the FAIRification process. Going further, we have also developed an R package that automates the analysis of the quality of a data set through the designation of letter grades. To solve the issue of data missingness, we propose the use of St-GNN autoencoders to detect and impute missing values from a data set by utilizing data from power plants nearby.

14 SOLAR ENERGY↗