Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Imputation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Linkage map construction using limited parental genotypic information

Abstract Genetic linkage maps based on single nucleotide polymorphisms (SNPs) represent an essential tool for a variety of genomic analyses. Today, next-generation sequencing (NGS) enables rapid genotyping of different mapping populations based on thousands of SNPs and the construction of highly saturated linkage maps. Nevertheless, missing data in the genotyping of the parental lines creates a bottleneck that determines the number of SNPs that can be used for the linkage map. As a proof of concept, a highly saturated genetic linkage map was constructed using the imputed genotypic data of a recombinant inbred line (RIL) population and the limited genotypic information of its parental lines. Two ABH genotype files were created from a pseudo-parental genotypic data set that includes all the SNPs present in the RIL population. In the first ABH file pseudo-parental 1 was considered parental A, while in the second pseudo-parental 1 was considered parental B. These two duplicate ABH genotype files were merged by chromosome and subjected to linkage map analysis. Since the ABH data were duplicated, two mirrored linkage groups were generated per chromosome. The correct linkage map was identified and selected based on the partial genotypic data of the parental lines. This strategy was effective for constructing a highly saturated linkage map of 33,421 SNPs based on the genotyping of 205 RILs and a limited number of 100 SNPs present in the parental lines. This strategy enables the use of all the NGS SNP data obtained from a low-coverage sequencing experiment in the mapping population.

59 BASIC BIOLOGICAL SCIENCES↗

Predictive analytics of selections of russet potatoes

We explore the application of machine learning algorithms specifically to enhance the selection process of Russet potato (Solanum tuberosum L.) clones in breeding trials by predicting their suitability for advancement. This study addresses the challenge of efficiently identifying high-yield, disease-resistant, and climate-resilient potato varieties that meet processing industry standards. Leveraging manually collected data from trials in the state of Oregon, we investigate the potential of a wide variety of state-of-the-art binary classification models. The dataset includes 1086 clones, with data on 38 attributes recorded for each clone, focusing on yield, size, appearance, and frying characteristics, with several control varieties planted consistently across four Oregon regions from 2013 to 2021. We conduct a comprehensive analysis of the dataset that includes preprocessing, feature engineering, and imputation to address missing values. We focus on several key metrics such as accuracy, F1-score, and Matthews correlation coefficient (MCC) for model evaluation. The top-performing models, namely a feedforward neural network classifier (Neural Net), a histogram-based gradient boosting classifier (HGBC), and a support vector machine classifier (SVM), demonstrate consistent and significant results. To further validate our findings, we conducted a simulation study using the aims, data-generating mechanisms, estimands, methods, and performance measures (ADEMP) framework, simulating different data-generating scenarios to assess model robustness and performance through true positive, true negative, false positive, and false negative distributions, area under the receiver operating characteristic curve (AUC-ROC) and MCC. The simulation results highlight that non-linear models like SVM and HGBC consistently show higher AUC-ROC and MCC than logistic regression, thus outperforming the traditional linear model across various distributions, and emphasizing the importance of model selection and tuning in agricultural trials. Variable selection further enhances model performance and identifies influential features in predicting trial outcomes. The findings emphasize the potential of machine learning in streamlining the selection process for potato varieties, offering benefits such as increased efficiency, substantial cost savings, and judicious resource utilization. Our study contributes insights into precision agriculture and showcases the relevance of advanced technologies for informed decision-making in breeding programs.

60 APPLIED LIFE SCIENCES↗

Challenging problems of quality assurance and quality control (QA/QC) of meteorological time series data

Abstract Representativeness and quality of collected meteorological data impact accuracy and precision of climate, hydrological, and biogeochemical analyses and predictions. We developed a comprehensive Quality Assurance (QA) and Quality Control (QC) statistical framework, consisting of three major phases: Phase I—Preliminary data exploration, i.e., processing of raw datasets, with the challenging problems of time formatting and combining datasets of different lengths and different time intervals; Phase II—QA of the datasets, including detecting and flagging of duplicates, outliers, and extreme data; and Phase III—the development of time series of a desired frequency, imputation of missing values, visualization and a final statistical summary. The paper includes two use cases based on the time series data collected at the Billy Barr meteorological station (East River Watershed, Colorado), and the Barro Colorado Island (BCI, Panama) meteorological station. The developed statistical framework is suitable for both real-time and post-data-collection QA/QC analysis of meteorological datasets.

54 ENVIRONMENTAL SCIENCES↗

Visibility-enhanced model-free deep reinforcement learning algorithm for voltage control in realistic distribution systems using smart inverters

Increasing integration of distributed solar photovoltaic (PV) into distribution networks could result in adverse effects on grid operation. Traditional model-based control algorithms require accurate model information that is difficult to acquire and thus are challenging to implement in practice. Here, this paper proposes a surrogate model-enabled grid visibility scheme to empower deep reinforcement learning (DRL) approach for distribution network voltage regulation using PV inverters with minimal system knowledge. In contrast to existing DRL methods, this paper presents and corroborates the adverse impact of missing load information on DRL performance and, based on this finding, proposes a surrogate model methodology to impute load information utilizing observable data. Additionally, a multi-fidelity neural network is utilized to construct the DRL training environment, chosen for its efficient data utilization and enhanced robustness to data uncertainty. The feasibility and effectiveness of the proposed algorithm are assessed by considering DRL testing across varying degrees of observable load information and diverse training environments on a realistic power system.

14 SOLAR ENERGY↗

Machine learning models of intermittent operation of RO wellhead water treatment for salinity reduction and nitrate removal

Machine learning models were developed for intermittent multi-mode operation of a wellhead reverse osmosis water purification and desalination system to predict salt passage, nitrate passage, and permeate flux. The models, based on long short-term memory (LSTM) recurrent neural network (RNN) architecture, included an attention mechanism to increase model performance in proximity of the regulatory limit for nitrate. Training and testing of the models for the Startup, Production, Shutdown and Flushing operational modes were based on operational data (consisting of 22 process variables per data sample) acquired every 2–5 s over a six-month period. The significant sets of model input attributes for the different operational modes were assessed via Spearman ranking correlation, Self-Organizing Map (SOM) analysis and feed forward feature selection (FFFS). Although the variability of nitrate passage, salt passage and permeate flux was significant over the four operational modes, prediction performance for the three outcomes were with R2 and Average Absolute Relative Error (AARE) of 0.78–0.95 and 2.96–6.16 %, respectively. Model updates post membrane elements replacement demonstrated similar levels of prediction accuracy. The study results suggest that there is merit in exploring the utility of multi-mode models for sensor fault detection, data imputation, and for potential use in model-predictive control.

Intermittent RO operation↗

Conditional distribution estimation of building characteristics with diffusion models for urban energy modeling

Understanding current energy consumption behavior in communities is critical for informing future energy use decisions and enabling efficient energy management. Urban energy models, which are used to simulate these energy use patterns, require large datasets with detailed building characteristics for accurate outcomes. However, such detailed characteristics at the individual building level are often unknown and costly to acquire, or unavailable. Through this work, we propose using a generative modeling approach to generate realistic building attributes to fill in the data gaps and finally provide complete characteristics as inputs to energy models. Our model learns complex, building-level patterns from training on a large-scale residential building stock model containing 2.2 million buildings. We employ a tabular diffusion-based framework that is designed to handle heterogeneous (discrete and continuous) features in tabular building data, such as occupancy, floor area, heating, cooling, and other equipment details. We develop a capability for conditional diffusion, enabling the imputation of missing building characteristics conditioned on known attributes. We conduct a comprehensive validation of our conditional diffusion model, firstly by comparing the generated conditional distributions against the underlying data distribution, and secondly, by performing a case study for a Baltimore residential region, showing the practical utility of our approach. Our work is one of the first to demonstrate the potential of generative modeling to accelerate building energy modeling workflows.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Hydrologic connectivity and dynamics of solute transport in a mountain stream: Insights from a long-term tracer test and multiscale transport modeling informed by machine learning

The movement of solutes in a watershed is a complex process with multiple interactions and feedbacks across spatial and temporal scales. Modeling the dynamics of solute transport along diverse hydrologic pathways within watersheds – from hillslopes to stream channels and in and out of the hyporheic zones – is challenging but critically important, as these processes integrate and contribute to the biogeochemical functioning of the river corridor up to the river network scale. Here we use results from a long-term network-scale tracer test at the H.J. Andrews experimental forest in western Cascade Mountains, Oregon, USA to inform a multiscale framework for transport in stream corridors. The framework uses a Lagrangian-based subgrid model to represent the effects of hyporheic exchange flow and advective transport at stream network scales. The spatially and temporally resolved stream discharge needed for the transport model is imputed across the river system by an entity-aware long short-term memory network. Modeled concentrations show good agreements with the observations and exhibit power scaling laws indicative of a very wide range of timescales over which hyporheic exchange flow occurs. Our results demonstrate a data-informed modeling framework that links dynamical processes occurring at small scales to a network context to help understand how changes at reach scale cascade into network-scale effects, providing a useful tool for sustainable river basin management.

54 ENVIRONMENTAL SCIENCES↗

Impact of duration and missing data on the long-term photovoltaic degradation rate estimation

Accurate quantification of photovoltaic (PV) system degradation rate (R D ) is essential for lifetime yield predictions. Although R D is a critical parameter, its estimation lacks a standardized methodology that can be applied on outdoor field data. The purpose of this paper is to investigate the impact of time period duration and missing data on R D by analyzing the performance of different techniques applied to synthetic PV system data at different linear R D patterns and known noise conditions. The analysis includes the application of different techniques to a 10-year synthetic dataset of a crystalline Silicon PV system, with emulated degradation levels and imputed missing data. Here, the analysis demonstrated that the accuracy of ordinary least squares (OLS), year-on-year (YOY), autoregressive integrated moving average (ARIMA) and robust principal component analysis (RPCA) techniques is affected by the evaluation duration with all techniques converging to lower R D deviations over the 10-year evaluation, apart from RPCA at high degradation levels. Moreover, the estimated R D is strongly affected by the amount of missing data. Filtering out the corrupted data yielded more accurate R D results for all techniques. It is proven that the application of a change-point detection stage is necessary and guidelines for accurate R D estimation are provided.

14 SOLAR ENERGY↗

OmicsMLMentor: A Web Application for Guided Machine Learning Analysis of Omics Data

Expression-based omics technologies (e.g. proteomics, metabolomics, transcriptomics, etc.) increasingly rely on supervised and unsupervised machine learning (ML) models to find key biomolecules distinguishing conditions, identify natural groupings in biological data, or generate predictions for outcomes of interest. Fitting ML models to omics data presents several challenges, including handling missing data, selecting a normalization method, choosing a valid model, and optimizing hyperparameters, all requiring statistical programming skills to address these challenges. Thus, the open-source web application SLOPE was designed to lower the barrier to ML modeling for omics data. SLOPE supports the fitting of 15 ML models (10 supervised and 5 unsupervised) tailored to omics datasets, such as proteomics, metabolomics, lipidomics, and transcriptomics. SLOPE offers several omics-specific features, including methods for handling missingness (imputation, conversion, removal), normalization tests, ranking of models based on the structure of a user’s data and user input, and optimal hyperparameter selections using cross-validation splits. By streamlining ML workflows for omics analysis, SLOPE address critical gaps in existing online web tools, facilitating a broader adoption of these models for omics research. Here, SLOPE is applied to data from a lignin exposure study to highlight the workflow for fitting both supervised and unsupervised models to data.

lipidomics↗

Insights into Prismatic Loop Formation in Irradiated Fe–Cr Alloys from Hypothesis-Driven Active Learning and Causal Analysis

Neutron and electron irradiation experimental studies conducted on body-centered cubic Fe and Fe–Cr alloys have established two prismatic dislocation loop populations, which have Burgers vectors of either a/2$\langle$111$\rangle$ or a$\langle$100$\rangle$. Here, the loop formation depends on factors such as dose (D), dose rate (D rt ), temperature (T), chromium content (Cr%), and other alloying elements. Hence, it is important to understand how irradiation-induced dislocation loops evolve conditional upon the loop characteristics, such as loop density (DD), average loop size d̅, and irradiation parameters (D, D rt , T, and irradiation type), which is still an active area of research. To understand these complex structure–property relationships, machine learning (ML) is employed in a three-step approach. This includes imputing missing data with a k-nearest neighbor, generating functionalized features, and assessing feature importance with random forest classification and regression. Physics-based features are incorporated in a hypothesis-driven active learning scheme to overcome data unavailability challenges. Insights obtained from ML models (i) to categorize dislocation loop types, show the highest correlation with d̅; (ii) Log(DD), obtained through mathematical formulations involving D, Cr%, d̅, and T (e.g., Log(DD) ~ D + exp(-Cr%) + 1/d̅ and log(DD) ~ D + exp(-Cr%) + 1/T). Hypothesis-driven active learning is able to predict Log(DD) in which the experimental date is not known. Causal models verify cause–effect relationships for dislocation loop classification and irradiation factors in FeCr alloys.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Automated Gold Nanorod Spectral Morphology Analysis Pipeline

The development of a colloidal synthesis procedure to produce nanomaterials with high shape and size purity is often a time-consuming, iterative process. This is often due to quantitative uncertainties in the required reaction conditions and the time, resources, and expertise intensive characterization methods required for quantitative determination of nanomaterial size and shape. Absorption spectroscopy is often the easiest method for colloidal nanomaterial characterization. However, due to the lack of a reliable method to extract nanoparticle shapes from absorption spectroscopy, it is generally treated as a more qualitative measure for metal nanoparticles. This work demonstrates a gold nanorod (AuNR) spectral morphology analysis tool, called AuNR-SMA, which is a fast and accurate method to extract quantitative structural information from colloidal AuNR absorption spectra. To demonstrate the practical utility of this model, we apply it to three distinct applications. First, we demonstrate this model's utility as an automated analysis tool in a high-throughput AuNR synthesis procedure by generating quantitative size information from optical spectra. Second, we use the predictions generated by this model to train a machine learning model to predict the resulting AuNR size distributions under specified reaction conditions. Third, we apply this model to spectra extracted from the literature where no size distributions are reported and impute unreported quantitative information on AuNR synthesis. This approach can potentially be extended to any other nanocrystal system where absorption spectra are size dependent, and accurate numerical simulation of absorption spectra is possible. In addition, this pipeline could be integrated into automated synthesis apparatuses to provide interpretable data from simple measurements, help explore the synthesis science of nanoparticles in a rational manner, or facilitate closed-loop workflows.

36 MATERIALS SCIENCE↗

A Coupled Deep Learning Model for Estimating Surface NO 2 Levels from Remote Sensing Data: 15-Year Study Over the Contiguous United States

This study proposes a novel two-step deep learning (DL) model for estimating surface NO 2 concentrations using satellite data over the contiguous United States (CONUS) from 2005 to 2019. The first phase of the model uses partial convolutional neural network (PCNN), an advanced DL model that accurately imputes gaps between surface NO 2 stations and creates 5,478 daily-mean NO 2 grids (PCNN-NO 2 ) of the 2005-2019 period over the study area. We then feed the PCNN-NO 2 , along with other predictor variables, into a deep neural network (DNN) to estimate surface NO 2 levels, achieving exceptional performance with a correlation coefficient of 0.975 to 0.978, a mean absolute bias of 0.99 ppb to 1.38 ppb, and a root mean square error of 1.47 ppb to 1.97 ppb. Spatial cross-validation results also indicate strong spatial performance of PCNN-DNN surface NO 2 estimates. In addition to its accurate estimates, the PCNN-DNN model consistently generates estimated NO 2 grids without any missing values, improving the quality of various applications such as emission reduction strategies and public health studies. Between 2005 and 2019, the 5,478 daily estimated NO 2 grids over the CONUS reveal significant reductions in NO 2 levels in fourteen major urban environments: Washington D.C. (-43%), New York (-45%), Los Angeles (-38%), Chicago (-25%), Boston (-43%), Houston (-34%), Dallas (-40%), Philadelphia (-41%), Phoenix (-38%), Detroit (-20%), Denver (-23%), Atlanta (-0.7%), Cincinnati (-38%), and Pittsburgh (-56%). Furthermore, the study shows that the denser urban regions that in-situ stations are installed in, the higher the difference between in-situ observations and regional-mean NO 2 levels.

54 ENVIRONMENTAL SCIENCES↗

Predicting Large‐Scale Systematic Missing Pipe Attributes in Water Distribution Networks

Water distribution network (WDN) models are an essential tool used by water utilities for hydraulic analysis. Unfortunately, missing data and insufficient resources often make creating and maintaining these models unfeasible. Existing methods to address missing pipe properties, like sequential imputation for missing values and reconstruction using graph metrics, are designed to accommodate random patterns of missing information and require a significant percentage of the system's attributes to be known. However, these data completeness assumptions do not always align with real‐world scenarios where large sections of the WDN model have missing data. To address this challenge, this study proposes a data‐driven approach for estimating pipe diameter when considering different spatial patterns and degrees of data completeness (i.e., 0%–90%). Using data from 16 WDNs in Kentucky, this study compares the use of machine learning (ML) using topological and geospatial features against an existing deterministic approach. Results demonstrate that WDN models with pipe diameters predicted by the proposed ML method had comparable hydraulic performance to the ground truth models. Moreover, results showed that ML method performance varies between WDNs of differing topological classification. Insights from this study help advance the ability to leverage partial data to create and maintain WDN models amid uncertainty and inadequate resources.

Poff, Jason W. [Oregon State Univ., Corvallis, OR ↗

A genomic data archive from the Network for Pancreatic Organ donors with Diabetes

The Network for Pancreatic Organ donors with Diabetes (nPOD) is the largest biorepository of human pancreata and associated immune organs from donors with type 1 diabetes (T1D), maturity-onset diabetes of the young (MODY), cystic fibrosis-related diabetes (CFRD), type 2 diabetes (T2D), gestational diabetes, islet autoantibody positivity (AAb+), and without diabetes. nPOD recovers, processes, analyzes, and distributes high-quality biospecimens, collected using optimized standard operating procedures, and associated de-identified data/metadata to researchers around the world. Herein describes the release of high-parameter genotyping data from this collection. 372 donors were genotyped using a custom precision medicine single nucleotide polymorphism (SNP) microarray. Data were technically validated using published algorithms to evaluate donor relatedness, ancestry, imputed HLA, and T1D genetic risk score. Additionally, 207 donors were assessed for rare known and novel coding region variants via whole exome sequencing (WES). These data are publicly-available to enable genotype-specific sample requests and the study of novel genotype:phenotype associations, aiding in the mission of nPOD to enhance understanding of diabetes pathogenesis to promote the development of novel therapies.

59 BASIC BIOLOGICAL SCIENCES↗

A novel data gaps filling method for solar PV output forecasting

This study proposes a modified gaps filling method, expanding the column mean imputation method and evaluated using randomly generated missing values comprising 5%, 10%, 15%, and 20% of the original data on power output. The XGBoost algorithm was implemented as a forecasting model using the original and processed datasets and two sources of solar radiation data, namely, Shortwave Radiation (SWR) from Advanced Himawari Imager 8 (AHI-8) and Surface Solar Radiation Downward (SSRD) from ERA5 global reanalysis data. Further, the accuracy of the two sets of forecasted power output was evaluated using Root Mean Square Error (RMSE) and Mean Absolute Error (MAE). Results show that by applying the proposed gap filling method and using SWR in forecasting solar photovoltaic (PV) output, the improvement in the RMSE and MAE values range from 12.52% to 24.30% and from 21.10% to 31.31%, respectively. Meanwhile, using SSRD, the improvement in the RMSE values range from 14.01% to 28.54% and MAE values from 22.39% to 35.53%. To further evaluate the accuracy of the proposed gap-filling method, the proposed method could be validated using different datasets and other forecasting methods. Future studies could also consider applying the said method to datasets with data gaps higher than 20%.

Energy & Fuels↗

Constructing a Simulation Surrogate with Partially Observed Output

Gaussian process surrogates are a popular alternative to directly using computationally expensive simulation models. When the simulation output consists of many responses, dimension-reduction techniques are often employed to construct these surrogates. However, surrogate methods with dimension reduction generally rely on complete output training data. This article proposes a new Gaussian process surrogate method that permits the use of partially observed output while remaining computationally efficient. The new method involves the imputation of missing values and the adjustment of the covariance matrix used for Gaussian process inference. The resulting surrogate represents the available responses, disregards the missing responses, and provides meaningful uncertainty quantification. In conclusion, the proposed approach is shown to offer sharper inference than alternatives in a simulation study and a case study where an energy density functional model that frequently returns incomplete output is calibrated.

42 ENGINEERING↗

Generating synthetic occupants for use in building performance simulation

Occupant behaviour simulation frameworks can employ synthetic populations to characterize occupancy and behavioural patterns in buildings based on observed demographic data at a certain geographical location. For buildings, very few synthetic occupant populations have been generated. This paper uses a Bayesian Networks (BN) structural learning approach to synthesize populations of occupants in a multi-family housing case study. Two additional cases of office occupants and senior housing residents are considered as a cross-case comparison. Furthermore, we draw upon the extended version of drivers-needs-actions-systems (DNAS) framework to guide the selection of variables and data imputation. Our results show that the BN approach is powerful in learning the structure of data sets. The synthetic data sets successfully match the joint distributions of the underlying combined data sets. Experiments on the multi-family housing particularly show better performance than the office and senior housing cases.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Pharmacoepidemiology, Machine Learning and COVID-19: An intent-to-treat analysis of hydroxychloroquine, with or without azithromycin, and COVID-19 outcomes amongst hospitalized US Veterans

Hydroxychloroquine (HCQ) was proposed as an early therapy for coronavirus disease 2019 (COVID-19) after in vitro studies indicated possible benefit. Previous in vivo observational studies have presented conflicting results, though recent randomized clinical trials have reported no benefit from HCQ amongst hospitalized COVID-19 patients. In this work, we examined the effects of HCQ alone, and in combination with azithromycin, in a hospitalized COVID-19 positive, United States (US) Veteran population using a propensity score adjusted survival analysis with imputation of missing data. From March 1, 2020 through April 30, 2020, 64,055 US Veterans were tested for COVID-19 based on Veteran Affairs Healthcare Administration electronic health record data. Of the 7,193 positive cases, 2,809 were hospitalized, and 657 individuals were prescribed HCQ within the first 48-hours of hospitalization for the treatment of COVID-19. There was no apparent benefit associated with HCQ receipt, alone or in combination with azithromycin, and an increased risk of intubation when used in combination with azithromycin [Hazard Ratio (95% Confidence Interval): 1.55 (1.07, 2.24)]. In conclusion, we assessed the effectiveness of HCQ with or without azithromycin in treating patients hospitalized with COVID-19 using a national sample of the US Veteran population. Using rigorous study design and analytic methods to reduce confounding and bias, we found no evidence of a survival benefit from the administration of HCQ.

60 APPLIED LIFE SCIENCES↗