Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data imputation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Improved Particle Heat Transfer by way of Bimodal Particle Distributions for High Temperature Solar Thermal Energy

High temperature solar thermal facilities are looking to increase operating temperatures through novel heat transfer media, one such being solid particles. These particles operating at high temperatures will require transferring their thermal energy into another working fluid like supercritical carbon dioxide which can be used in advanced power cycles. Achieving high heat transfer between the particles and supercritical carbon dioxide is essential to high efficiency and low-cost operation. Therefore, optimizing the thermal conductivity of these particles is one potential way to ensure high performance. Traditionally, unimodal particle distributions have been employed in high temperature particle solar power plants. However, ambient temperature testing of bimodal particle distributions has revealed a superior thermal conductivity when compared to its unimodal counterpart at the same temperature. This data was obtained by certified, off-the-shelf instruments that can effectively simulate the conditions a particle would be exposed to in a high temperature solar thermal system. Data obtained in this way suggests that the increased thermal conductively imputed by a bimodal particle distribution is significant at working temperatures in solar facilities. Furthermore, the thermal conductivity of these bimodal particle distributions peaks when the best combination of large and small particles is applied. At high temperatures, binary particle distributions are compared to monodispersed distributions of larger particles where heat transfer is more prolific due to the increased surface radiation. Various thermal conductivity, porosity and heat exchanger models are explored in conjunction with data acquired up to 700 C.

Stout, Dallin (ORCID:0009000294586091)↗

Fission Product Gas Monitoring During Fuel Drying Operations - 20094

The goal of the project was to devise and implement a method of accurately qualifying any noble gas emitted during fuel drying operations (part of the dry fuel storage process) in order to provide the data needed to validate the off-site dose calculations. The imputes of the project was the NRC publishing Information Notice 18-01, 'Noble Fission Gas Releases During Spent Fuel Cask Loading Operations' in February 2018. The information notice describes some industry events that occurred during vacuum drying operations as part of a dry fuel campaign. The solution was devised by Fort Calhoun Station staff in partnership with Mirion Technologies technical staff. The monitoring was accomplished by directing all the effluent from the fuel storage cask vacuum drying machine through a sample chamber containing a compact ISOCS characterized CZT based gamma spectrometer connected to a Mirion Data Analyst. This resulted near real-time quantification of the radionuclides of concern. (authors)

07 ISOTOPE AND RADIATION SOURCES↗

Quality Assurance and Quality Control (QA/QC) of Meteorological Time Series Data for Billy Barr, East River, Colorado USA

A comprehensive Quality Assurance (QA) and Quality Control (QC) statistical framework consists of three major phases: Phase 1—Preliminary raw data sets exploration, including time formatting and combining datasets of different lengths and different time intervals; Phase 2—QA of the datasets, including detecting and flagging of duplicates, outliers, and extreme values; and Phase 3—the development of time series of a desired frequency, imputation of missing values, visualization and a final statistical summary. The time series data collected at the Billy Barr meteorological station (East River Watershed, Colorado) were analyzed. The developed statistical framework is suitable for both real-time and post-data-collection QA/QC analysis of meteorological datasets.The files that are in this data package include one excel file, converted to CSV format (Billy_Barr_raw_qaqc.csv) that contains the raw meteorological data, i.e., input data used for the QA/QC analysis. The second CSV file (Billy_Barr_1hr.csv) is the QA/QC and flagged meteorological data, i.e., output data from the QA/QC analysis. The last file (QAQC_Billy_Barr_2021-03-22.R) is a script written in R that implements the QA/QC and flagging process. The purpose of the CSV data files included in this package is to provide input and output files implemented in the R script.

54 ENVIRONMENTAL SCIENCES↗

OmicsMLMentor: A Web Application for Guided Machine Learning Analysis of Omics Data

Expression-based omics technologies (e.g. proteomics, metabolomics, transcriptomics, etc.) increasingly rely on supervised and unsupervised machine learning (ML) models to find key biomolecules distinguishing conditions, identify natural groupings in biological data, or generate predictions for outcomes of interest. Fitting ML models to omics data presents several challenges, including handling missing data, selecting a normalization method, choosing a valid model, and optimizing hyperparameters, all requiring statistical programming skills to address these challenges. Thus, the open-source web application SLOPE was designed to lower the barrier to ML modeling for omics data. SLOPE supports the fitting of 15 ML models (10 supervised and 5 unsupervised) tailored to omics datasets, such as proteomics, metabolomics, lipidomics, and transcriptomics. SLOPE offers several omics-specific features, including methods for handling missingness (imputation, conversion, removal), normalization tests, ranking of models based on the structure of a user’s data and user input, and optimal hyperparameter selections using cross-validation splits. By streamlining ML workflows for omics analysis, SLOPE address critical gaps in existing online web tools, facilitating a broader adoption of these models for omics research. Here, SLOPE is applied to data from a lignin exposure study to highlight the workflow for fitting both supervised and unsupervised models to data.

lipidomics↗

A Random Forest Approach to Identifying Young Stellar Object Candidates in the Lupus Star-forming Region

The identification and characterization of stellar members within a star-forming region are critical to many aspects of star formation, including formalization of the initial mass function, circumstellar disk evolution, and star formation history. Previous surveys of the Lupus star-forming region have identified members through infrared excess and accretion signatures. We use machine learning to identify new candidate members of Lupus based on surveys from two space-based observatories: ESA’s Gaia and NASA’s Spitzer. Astrometric measurements from Gaia's Data Release 2 and astrometric and photometric data from the Infrared Array Camera on the Spitzer Space Telescope, as well as from other surveys, are compiled into a catalog for the random forest (RF) classifier. The RF classifiers are tested to find the best features, membership list, non-membership identification scheme, imputation method, training set class weighting, and method of dealing with class imbalance within the data. We list 27 candidate members of the Lupus star-forming region for spectroscopic follow-up. Most of the candidates lie in Clouds V and VI, where only one confirmed member of Lupus was previously known. These clouds likely represent a slightly older population of star formation.

79 ASTRONOMY AND ASTROPHYSICS↗

Conditional distribution estimation of building characteristics with diffusion models for urban energy modeling

Understanding current energy consumption behavior in communities is critical for informing future energy use decisions and enabling efficient energy management. Urban energy models, which are used to simulate these energy use patterns, require large datasets with detailed building characteristics for accurate outcomes. However, such detailed characteristics at the individual building level are often unknown and costly to acquire, or unavailable. Through this work, we propose using a generative modeling approach to generate realistic building attributes to fill in the data gaps and finally provide complete characteristics as inputs to energy models. Our model learns complex, building-level patterns from training on a large-scale residential building stock model containing 2.2 million buildings. We employ a tabular diffusion-based framework that is designed to handle heterogeneous (discrete and continuous) features in tabular building data, such as occupancy, floor area, heating, cooling, and other equipment details. We develop a capability for conditional diffusion, enabling the imputation of missing building characteristics conditioned on known attributes. We conduct a comprehensive validation of our conditional diffusion model, firstly by comparing the generated conditional distributions against the underlying data distribution, and secondly, by performing a case study for a Baltimore residential region, showing the practical utility of our approach. Our work is one of the first to demonstrate the potential of generative modeling to accelerate building energy modeling workflows.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Predicting Large‐Scale Systematic Missing Pipe Attributes in Water Distribution Networks

Water distribution network (WDN) models are an essential tool used by water utilities for hydraulic analysis. Unfortunately, missing data and insufficient resources often make creating and maintaining these models unfeasible. Existing methods to address missing pipe properties, like sequential imputation for missing values and reconstruction using graph metrics, are designed to accommodate random patterns of missing information and require a significant percentage of the system's attributes to be known. However, these data completeness assumptions do not always align with real‐world scenarios where large sections of the WDN model have missing data. To address this challenge, this study proposes a data‐driven approach for estimating pipe diameter when considering different spatial patterns and degrees of data completeness (i.e., 0%–90%). Using data from 16 WDNs in Kentucky, this study compares the use of machine learning (ML) using topological and geospatial features against an existing deterministic approach. Results demonstrate that WDN models with pipe diameters predicted by the proposed ML method had comparable hydraulic performance to the ground truth models. Moreover, results showed that ML method performance varies between WDNs of differing topological classification. Insights from this study help advance the ability to leverage partial data to create and maintain WDN models amid uncertainty and inadequate resources.

Poff, Jason W. [Oregon State Univ., Corvallis, OR ↗

A consistent dataset for the net income distribution for 190 countries and aggregated to 32 geographical regions from 1958 to 2015

Abstract. Data on income distributions within and across countries are becoming increasingly important for informing analysis of income inequality and understanding the distributional consequences of climate change. While datasets on income distribution collected from household surveys are available for multiple countries, these datasets often do not represent the same concept of inequality (or income concept) and therefore make comparisons across countries, over time and across datasets difficult. Here, we present a consistent dataset of income distributions across 190 countries from 1958 to 2015 measured in terms of net income. We complement the observed values in this dataset with values imputed from a summary measure of the income distribution, specifically the Gini coefficient. For the imputation, we use a recently developed nonparametric principal-component-based approach that shows an excellent fit to data on income distributions compared to other approaches. We also present another version of this dataset aggregated from the country level to 32 geographical regions. Our dataset is developed for the purpose of calibrating models such as integrated human–Earth system models with detailed data on income distributions. This dataset will enable more robust analysis of income distribution at multiple scales. The latest version of our data are available on Zenodo: https://doi.org/10.5281/zenodo.7093997 (Narayan et al., 2022b).

97 MATHEMATICS AND COMPUTING↗

MultiPEM Toolbox: User Manual [Rev. 2]

This document explains use of the Multi-Phenomenology Explosion Monitoring (Multi PEM) Toolbox, a collection of R scripts for estimating the unknown device parameters of a new event with uncertainty quantification. The methodology and application used for illustration in this user manual are fully documented in a Los Alamos National Laboratory technical report hereafter designated “WPA” for reference. Additional details on the application are found in a recent journal article. Two assessment types are available: rapid and complete. Rapid assessments are conducted in two stages, as described in Section 2. In the first stage, calibration data are used to estimate forward and error model parameters (WPA, §5.1) and (if relevant) errors-in-variables yield values for calibration sources (WPA, §3, Equation (3)). In the second stage, new event data are used to estimate the unknown new event device parameters (WPA, §5.2) with uncertainty quantification. Two options for treating the inferred first stage parameters in second stage Bayesian analysis are available: fixing them at their maximum likelihood estimate (default), or multiple imputation. Multiple imputation involves utilizing several posterior samples (imputations) of the first stage parameters as fixed values in the second stage posterior sampling of the new event device parameters. Second stage sampling is conducted across imputations in parallel to improve computational efficiency. This method produces improved uncertainty quantification of the new event device parameters compared with the default treatment of the first stage parameters, at the expense of additional computation. Complete assessments are conducted in a single stage, as described in Section 3. Calibration and (if relevant) new event data are used simultaneously to estimate all forward model, error model, and (if relevant) new event device parameters with uncertainty quantification on the latter. As the name suggests, rapid assessments generally run substantially faster than complete assessments (even with multiple imputation), because the results of first stage analysis can be stored and incorporated into estimating a relatively low-dimensional space of new event device parameters whenever relevant new event data becomes available. On the other hand, complete assessments must be run on the full set of model and device parameters with calibration and new event data every time the latter becomes available.

97 MATHEMATICS AND COMPUTING↗

A graph neural network (GNN) approach to basin-scale river network learning: the role of physics-based connectivity and data fusion

Abstract. Rivers and river habitats around the world are under sustained pressure from human activities and the changing global environment. Our ability to quantify and manage the river states in a timely manner is critical for protecting the public safety and natural resources. In recent years, vector-based river network models have enabled modeling of large river basins at increasingly fine resolutions, but are computationally demanding. This work presents a multistage, physics-guided, graph neural network (GNN) approach for basin-scale river network learning and streamflow forecasting. During training, we train a GNN model to approximate outputs of a high-resolution vector-based river network model; we then fine-tune the pretrained GNN model with streamflow observations. We further apply a graph-based, data-fusion step to correct prediction biases. The GNN-based framework is first demonstrated over a snow-dominated watershed in the western United States. A series of experiments are performed to test different training and imputation strategies. Results show that the trained GNN model can effectively serve as a surrogate of the process-based model with high accuracy, with median Kling–Gupta efficiency (KGE) greater than 0.97. Application of the graph-based data fusion further reduces mismatch between the GNN model and observations, with as much as 50 % KGE improvement over some cross-validation gages. To improve scalability, a graph-coarsening procedure is introduced and is demonstrated over a much larger basin. Results show that graph coarsening achieves comparable prediction skills at only a fraction of training cost, thus providing important insights into the degree of physical realism needed for developing large-scale GNN-based river network models.

54 ENVIRONMENTAL SCIENCES↗

Jobs, jobs, jobs: what’s an analyst to do?

Analysts and economists often face the task of using employment metrics to characterize industries of interest. Some key challenges can be understanding where to find employment metrics, the differences in various employment metrics, and when each metric should be used. This article analyzes a variety of publicly available employment data for the United States and compares these data. A detailed description of the intricacies of each data source is provided, which covers factors such as regionality, industry breakout, periodicity, and the types of jobs included. This article provides several case study examples, using the oil and gas extraction, coal mining, and chemical manufacturing sectors to portray challenges data users may face when developing employment estimates that suit their needs. Data users should be aware of a variety of data sources to understand alternative analysis options when data limitations are present and to determine which data source best meets their needs. Instances may occur in which information from one dataset may be used to help impute missing values.

99 GENERAL AND MISCELLANEOUS↗

Constructing a Simulation Surrogate with Partially Observed Output

Gaussian process surrogates are a popular alternative to directly using computationally expensive simulation models. When the simulation output consists of many responses, dimension-reduction techniques are often employed to construct these surrogates. However, surrogate methods with dimension reduction generally rely on complete output training data. This article proposes a new Gaussian process surrogate method that permits the use of partially observed output while remaining computationally efficient. The new method involves the imputation of missing values and the adjustment of the covariance matrix used for Gaussian process inference. The resulting surrogate represents the available responses, disregards the missing responses, and provides meaningful uncertainty quantification. In conclusion, the proposed approach is shown to offer sharper inference than alternatives in a simulation study and a case study where an energy density functional model that frequently returns incomplete output is calibrated.

42 ENGINEERING↗

Predictive analytics of selections of russet potatoes

We explore the application of machine learning algorithms specifically to enhance the selection process of Russet potato (Solanum tuberosum L.) clones in breeding trials by predicting their suitability for advancement. This study addresses the challenge of efficiently identifying high-yield, disease-resistant, and climate-resilient potato varieties that meet processing industry standards. Leveraging manually collected data from trials in the state of Oregon, we investigate the potential of a wide variety of state-of-the-art binary classification models. The dataset includes 1086 clones, with data on 38 attributes recorded for each clone, focusing on yield, size, appearance, and frying characteristics, with several control varieties planted consistently across four Oregon regions from 2013 to 2021. We conduct a comprehensive analysis of the dataset that includes preprocessing, feature engineering, and imputation to address missing values. We focus on several key metrics such as accuracy, F1-score, and Matthews correlation coefficient (MCC) for model evaluation. The top-performing models, namely a feedforward neural network classifier (Neural Net), a histogram-based gradient boosting classifier (HGBC), and a support vector machine classifier (SVM), demonstrate consistent and significant results. To further validate our findings, we conducted a simulation study using the aims, data-generating mechanisms, estimands, methods, and performance measures (ADEMP) framework, simulating different data-generating scenarios to assess model robustness and performance through true positive, true negative, false positive, and false negative distributions, area under the receiver operating characteristic curve (AUC-ROC) and MCC. The simulation results highlight that non-linear models like SVM and HGBC consistently show higher AUC-ROC and MCC than logistic regression, thus outperforming the traditional linear model across various distributions, and emphasizing the importance of model selection and tuning in agricultural trials. Variable selection further enhances model performance and identifies influential features in predicting trial outcomes. The findings emphasize the potential of machine learning in streamlining the selection process for potato varieties, offering benefits such as increased efficiency, substantial cost savings, and judicious resource utilization. Our study contributes insights into precision agriculture and showcases the relevance of advanced technologies for informed decision-making in breeding programs.

60 APPLIED LIFE SCIENCES↗

A Coupled Deep Learning Model for Estimating Surface NO 2 Levels from Remote Sensing Data: 15-Year Study Over the Contiguous United States

This study proposes a novel two-step deep learning (DL) model for estimating surface NO 2 concentrations using satellite data over the contiguous United States (CONUS) from 2005 to 2019. The first phase of the model uses partial convolutional neural network (PCNN), an advanced DL model that accurately imputes gaps between surface NO 2 stations and creates 5,478 daily-mean NO 2 grids (PCNN-NO 2 ) of the 2005-2019 period over the study area. We then feed the PCNN-NO 2 , along with other predictor variables, into a deep neural network (DNN) to estimate surface NO 2 levels, achieving exceptional performance with a correlation coefficient of 0.975 to 0.978, a mean absolute bias of 0.99 ppb to 1.38 ppb, and a root mean square error of 1.47 ppb to 1.97 ppb. Spatial cross-validation results also indicate strong spatial performance of PCNN-DNN surface NO 2 estimates. In addition to its accurate estimates, the PCNN-DNN model consistently generates estimated NO 2 grids without any missing values, improving the quality of various applications such as emission reduction strategies and public health studies. Between 2005 and 2019, the 5,478 daily estimated NO 2 grids over the CONUS reveal significant reductions in NO 2 levels in fourteen major urban environments: Washington D.C. (-43%), New York (-45%), Los Angeles (-38%), Chicago (-25%), Boston (-43%), Houston (-34%), Dallas (-40%), Philadelphia (-41%), Phoenix (-38%), Detroit (-20%), Denver (-23%), Atlanta (-0.7%), Cincinnati (-38%), and Pittsburgh (-56%). Furthermore, the study shows that the denser urban regions that in-situ stations are installed in, the higher the difference between in-situ observations and regional-mean NO 2 levels.

54 ENVIRONMENTAL SCIENCES↗

Bayesian Framework for Multi-Timescale State Estimation in Low-Observable Distribution Systems

To support the smart grid paradigm, there has been a significant increase in sensor deployments and metering infrastructure in distribution systems. However, the measurements provided by these sensors and metering devices are typically sampled at different rates and could suffer from losses during the aggregation process. It is crucial to effectively reconcile the time-series measurements for a reliable state estimation. While weighted least squares has been the traditional approach for state estimation, sparsity-based approaches like matrix completion have become popular due to their superior performance in low-observability conditions. This paper proposes a Bayesian framework for both multi-timescale data aggregation and matrix completion based state estimation. Specifically, the multiscale time-series data aggregated from heterogenous sources are reconciled using a multitask Gaussian process that exploits the spatio-temporal correlations. Here, the resulting consistent timeseries alongwith the confidence bound on the imputations are fed into a Bayesian matrix completion method augmented with linearized power-flow constraints to accurately estimate the states in low-observability conditions. Results on three phase unbalanced IEEE 37 and IEEE 123 bus test systems reveal the superior performance of the proposed Bayesian framework. The computational complexity for the proposed Bayesian framework is also quantified.

42 ENGINEERING↗

Integrating Intermediate Traits in Phylogenetic Genotype-to-Phenotype Studies

A major goal of research in evolution and genetics is linking genotype to phenotype. This work could be direct, such as determining the genetic basis of a phenotype by leveraging genetic variation or divergence in a developmental, physiological, or behavioral trait. The work could also involve studying the evolutionary phenomena (e.g., reproductive isolation, adaptation, sexual dimorphism, behavior) that reveal an indirect link between genotype and a trait of interest. When the phenotype diverges across evolutionarily distinct lineages, this genotype-to-phenotype problem can be addressed using phylogenetic genotype-to-phenotype (PhyloG2P) mapping, which uses genetic signatures and convergent phenotypes on a phylogeny to infer the genetic bases of traits. The PhyloG2P approach has proven powerful in revealing key genetic changes associated with diverse traits, including the mammalian transition to marine environments and transitions between major mechanisms of photosynthesis. However, there are several intermediate traits layered in between genotype and the phenotype of interest, including but not limited to transcriptional profiles, chromatin states, protein abundances, structures, modifications, metabolites, and physiological parameters. Each intermediate trait is interesting and informative in its own right, but synthesis across data types has great promise for providing a deep, integrated, and predictive understanding of how genotypes drive phenotypic differences and convergence. We argue that an expanded PhyloG2P framework (the PhyloG2P matrix) that explicitly considers intermediate traits, and imputes those that are prohibitive to obtain, will allow a better mechanistic understanding of any trait of interest. Furthermore, this approach provides a proxy for functional validation and mechanistic understanding in organisms where laboratory manipulation is impractical.

59 BASIC BIOLOGICAL SCIENCES↗

STAIR 2.0: A Generic and Automatic Algorithm to Fuse Modis, Landsat, and Sentinel-2 to Generate 10 m, Daily, and Cloud-/Gap-Free Surface Reflectance Product

Remote sensing datasets with both high spatial and high temporal resolution are critical for monitoring and modeling the dynamics of land surfaces. However, no current satellite sensor could simultaneously achieve both high spatial resolution and high revisiting frequency. Therefore, the integration of different sources of satellite data to produce a fusion product has become a popular solution to address this challenge. Many methods have been proposed to generate synthetic images with rich spatial details and high temporal frequency by combining two types of satellite datasets—usually frequent coarse-resolution images (e.g., MODIS) and sparse fine-resolution images (e.g., Landsat). In this paper, we introduce STAIR 2.0, a new fusion method that extends the previous STAIR fusion framework, to fuse three types of satellite datasets, including MODIS, Landsat, and Sentinel-2. In STAIR 2.0, input images are first processed to impute missing-value pixels that are due to clouds or sensor mechanical issues using a gap-filling algorithm. The multiple refined time series are then integrated stepwisely, from coarse- to fine- and high-resolution, ultimately providing a synthetic daily, high-resolution surface reflectance observations. We applied STAIR 2.0 to generate a 10-m, daily, cloud-/gap-free time series that covers the 2017 growing season of Saunders County, Nebraska. Moreover, the framework is generic and can be extended to integrate more types of satellite data sources, further improving the quality of the fusion product. View Full-Text

47 OTHER INSTRUMENTATION↗

Integration of ultra-low coverage whole-genome sequences for reconstructing the evolutionary history of Galapagos giant tortoises

Genomic data from contemporary and historical samples often need to be coupled for evolutionary reconstructions of multitaxon complexes. However, the genetic data recovered from historical samples may result only in ultra-low coverage whole-genome sequences (ulcWGS; <0.15× depth), leading to inaccurate evolutionary inferences given a preponderance of missing data. Using the Galapagos giant tortoise radiation as a study system (Chelonoidis spp., composed of 13 extant and four extinct lineages), we assembled a novel methodological pipeline that removes potential noise introduced by the missing data and enhances the evolutionary signal from ulcWGS samples. We leveraged existing tools for phylogenomic placement (EPA-ng), population genomic structure (smartsnp) and admixture (Admixfrog, NGSadmix) to demonstrate that the evolutionary history of samples can be uncovered with sequencing depths as low as 0.008–0.139×. Importantly, these approaches do not use genotype imputation of the ulcWGS samples, which would require extensive reference datasets. Our application to two cases of extinct lineages of Galapagos giant tortoises, with and without references from the same lineage, demonstrates the general value of the approach. We confirm where the extinct lineages from San Cristóbal and Santa Fe islands fit into the Galapagos giant tortoise radiation, and that these lineages were evolutionarily distinct entities.

ancient DNA↗