Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data imputation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Fission Product Gas Monitoring During Fuel Drying Operations - 20094

The goal of the project was to devise and implement a method of accurately qualifying any noble gas emitted during fuel drying operations (part of the dry fuel storage process) in order to provide the data needed to validate the off-site dose calculations. The imputes of the project was the NRC publishing Information Notice 18-01, 'Noble Fission Gas Releases During Spent Fuel Cask Loading Operations' in February 2018. The information notice describes some industry events that occurred during vacuum drying operations as part of a dry fuel campaign. The solution was devised by Fort Calhoun Station staff in partnership with Mirion Technologies technical staff. The monitoring was accomplished by directing all the effluent from the fuel storage cask vacuum drying machine through a sample chamber containing a compact ISOCS characterized CZT based gamma spectrometer connected to a Mirion Data Analyst. This resulted near real-time quantification of the radionuclides of concern. (authors)

07 ISOTOPE AND RADIATION SOURCES↗

Quality Assurance and Quality Control (QA/QC) of Meteorological Time Series Data for Billy Barr, East River, Colorado USA

A comprehensive Quality Assurance (QA) and Quality Control (QC) statistical framework consists of three major phases: Phase 1—Preliminary raw data sets exploration, including time formatting and combining datasets of different lengths and different time intervals; Phase 2—QA of the datasets, including detecting and flagging of duplicates, outliers, and extreme values; and Phase 3—the development of time series of a desired frequency, imputation of missing values, visualization and a final statistical summary. The time series data collected at the Billy Barr meteorological station (East River Watershed, Colorado) were analyzed. The developed statistical framework is suitable for both real-time and post-data-collection QA/QC analysis of meteorological datasets.The files that are in this data package include one excel file, converted to CSV format (Billy_Barr_raw_qaqc.csv) that contains the raw meteorological data, i.e., input data used for the QA/QC analysis. The second CSV file (Billy_Barr_1hr.csv) is the QA/QC and flagged meteorological data, i.e., output data from the QA/QC analysis. The last file (QAQC_Billy_Barr_2021-03-22.R) is a script written in R that implements the QA/QC and flagging process. The purpose of the CSV data files included in this package is to provide input and output files implemented in the R script.

54 ENVIRONMENTAL SCIENCES↗

OmicsMLMentor: A Web Application for Guided Machine Learning Analysis of Omics Data

Expression-based omics technologies (e.g. proteomics, metabolomics, transcriptomics, etc.) increasingly rely on supervised and unsupervised machine learning (ML) models to find key biomolecules distinguishing conditions, identify natural groupings in biological data, or generate predictions for outcomes of interest. Fitting ML models to omics data presents several challenges, including handling missing data, selecting a normalization method, choosing a valid model, and optimizing hyperparameters, all requiring statistical programming skills to address these challenges. Thus, the open-source web application SLOPE was designed to lower the barrier to ML modeling for omics data. SLOPE supports the fitting of 15 ML models (10 supervised and 5 unsupervised) tailored to omics datasets, such as proteomics, metabolomics, lipidomics, and transcriptomics. SLOPE offers several omics-specific features, including methods for handling missingness (imputation, conversion, removal), normalization tests, ranking of models based on the structure of a user’s data and user input, and optimal hyperparameter selections using cross-validation splits. By streamlining ML workflows for omics analysis, SLOPE address critical gaps in existing online web tools, facilitating a broader adoption of these models for omics research. Here, SLOPE is applied to data from a lignin exposure study to highlight the workflow for fitting both supervised and unsupervised models to data.

lipidomics↗

A Random Forest Approach to Identifying Young Stellar Object Candidates in the Lupus Star-forming Region

The identification and characterization of stellar members within a star-forming region are critical to many aspects of star formation, including formalization of the initial mass function, circumstellar disk evolution, and star formation history. Previous surveys of the Lupus star-forming region have identified members through infrared excess and accretion signatures. We use machine learning to identify new candidate members of Lupus based on surveys from two space-based observatories: ESA’s Gaia and NASA’s Spitzer. Astrometric measurements from Gaia's Data Release 2 and astrometric and photometric data from the Infrared Array Camera on the Spitzer Space Telescope, as well as from other surveys, are compiled into a catalog for the random forest (RF) classifier. The RF classifiers are tested to find the best features, membership list, non-membership identification scheme, imputation method, training set class weighting, and method of dealing with class imbalance within the data. We list 27 candidate members of the Lupus star-forming region for spectroscopic follow-up. Most of the candidates lie in Clouds V and VI, where only one confirmed member of Lupus was previously known. These clouds likely represent a slightly older population of star formation.

79 ASTRONOMY AND ASTROPHYSICS↗

Exploration Analysis of Carbon Dioxide Levels and Ultrasound Measures of the Eye During ISS Missions

Enhanced screening for the Visual Impairment/Intracranial Pressure (VIIP) Syndrome, including in-flight ultrasound, was implemented in 2010 to better characterize the changes in vision observed in some long-duration crewmembers. Suggested possible risk factors for VIIP include cardiovascular changes, diet, anatomical and genetic factors, and environmental conditions. As a potent vasodilator, carbon dioxide (CO (sub 2)), which is chronically elevated on the International Space Station (ISS) relative to typical indoor and outdoor ambient levels on Earth, seems a plausible contributor to VIIP. In an effort to understand the possible associations between CO (sub 2) and VIIP, this study analyzes the relationship between ambient CO (sub 2) levels on ISS and ultrasound measures of the eye obtained from ISS fliers. CO (sub 2) measurements will be pulled directly from Operational Data Reduction Complex for the Lab and Node 3 major constituent analyzers (MCAs) on ISS or from sensors located in the European Columbus module, as available. CO (sub 2) measures between ultrasound sessions will be summarized using standard time series class metrics in MATLAB including time-weighted means and variances. Cumulative CO (sub 2) exposure metrics will also be developed. Regression analyses will be used to quantify the relationships between the CO (sub 2) metrics and specific ultrasound measures. Generalized estimating equations will adjust for the repeated measures within individuals. Multiple imputation techniques will be used to adjust for any possible biases in missing data for either CO (sub 2) or ultrasound measures. These analyses will elucidate the possible relationship between CO (sub 2) and changes in vision and also inform future analysis of inflight VIIP data.

Young, M.↗

Conditional distribution estimation of building characteristics with diffusion models for urban energy modeling

Understanding current energy consumption behavior in communities is critical for informing future energy use decisions and enabling efficient energy management. Urban energy models, which are used to simulate these energy use patterns, require large datasets with detailed building characteristics for accurate outcomes. However, such detailed characteristics at the individual building level are often unknown and costly to acquire, or unavailable. Through this work, we propose using a generative modeling approach to generate realistic building attributes to fill in the data gaps and finally provide complete characteristics as inputs to energy models. Our model learns complex, building-level patterns from training on a large-scale residential building stock model containing 2.2 million buildings. We employ a tabular diffusion-based framework that is designed to handle heterogeneous (discrete and continuous) features in tabular building data, such as occupancy, floor area, heating, cooling, and other equipment details. We develop a capability for conditional diffusion, enabling the imputation of missing building characteristics conditioned on known attributes. We conduct a comprehensive validation of our conditional diffusion model, firstly by comparing the generated conditional distributions against the underlying data distribution, and secondly, by performing a case study for a Baltimore residential region, showing the practical utility of our approach. Our work is one of the first to demonstrate the potential of generative modeling to accelerate building energy modeling workflows.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Predicting Large‐Scale Systematic Missing Pipe Attributes in Water Distribution Networks

Water distribution network (WDN) models are an essential tool used by water utilities for hydraulic analysis. Unfortunately, missing data and insufficient resources often make creating and maintaining these models unfeasible. Existing methods to address missing pipe properties, like sequential imputation for missing values and reconstruction using graph metrics, are designed to accommodate random patterns of missing information and require a significant percentage of the system's attributes to be known. However, these data completeness assumptions do not always align with real‐world scenarios where large sections of the WDN model have missing data. To address this challenge, this study proposes a data‐driven approach for estimating pipe diameter when considering different spatial patterns and degrees of data completeness (i.e., 0%–90%). Using data from 16 WDNs in Kentucky, this study compares the use of machine learning (ML) using topological and geospatial features against an existing deterministic approach. Results demonstrate that WDN models with pipe diameters predicted by the proposed ML method had comparable hydraulic performance to the ground truth models. Moreover, results showed that ML method performance varies between WDNs of differing topological classification. Insights from this study help advance the ability to leverage partial data to create and maintain WDN models amid uncertainty and inadequate resources.

Poff, Jason W. [Oregon State Univ., Corvallis, OR ↗

A consistent dataset for the net income distribution for 190 countries and aggregated to 32 geographical regions from 1958 to 2015

Abstract. Data on income distributions within and across countries are becoming increasingly important for informing analysis of income inequality and understanding the distributional consequences of climate change. While datasets on income distribution collected from household surveys are available for multiple countries, these datasets often do not represent the same concept of inequality (or income concept) and therefore make comparisons across countries, over time and across datasets difficult. Here, we present a consistent dataset of income distributions across 190 countries from 1958 to 2015 measured in terms of net income. We complement the observed values in this dataset with values imputed from a summary measure of the income distribution, specifically the Gini coefficient. For the imputation, we use a recently developed nonparametric principal-component-based approach that shows an excellent fit to data on income distributions compared to other approaches. We also present another version of this dataset aggregated from the country level to 32 geographical regions. Our dataset is developed for the purpose of calibrating models such as integrated human–Earth system models with detailed data on income distributions. This dataset will enable more robust analysis of income distribution at multiple scales. The latest version of our data are available on Zenodo: https://doi.org/10.5281/zenodo.7093997 (Narayan et al., 2022b).

97 MATHEMATICS AND COMPUTING↗

Introductory comments on the USGS geographic applications program

The third phase of remote sensing technologies and potentials applied to the operations of the U.S. Geological Survey is introduced. Remote sensing data with multidisciplinary spatial data from traditional sources is combined with geographic theory and techniques of environmental modeling. These combined imputs are subject to four sequential activities that involve: (1) thermatic mapping of land use and environmental factors; (2) the dynamics of change detection; (3) environmental surveillance to identify sudden changes and general trends; and (4) preparation of statistical model and analytical reports. Geography program functions, products, clients, and goals are presented in graphical form, along with aircraft photo missions, geography test sites, and FY-70.

Gerlach, A. C.↗

Antenna arraying of Voyager telemetry signals by symbol stream combining

Telemetry signals received from the Voyager 2 spacecraft at Deep Space Stations at Parkes and Canberra, Australia, on February 6, 1986, were combined by the method of symbol stream combining. This second demonstration of symbol stream combining followed the International Cometary Explorer (ICE) demonstration at Giacobini-Zinner encounter in September 1985. The Voyager demonstration was at a symbol rate of 43.2 ksymb/s, compared to 2 ksymb/s for ICE. Recording, playback, and combining at this higher rate were demonstrated. The average symbol signal-to-noise ratio (SNR) of the combined data was 2.84 dB, or 0.23 dB less than the sum of the SNRs of the two imput symbol streams. This 0.23 loss from ideal combining was due to use of 4-bit quantization of the input symbol stream and imperfect scaling. A practical implementation with 8-bit quantization could achieve combining losses of under 0.05 dB over a wide dynamic range of input signal levels.

Hurd, W. J.↗

MultiPEM Toolbox: User Manual [Rev. 2]

This document explains use of the Multi-Phenomenology Explosion Monitoring (Multi PEM) Toolbox, a collection of R scripts for estimating the unknown device parameters of a new event with uncertainty quantification. The methodology and application used for illustration in this user manual are fully documented in a Los Alamos National Laboratory technical report hereafter designated “WPA” for reference. Additional details on the application are found in a recent journal article. Two assessment types are available: rapid and complete. Rapid assessments are conducted in two stages, as described in Section 2. In the first stage, calibration data are used to estimate forward and error model parameters (WPA, §5.1) and (if relevant) errors-in-variables yield values for calibration sources (WPA, §3, Equation (3)). In the second stage, new event data are used to estimate the unknown new event device parameters (WPA, §5.2) with uncertainty quantification. Two options for treating the inferred first stage parameters in second stage Bayesian analysis are available: fixing them at their maximum likelihood estimate (default), or multiple imputation. Multiple imputation involves utilizing several posterior samples (imputations) of the first stage parameters as fixed values in the second stage posterior sampling of the new event device parameters. Second stage sampling is conducted across imputations in parallel to improve computational efficiency. This method produces improved uncertainty quantification of the new event device parameters compared with the default treatment of the first stage parameters, at the expense of additional computation. Complete assessments are conducted in a single stage, as described in Section 3. Calibration and (if relevant) new event data are used simultaneously to estimate all forward model, error model, and (if relevant) new event device parameters with uncertainty quantification on the latter. As the name suggests, rapid assessments generally run substantially faster than complete assessments (even with multiple imputation), because the results of first stage analysis can be stored and incorporated into estimating a relatively low-dimensional space of new event device parameters whenever relevant new event data becomes available. On the other hand, complete assessments must be run on the full set of model and device parameters with calibration and new event data every time the latter becomes available.

97 MATHEMATICS AND COMPUTING↗

Development of a takeoff performance monitoring system

Discussed are the development and testing of a real-time takeoff performance monitoring algorithm. The algorithm is made up of two segments: a pretakeoff segment and a real-time segment. One-time imputs of ambient conditions and airplane configuration information are used in the pretakeoff segment to generate scheduled performance data for that takeoff. The real-time segment uses the scheduled performance data generated in the pretakeoff segment, runway length data, and measured parameters to monitor the performance of the airplane throughout the takeoff roll. Airplane and engine performance deficiencies are detected and annunciated. An important feature of this algorithm is the one-time estimation of the runway rolling friction coefficient. The algorithm was tested using a six-degree-of-freedom airplane model in a computer simulation. Results from a series of sensitivity analyses are also included.

Srivatsan, Raghavachari↗

A graph neural network (GNN) approach to basin-scale river network learning: the role of physics-based connectivity and data fusion

Abstract. Rivers and river habitats around the world are under sustained pressure from human activities and the changing global environment. Our ability to quantify and manage the river states in a timely manner is critical for protecting the public safety and natural resources. In recent years, vector-based river network models have enabled modeling of large river basins at increasingly fine resolutions, but are computationally demanding. This work presents a multistage, physics-guided, graph neural network (GNN) approach for basin-scale river network learning and streamflow forecasting. During training, we train a GNN model to approximate outputs of a high-resolution vector-based river network model; we then fine-tune the pretrained GNN model with streamflow observations. We further apply a graph-based, data-fusion step to correct prediction biases. The GNN-based framework is first demonstrated over a snow-dominated watershed in the western United States. A series of experiments are performed to test different training and imputation strategies. Results show that the trained GNN model can effectively serve as a surrogate of the process-based model with high accuracy, with median Kling–Gupta efficiency (KGE) greater than 0.97. Application of the graph-based data fusion further reduces mismatch between the GNN model and observations, with as much as 50 % KGE improvement over some cross-validation gages. To improve scalability, a graph-coarsening procedure is introduced and is demonstrated over a much larger basin. Results show that graph coarsening achieves comparable prediction skills at only a fraction of training cost, thus providing important insights into the degree of physical realism needed for developing large-scale GNN-based river network models.

54 ENVIRONMENTAL SCIENCES↗

Jobs, jobs, jobs: what’s an analyst to do?

Analysts and economists often face the task of using employment metrics to characterize industries of interest. Some key challenges can be understanding where to find employment metrics, the differences in various employment metrics, and when each metric should be used. This article analyzes a variety of publicly available employment data for the United States and compares these data. A detailed description of the intricacies of each data source is provided, which covers factors such as regionality, industry breakout, periodicity, and the types of jobs included. This article provides several case study examples, using the oil and gas extraction, coal mining, and chemical manufacturing sectors to portray challenges data users may face when developing employment estimates that suit their needs. Data users should be aware of a variety of data sources to understand alternative analysis options when data limitations are present and to determine which data source best meets their needs. Instances may occur in which information from one dataset may be used to help impute missing values.

99 GENERAL AND MISCELLANEOUS↗

Constructing a Simulation Surrogate with Partially Observed Output

Gaussian process surrogates are a popular alternative to directly using computationally expensive simulation models. When the simulation output consists of many responses, dimension-reduction techniques are often employed to construct these surrogates. However, surrogate methods with dimension reduction generally rely on complete output training data. This article proposes a new Gaussian process surrogate method that permits the use of partially observed output while remaining computationally efficient. The new method involves the imputation of missing values and the adjustment of the covariance matrix used for Gaussian process inference. The resulting surrogate represents the available responses, disregards the missing responses, and provides meaningful uncertainty quantification. In conclusion, the proposed approach is shown to offer sharper inference than alternatives in a simulation study and a case study where an energy density functional model that frequently returns incomplete output is calibrated.

42 ENGINEERING↗

An analysis of the human as a predictor model

A hybrid sampled data-continuous impulse response of a preview tracker using a fast time model predictor aid is derived. The model can accept transient and nontransient imputs and is suitable for studying a human preview tracker using an internalized or externalized predictor display. The model is shown to behave reasonably in a simple example and its implications for modeling studies are discussed.

Kreifeldt, J. G.↗

Analysis of reactively loaded microstrip disk antenna

The moment method solution to the problem of a reactively loaded circular patch is presented. Using the reaction integral equation in conjuction with the method of moments, parameters of the Thevenin's equivalent network for the loaded patch are obtained. From the equivalent network parameters an expression for the imput impedance of the loaded patch is derived. A design procedure for a circularly polarized disk antenna is presented. Computed results are compared with the experimental data.

Deshpande, M. D.↗

Predictive analytics of selections of russet potatoes

We explore the application of machine learning algorithms specifically to enhance the selection process of Russet potato (Solanum tuberosum L.) clones in breeding trials by predicting their suitability for advancement. This study addresses the challenge of efficiently identifying high-yield, disease-resistant, and climate-resilient potato varieties that meet processing industry standards. Leveraging manually collected data from trials in the state of Oregon, we investigate the potential of a wide variety of state-of-the-art binary classification models. The dataset includes 1086 clones, with data on 38 attributes recorded for each clone, focusing on yield, size, appearance, and frying characteristics, with several control varieties planted consistently across four Oregon regions from 2013 to 2021. We conduct a comprehensive analysis of the dataset that includes preprocessing, feature engineering, and imputation to address missing values. We focus on several key metrics such as accuracy, F1-score, and Matthews correlation coefficient (MCC) for model evaluation. The top-performing models, namely a feedforward neural network classifier (Neural Net), a histogram-based gradient boosting classifier (HGBC), and a support vector machine classifier (SVM), demonstrate consistent and significant results. To further validate our findings, we conducted a simulation study using the aims, data-generating mechanisms, estimands, methods, and performance measures (ADEMP) framework, simulating different data-generating scenarios to assess model robustness and performance through true positive, true negative, false positive, and false negative distributions, area under the receiver operating characteristic curve (AUC-ROC) and MCC. The simulation results highlight that non-linear models like SVM and HGBC consistently show higher AUC-ROC and MCC than logistic regression, thus outperforming the traditional linear model across various distributions, and emphasizing the importance of model selection and tuning in agricultural trials. Variable selection further enhances model performance and identifies influential features in predicting trial outcomes. The findings emphasize the potential of machine learning in streamlining the selection process for potato varieties, offering benefits such as increased efficiency, substantial cost savings, and judicious resource utilization. Our study contributes insights into precision agriculture and showcases the relevance of advanced technologies for informed decision-making in breeding programs.

60 APPLIED LIFE SCIENCES↗