Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “imputation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

136 records · Page 8

Bayesian Rules of Thumb: Robust Uncertainty Quantification in Early Project Cost Estimation

Systems engineers often make use of cost Rules ofThumb in order to estimate cost during early phases of projectformulation. These Rules of Thumb typically take the form ofa sequence of percentages over which a total cost is allocatedacross NASA WBS elements. Rules of Thumb can then be usedto extrapolate cost from one or more known WBS elements tothe remaining unknown WBS elements, assisting early projectformulation architecture studies (such as those in JPL’s Team Xand A Team).A number of issues can arise when generating and using costRules of Thumb. For example, many records of project costsconsist of incomplete data. Typical methods of dealing withincomplete cost allocation data include (a) ignoring missionswith incomplete data, or (b) taking averages of the non-zero percentagesacross missions, but both of these methods can result inbiased estimates if the existence of incomplete data correlateswith total mission cost or any particular WBS element. Anothercommon example is cost reported in one or more incorrect WBSelements. This is especially prevalent in smaller missions whereit is more common for engineers to perform tasks that fall underthe purview of multiple WBS elements.Furthermore, a Rule of Thumb estimate is typically reported asa point estimate; there is no reported uncertainty around thepercentages used to generate an allocation. Even in the rarecase in which confidence intervals around mean percentages areprovided, there may be positive or negative correlations betweenWBS elements which can skew estimates.Here we attempt to address these problems by formulatingprobabilistic Rules of Thumb in which a distribution of allocationschemes, rather than a single allocation scheme, is generated.We use a bootstrap imputation method to simultaneouslyaccount for uncertainty in the missing data while using allavailable information contained in the dataset. The imputeddatasets are then input into a multivariate Bayesian modelwhich accounts for correlations between WBS elements andproperly accounts for uncertainty in the final Rule of Thumbpercentages and predictions. We describe the mathematicalmodel and provides snippets of R code utilizing the brms(Bayesian Regression Models using Stan) package. To illustratethis model, we generate a Bayesian Level 2 WBS Cost Rule ofThumb for MIDEX (Medium-Class Explorers) missions withdata extracted from NASA’s CADRe. We then compare thismethod’s performance with the classical Rule of Thumb method.

Hooke, Melissa A↗

Characterizing Spatiotemporal Uncertainty in Interpolated Meteorological Data

Interpolated meteorological data invariably contain errors. These errors have structure in time and space, particularly autocorrelation, which can cause the effects of errors to compound when model outputs are aggregated temporally or spatially. One way to account for this uncertainty is with a probabilistic model from which samples can be drawn that are coherent with respect to underlying spatial and temporal covariance structure. This work describes a probabilistic method for spatial interpolation of point-wise meteorological time series. Observational data from weather stations are generally sparse in space and dense in time (but sometimes missing). The method works by projecting time series onto orthogonal basis vectors and spatially interpolating each resulting component independently. Under suitable assumptions, and data transformations to better satisfy those assumptions, Gaussian process regression provides a complete description of the joint predictive distribution over a Gaussian random field. Spatiotemporally coherent realizations are generated as the sum of conditional (spatial) simulations of each orthogonal (temporal) component. Data-derived and generic orthogonal bases are considered. In addition to spatial interpolation, imputation of missing observational data is examined. The method is applied using near-surface air temperature over the Western United States and validated by comparing theoretical versus actual coverage of predictive distributions and analyzing the degree to which spatial and temporal covariance structure is reproduced. Computational considerations, relating to conditional simulation of random fields, are also addressed.

Conor T Doherty↗

AUTONOMIE VID

Autonomie Vehicle Information Database (VID) offers a comprehensive list of vehicle specifications since 1990. The database details more than 65,000 vehicles with hundreds of attributes. The database is the result of the development of a general automated data collection framework as well as the development of building blocks for processing, cleaning, integrating and analyzing complex data. The data has undergone several layers of outlier detections processes, machine learning based imputations methods have been used to deal with missing data problems, and new fields have been created according to the rules of feature engineering. Thanks to this streamlined data pipelines, the resulting processed aggregated data should deliver a unique level of information to the user in which the content can be efficiently maintained and updated.

Moswd, Ayman↗

Fission Product Gas Monitoring During Fuel Drying Operations - 20094

The goal of the project was to devise and implement a method of accurately qualifying any noble gas emitted during fuel drying operations (part of the dry fuel storage process) in order to provide the data needed to validate the off-site dose calculations. The imputes of the project was the NRC publishing Information Notice 18-01, 'Noble Fission Gas Releases During Spent Fuel Cask Loading Operations' in February 2018. The information notice describes some industry events that occurred during vacuum drying operations as part of a dry fuel campaign. The solution was devised by Fort Calhoun Station staff in partnership with Mirion Technologies technical staff. The monitoring was accomplished by directing all the effluent from the fuel storage cask vacuum drying machine through a sample chamber containing a compact ISOCS characterized CZT based gamma spectrometer connected to a Mirion Data Analyst. This resulted near real-time quantification of the radionuclides of concern. (authors)

07 ISOTOPE AND RADIATION SOURCES↗

Design Choices in Anomaly Detection for Industrial Control Systems: Insights from Gas Pipeline Data

Industrial control systems (ICS) remain vulnerable to increasingly sophisticated cyberattacks, yet evaluating anomaly detection models in these environments is challenging due to temporal dependencies, missing-not-at-random patterns, and extremely imbalanced datasets. These factors make common practices—especially random data splits and naïve imputation—prone to severe temporal leakage, which can inflate reported performance and obscure real-world limitations. In this work, we systematically examine classical machine learning models, temporal deep learning architecture, and tensor-decomposition–based methods on a gas-pipeline dataset using a fully temporally separated evaluation pipeline designed to mimic realistic deployment conditions. Our findings show that proper temporal handling and MNAR-aware preprocessing significantly alter the relative performance of popular anomaly-detection methods, providing practical guidance for designing reliable, leakage-resistant ICS intrusion-detection systems.

97 MATHEMATICS AND COMPUTING↗

Reliable machine prognostic health management in the presence of missing data

Prognostics and health management enables the prediction of future degradation and remaining useful life (RUL) for in-service systems based on historical and contemporary data, showing promise for many practical applications. One major challenge for prognostics is the common occurrence of missing values in time-series data, often caused by disruptions in sensor communication or hardware/software failures. Another major concern is that the sufficient prior knowledge of critical component degradation with a clear failure threshold is often not readily available in practice. These issues can significantly hinder the application of advanced signal and data analysis methods and consequently degrade the health management performance. In this article, we propose a novel data-driven framework that is capable of providing accurate and reliable predictions of degradation and RUL. In this approach, one-hot health state indicators are appended to the historical time series so that the model learns end-of-life automatically. A modified gate recurrent unit based variational autoencoder is employed in generative adversarial networks to model the temporal irregularity of the incomplete time series. Furthermore, experiments on multivariate time-series datasets collected from real-world aeroengines verify that significant performance improvement can be achieved using the proposed model for robust long-term prognostics.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Gap-filling eddy covariance methane fluxes: Comparison of machine learning model predictions and uncertainties at FLUXNET-CH4 wetlands

Time series of methane fluxes measured by eddy-covariance require gap-filling to estimate annual emissions. Gap-filling methane fluxes is challenging because of high variability and complex responses to multiple drivers. To date, there is no widely established gap-filling standard for methane, with regards both to the best model algorithms and predictors. In this study, we address the need for standardization by synthesizing results of gap-filling methods applied at 17 wetland sites spanning boreal to tropical regions including all major wetlands classes and two rice paddies. We introduce new procedures for: 1) creating realistic artificial gap scenarios, 2) training and evaluating gap-filling models without overstating performance, and 3) predicting half-hourly methane fluxes and annual emissions with robust uncertainty estimates. We tested a conventional method (marginal distribution sampling) and four machine learning algorithms - penalized linear regression, artificial neural networks, random forests, and boosted decision trees - and four predictor sets, including temporal, meteorological, ecosystem carbon and energy flux, and soil predictors. We find that the conventional method can achieve similar median performance to the machine learning models but is worse than the best machine learning models and relatively insensitive to predictor choices. Of the machine learning models, decision tree algorithms performed the best in cross-validation experiments, even with a baseline predictor set, and artificial neural networks showed comparable performance when using all predictors. Soil temperature was frequently the most important predictor whilst water table depth was important at sites with substantial water table fluctuations, highlighting the value of data on soil conditions. Raw gap-filling uncertainties from the machine learning models were underestimated and we propose a method to calibrate uncertainties to observations. Finally, we gap-fill and provide summary evaluation metrics for all 81 sites in the FLUXNET-CH4 community dataset and publicly release the python code for model development, evaluation, and uncertainty estimation.

42 ENGINEERING↗

High-dimensional data analytics in civil engineering: A review on matrix and tensor decomposition

Recent developments in sensing and monitoring techniques have led to the generation of high-dimensional data in the field of civil engineering. High-dimensional data analytics methods have thus been developed to interpret such complex data. Among the different high-dimensional data analytics techniques, matrix and tensor decomposition methods have acquired a notable interest in the civil engineering community over the past decade. Due to their unique ability to deal with highly redundant and correlated data, these methods are establishing themselves as promising and efficient tools to analyze high-dimensional data in the civil engineering arena. In this paper, high-dimensional data is referred to as a data set in which the number of features is comparable or larger than the number of observations. This review paper aims to summarize the applications of matrix and tensor decomposition methods in civil engineering over the last decade. The survey begins with a general overview of matrix and tensor decomposition followed by highlighting their significance in the field. Afterward, various applications of these high-dimensional data analytics methods in civil engineering are presented, while the advantages offered by these methods are discussed. Lastly, challenges and potential research avenues for employing matrix and tensor decomposition and future emerging trends for their novel use are highlighted.

42 ENGINEERING↗

Identifying Light-Duty Vehicle Travel from Large-Scale Multimodal Wearable GPS Data with Novelty Detection Algorithms

Identifying travel mode within travel survey data sets, especially light-duty vehicle (LDV) travel, is foundational, though nontrivial, to travel behavior analysis and fuel consumption estimation. Current travel mode detection approaches require well-sampled and balanced data sets with ground truth travel mode labels. They are rarely applied and validated on large-scale, real-world data sets, which may not satisfy the data requirements. This paper proposes an LDV travel mode detection model as a supplement to current travel mode detection methods, for the case when the training set is highly (and/or completely) unbalanced, to the extent that classical machine-learning approaches become difficult or impossible to deploy. The proposed model uses a novelty detection technique-one-class support vector machines (OCSVMs)-and a novel exhaustive feature extraction (EFE) technique on continuous time series data (i.e., Global Positioning System [GPS] speed profiles) for single-mode trip trajectories. Training and validation of the model are conducted on a large-scale, real-world data set. The proposed method accurately identifies LDV trips from a broad set of multimodal trips by leveraging a wealth of preexisting in-vehicle GPS travel data. Additional sensitivity analysis sheds light on the optimal training size, which will benefit applications limited by highly imbalanced data. The paper also discusses performance comparison with regular machine-learning approaches, the model's robustness, and the potential to extend the proposed model to multimodal prediction.

47 OTHER INSTRUMENTATION↗

Online PMU Missing Value Replacement Via Event-Participation Decomposition

We introduce a new method for online Phasor Measurement Unit (PMU) missing value replacement. Our approach allows us to decompose PMU event responses into a non-dynamic component (denoted the participation factor) that can be inferred directly from the past and a dynamic component that can be inferred directly from all other PMUs (denoted the event strength). When missing values occur, we can use these two components, which do not rely on the missing index, to estimate the correct value. The method is extremely fast and can easily be used for online applications. Furthermore, extensive testing on real power system event data reveals that our approach achieves state-of-the-art performance in terms of Mean Absolute Percent Errors (MAPEs) for PMU data dropped during event periods. Here, the method also yields an interpretable and simplified view of events for further analysis and applications. The method relies only on PMU data and does not take outside information such as network topology.

24 POWER TRANSMISSION AND DISTRIBUTION↗