Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Missing data problem”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Inpainting radar missing data regions with deep learning

Abstract. Missing and low-quality data regions are a frequent problem for weather radars. They stem from a variety of sources: beam blockage, instrument failure, near-ground blind zones, and many others. Filling in missing data regions is often useful for estimating local atmospheric properties and the application of high-level data processing schemes without the need for preprocessing and error-handling steps – feature detection and tracking, for instance. Interpolation schemes are typically used for this task, though they tend to produce unrealistically spatially smoothed results that are not representative of the atmospheric turbulence and variability that are usually resolved by weather radars. Recently, generative adversarial networks (GANs) have achieved impressive results in the area of photo inpainting. Here, they are demonstrated as a tool for infilling radar missing data regions. These neural networks are capable of extending large-scale cloud and precipitation features that border missing data regions into the regions while hallucinating plausible small-scale variability. In other words, they can inpaint missing data with accurate large-scale features and plausible local small-scale features. This method is demonstrated on a scanning C-band and vertically pointing Ka-band radar that were deployed as part of the Cloud Aerosol and Complex Terrain Interactions (CACTI) field campaign. Three missing data scenarios are explored: infilling low-level blind zones and short outage periods for the Ka-band radar and infilling beam blockage areas for the C-band radar. Two deep-learning-based approaches are tested, a convolutional neural network (CNN) and a GAN that optimize pixel-level error or combined pixel-level error and adversarial loss respectively. Both deep-learning approaches significantly outperform traditional inpainting schemes under several pixel-level and perceptual quality metrics.

54 ENVIRONMENTAL SCIENCES↗

Tomographic methods in flow diagnostics

This report presents a viewpoint of tomography that should be well adapted to currently available optical measurement technology as well as the needs of computational and experimental fluid dynamists. The goals in mind are to record data with the fastest optical array sensors; process the data with the fastest parallel processing technology available for small computers; and generate results for both experimental and theoretical data. An in-depth example treats interferometric data as it might be recorded in an aeronautics test facility, but the results are applicable whenever fluid properties are to be measured or applied from projections of those properties. The paper discusses both computed and neural net calibration tomography. The report also contains an overview of key definitions and computational methods, key references, computational problems such as ill-posedness, artifacts, missing data, and some possible and current research topics.

Decker, Arthur J.↗

Holographic interferometry of transparent media with reflection from imbedded test objects

In applying holographic interferometry, opaque objects blocking a portion of the optical beam used to form the interferogram give rise to incomplete data for standard computer tomography algorithms. An experimental technique for circumventing the problem of data blocked by opaque objects is presented. The missing data are completed by forming an interferogram using light backscattered from the opaque object, which is assumed to be diffuse. The problem of fringe localization is considered.

Prikryl, I.↗

Effective Missing Value Imputation Methods for Building Monitoring Data

To understand behaviors of natural and man-made events, such as energy consumption of buildings, which accounts for 40% of energy uses in the US, we deploy automated monitoring devices to record periodic observations. However, such experimental and observation data often contains problems and irregularities that have to be cleaned up before analyses. Due to various conditions affecting sensor operations, the communication channels, recording steps, or the recording media, the recorded data might have missing values, errors, or anomalous values. An effective way to clean up these problems is to replace these missing values, errors and anomalous values with expected values, a process generally known as imputation. In this work, we survey commonly used missing value imputation techniques and compare their performance on a set of building monitoring data. To compare the different types of sensor measurements with widely varying characteristics, we use normalized root mean squared error (NRMSE) as the key metric for the effectiveness of the imputation methods. We additionally consider periodicity and run time when considering comparing methods. Through extensive testing, we find that for small gap sizes, up to 8 consecutive missing values, linear interpolation performs the best; for larger gaps stretching up to 48 consecutive missing values, K-nearest neighbors provides the most accurate imputations; for even larger gaps, more computational intensive methods, such as matrix factorization, achieve the smallest NRMSE. Additionally, we observe that these computationally intensive algorithms not only provide accurate imputations for large gaps, but are also more robust across all types of sensors.

Cho, B↗

Identifying recharge sources and their impacts on a North Central New Mexico shallow aquifer using unsupervised machine learning

In this article, shallow aquifers are important but highly variable resources in arid to semi-arid regions. Limited shallow aquifer volume results in high sensitivity to recharge fluctuations, which can impact the local fauna and flora, and transport of contaminants in the aquifer or vadose zone. Aquifer response to external forcing (e.g., precipitation) is usually solved by estimating aquifer parameters and running physics-based models to match known fluctuations of hydraulic head. However, this technique is time and computationally expensive. Furthermore, high aquifer complexity decreases precision in physics-based models. Alternatively supervised machine learning is used to predict aquifer dynamics. However, these techniques rely on input data and struggle to interpret aquifer response for missing sources (i.e., snowpack data). To counter these problems, we propose an unsupervised machine learning technique (NMFk) to estimate the impact of different sources on aquifer recharge. NMFk is used to understand the influence of external forcing on shallow aquifer recharge in the Pajarito Plateau (Los Alamos, NM, USA). The results show how NMFk can be used to reduce the data dimension in a complex field dataset to three recharge signals that cause fluctuations within the field data. Here, the source signals are interpreted as rainfall, snowmelt, and a delayed aquifer response to the previous two signals. These results evidence how heterogeneous aquifers delimited by canyons incised into the Pajarito Plateau respond in similar ways to the source signals identified by NMFk. Furthermore, results show the importance of the local geology where faults act as sinks, and anthropogenic disturbances can facilitate infiltration amplifying the interpreted signal.

54 ENVIRONMENTAL SCIENCES↗

An evidential approach to problem solving when a large number of knowledge systems is available

Some recent problems are no longer formulated in terms of imprecise facts, missing data or inadequate measuring devices. Instead, questions pertaining to knowledge and information itself arise and can be phrased independently of any particular area of knowledge. The problem considered in the present work is how to model a problem solver that is trying to find the answer to some query. The problem solver has access to a large number of knowledge systems that specialize in diverse features. In this context, feature means an indicator of what the possibilities for the answer are. The knowledge systems should not be accessed more than once, in order to have truly independent sources of information. Moreover, these systems are allowed to run in parallel. Since access might be expensive, it is necessary to construct a management policy for accessing these knowledge systems. To help in the access policy, some control knowledge systems are available. Control knowledge systems have knowledge about the performance parameters status of the knowledge systems. In order to carry out the double goal of estimating what units to access and to answer the given query, diverse pieces of evidence must be fused. The Dempster-Shafer Theory of Evidence is used to pool the knowledge bases.

Dekorvin, Andre↗

Data Accountability and Uncertainty Analysis for the Mars Science Laboratory

This paper presents machine learning-based approaches to automate and optimize the detection of volume loss for the downlink process of telemetry data from the Mars Curiosity Rover. The Curiosity observes volume loss and data corruption, requiring re-transmits from the rover and Ground Data System Analysts (GDSA) to monitor the data flow. To resolve this issue, we created a data pipeline to accumulate data from various data sources in the downlink process and detect where the data is missed. In this paper, we benchmarked different methodologies based on the accuracy and excitability of them to identify whether a downlink data that is received to the ground system is complete or incomplete. Our results show that machine learning methods can improve the performance of the GDSA by 55% while the user can diagnose why data is missed and provide an explanation for the data accountability problem.

Chowdhury, Ameera↗

Diagnostic tolerance for missing sensor data

For practical automated diagnostic systems to continue functioning after failure, they must not only be able to diagnose sensor failures but also be able to tolerate the absence of data from the faulty sensors. It is shown that conventional (associational) diagnostic methods will have combinatoric problems when trying to isolate faulty sensors, even if they adequately diagnose other components. Moreover, attempts to extend the operation of diagnostic capability past sensor failure will necessarily compound those difficulties. Model-based reasoning offers a structured alternative that has no special problems diagnosing faulty sensors and can operate gracefully when sensor data is missing.

Scarl, Ethan A.↗

Missing Wedge Completion via Unsupervised Learning with Coordinate Networks

Cryogenic electron tomography (cryoET) is a powerful tool in structural biology, enabling detailed 3D imaging of biological specimens at a resolution of nanometers. Despite its potential, cryoET faces challenges such as the missing wedge problem, which limits reconstruction quality due to incomplete data collection angles. Recently, supervised deep learning methods leveraging convolutional neural networks (CNNs) have considerably addressed this issue; however, their pretraining requirements render them susceptible to inaccuracies and artifacts, particularly when representative training data is scarce. To overcome these limitations, we introduce a proof-of-concept unsupervised learning approach using coordinate networks (CNs) that optimizes network weights directly against input projections. This eliminates the need for pretraining, reducing reconstruction runtime by 3–20× compared to supervised methods. Our in silico results show improved shape completion and reduction of missing wedge artifacts, assessed through several voxel-based image quality metrics in real space and a novel directional Fourier Shell Correlation (FSC) metric. Our study illuminates benefits and considerations of both supervised and unsupervised approaches, guiding the development of improved reconstruction strategies.

42 ENGINEERING↗

Optimal Codes for the Burst Erasure Channel

Deep space communications over noisy channels lead to certain packets that are not decodable. These packets leave gaps, or bursts of erasures, in the data stream. Burst erasure correcting codes overcome this problem. These are forward erasure correcting codes that allow one to recover the missing gaps of data. Much of the recent work on this topic concentrated on Low-Density Parity-Check (LDPC) codes. These are more complicated to encode and decode than Single Parity Check (SPC) codes or Reed-Solomon (RS) codes, and so far have not been able to achieve the theoretical limit for burst erasure protection. A block interleaved maximum distance separable (MDS) code (e.g., an SPC or RS code) offers near-optimal burst erasure protection, in the sense that no other scheme of equal total transmission length and code rate could improve the guaranteed correctible burst erasure length by more than one symbol. The optimality does not depend on the length of the code, i.e., a short MDS code block interleaved to a given length would perform as well as a longer MDS code interleaved to the same overall length. As a result, this approach offers lower decoding complexity with better burst erasure protection compared to other recent designs for the burst erasure channel (e.g., LDPC codes). A limitation of the design is its lack of robustness to channels that have impairments other than burst erasures (e.g., additive white Gaussian noise), making its application best suited for correcting data erasures in layers above the physical layer. The efficiency of a burst erasure code is the length of its burst erasure correction capability divided by the theoretical upper limit on this length. The inefficiency is one minus the efficiency. The illustration compares the inefficiency of interleaved RS codes to Quasi-Cyclic (QC) LDPC codes, Euclidean Geometry (EG) LDPC codes, extended Irregular Repeat Accumulate (eIRA) codes, array codes, and random LDPC codes previously proposed for burst erasure protection. As can be seen, the simple interleaved RS codes have substantially lower inefficiency over a wide range of transmission lengths.

Hamkins, Jon↗

miss-SNF: a multimodal patient similarity network integration approach to handle completely missing data sources

Abstract Motivation Precision medicine leverages patient-specific multimodal data to improve prevention, diagnosis, prognosis, and treatment of diseases. Advancing precision medicine requires the non-trivial integration of complex, heterogeneous, and potentially high-dimensional data sources, such as multi-omics and clinical data. In the literature, several approaches have been proposed to manage missing data, but are usually limited to the recovery of subsets of features for a subset of patients. A largely overlooked problem is the integration of multiple sources of data when one or more of them are completely missing for a subset of patients, a relatively common condition in clinical practice. Results We propose miss-Similarity Network Fusion (miss-SNF), a novel general-purpose data integration approach designed to manage completely missing data in the context of patient similarity networks. miss-SNF integrates incomplete unimodal patient similarity networks by leveraging a non-linear message-passing strategy borrowed from the SNF algorithm. miss-SNF is able to recover missing patient similarities and is “task agnostic”, in the sense that can integrate partial data for both unsupervised and supervised prediction tasks. Experimental analyses on nine cancer datasets from The Cancer Genome Atlas (TCGA) demonstrate that miss-SNF achieves state-of-the-art results in recovering similarities and in identifying patients subgroups enriched in clinically relevant variables and having differential survival. Moreover, amputation experiments show that miss-SNF supervised prediction of cancer clinical outcomes and Alzheimer’s disease diagnosis with completely missing data achieves results comparable to those obtained when all the data are available. Availability and implementation miss-SNF code, implemented in R, is available at https://github.com/AnacletoLAB/missSNF.

Biochemistry & Molecular Biology↗

AGR 5/6/7 Data Qualification Report for ATR Cycles 162B through 168A

This report provides the qualification status of experimental data for the Advanced Gas Reactor (AGR) 5/6/7 fuel irradiation. AGR-5/6/7 was conducted in the Advanced Test Reactor (ATR) at Idaho National Laboratory (INL) in support of development and qualification of tri-structural isotropic (TRISO) low-enriched fuel for use in high temperature gas-cooled reactors. The objectives of the AGR-5/6/7 experiments are to: (i) irradiate reference-design fuel particles to support fuel qualification, (ii) establish operating margins for the fuel beyond normal operating conditions, and (iii) provide irradiated-fuel performance data and irradiated-fuel samples for post-irradiation examination (PIE) and safety testing. The test train contains five separate capsules that were independently controlled and monitored. Each capsule contains multiple 12.51-mm-long compacts filled with low enriched uranium carbide/oxide (UCO) TRISO fuel particles. The primary objective of the AGR-5/6 test (Capsules 1, 2, 4, and 5) is to verify successful performance of the reference-design fuel under normal operating conditions. The AGR-7 test (Capsule 3) was designed to explore fuel performance at higher temperatures to demonstrate the capability of the fuel to withstand conditions beyond normal operating conditions in support of plant design and licensing. AGR 5/6/7 will also provide irradiated-fuel performance data on fission-gas release from failed particles during irradiation. The AGR-5/6/7 capsules were irradiated in the ATR northeast flux trap location. The experiment began on February 16, 2018 and ended on July 22, 2020, spanning nine ATR cycles over two and a half years. Thus, the AGR-5/6/7 fuel compacts were irradiated for a total of 360.9 effective full power days. The AGR 5/6/7 experiment was able to remain in the reactor core during all three Powered Axial Locator Mechanism (PALM) cycles (163A, 165A, and 167A) without overheating its fuel compacts. This report includes irradiation monitoring data from nine ATR Cycles: 162B, 163A, 164A, 164B, 165A, 166A, 166B, 167A, and 168A, as stored in the Nuclear Data Management and Analysis System (NDMAS). During irradiation, data records consisted of instantaneous measurements recorded every minute and provided by text files automatically every 2 hours. The AGR 5/6/7 data streams addressed in this report include thermocouple (TC) temperatures, sweep gas data (flow rates [capsule inlet, outlet, and downstream at detector], pressure, and moisture content), and Fission Product Monitoring System (FPMS) data (release rates and release to birth rate ratios [R/Bs]) for each of the five capsules. A total of 94,989,908 TC temperature and sweep gas data records were received and processed by NDMAS for AGR 5/6/7 irradiation. Of these records, 41,593,387 (or 43.7% of the total) met data collection and accuracy requirements and are labeled as Qualified. A total of 57,746,693 TC temperature readings were captured from 54 installed TCs. Among them, 10,034,676 TC temperature records (only 17.4%) were Qualified and 47,701,371 TC temperatures (or 82.6%) are Failed due to 48 TC failures (63.5%) and due to missing values (19.1%). To assess performance of the operational TCs, analysis of daily correlations between TCs found no evidence of virtual junction failure for any TCs. Analyses on control charts of TC temperature differences revealed trending in TC readings for TC2, 4, 5, and 13 in Capsule 3, but there is no conclusive indication of TC drift failure that caused those trends. Therefore, TC control charts are not used to disqualify TC data, but only for users’ consideration. For sweep gas flow rates, a total of 31,519,747 gas flow records (84.4%) are Qualified for use for AGR-5/6/7 experiment; 5,723,468 gas flow records (15.4%) are Failed due mostly to missing values; and 74,641 high sweep gas flow rates (0.2 %) are Trend. A large number of Failed missing TC temperature and gas flow values were caused by an error in the data output script that outputted a ‘NULL’ value when values were unchanged. This problem was fixed during the outage of Cycle 166B, which led to a substantially decreased number of missing values during the last three cycles. Nonetheless, a large amount of non-missing data remained because of the high data acquisition frequency (1-minute) and still provided sufficient data to effectively monitor the experiment as designed. For FPMS data, NDMAS received and processed fission product release and R/B data for nine ATR cycles, when ATR core reached full power during AGR 5/6/7 irradiation. These data consist of 110,388 release rate records and 110,388 R/B records for the twelve radionuclides (Kr 85m, Kr 87, Kr 88, Kr 89, Kr 90, Xe 131m, Xe 133, Xe 135, Xe 135m, Xe 137, Xe 138, and Xe 139) for each of the five capsules. Equivalent numbers of uncertainty records associated the release rates and R/B values were provided. To date, qualification status of the FPMS data stored in the NDMAS dat

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Machine learning for seismic low-frequency extrapolation

The cycle-skipping problem that plagues full waveform inversion (FWI) can be at least partially mitigated if low frequencies (which encode the kinematics of wave propagation in seismic data) are recorded. However, seismic sources and receivers are band-limited, so seismic data does not generally include signals down to 0 Hz. To improve our ability to solve the seismic inverse problem, one can synthesize this missing low-frequency (LF) content from the recorded high-frequency (HF) data using machine learning (ML) models. Deep learning models such as convolutional neural networks (CNNs) demonstrate impressive ability to perform low frequency extrapolation. However, such models require powerful hardware (GPU machines) and careful training. We assess the extrapolation capabilities of three different ML models that do not require GPU machines, namely, random forest, Gaussian process regression and gradient boosting, on both synthetic and real data. Experimental results on two synthetic data sets (generated from a low velocity lens embedded in a homogeneous medium, and the Marmousi model) demonstrate that FWI applied to the extrapolated data consistently improves inversion accuracy relative to FWI applied to the original data sets that do not contain low frequencies. Application of low-frequency extrapolation to real data from the Northwest Shelf of Australia demonstrates that tree-based ML models such as gradient boosting can outperform CNNs in terms of both accuracy and computational cost on non-GPU architectures.

58 GEOSCIENCES↗

Reprocessing Microflare Data

The report concerns work on detecting and cataloging solar microflares using an automated. An accompanying figure represents the solar microflare distribution during the period of April 1991 to November 1992, the height of solar activity after the launch of CGRO. It also shows the distribution extending below the distribution obtained at GSFC by manual means. We have implemented significant refinements in the search algorithm. The algorithm in its simplest form searches for transient events and based upon the distribution of the signal among the different BATSE detectors, we can assign it to be of solar origin if the signal distribution conforms to what one expects from a burst or transient from that direction. One of the major problems in an earlier effort was to search for microflares and large flares simultaneously. The requirement for a dynamic range of almost 10(exp 4) resulted in ambiguous identifications at the low side of the distribution. We have since restricted the search to events with peak count rates under 2000/s. Larger events are easily identified in the manual search, so we have chosen not to duplicate that work. The second problem was that missing counts existed below channel 0 in the BATSE Large Area Detector (LAD) data. These have been recovered and are now included in the search process. This provides data below 20 keV, and as we get closer to the thermal part of the spectrum, it provides greater sensitivity. The third problem was that too many BATSE detectors were used in the search. Detectors with pointing directions far from the Sun, although detecting the event, had poorly known responses. Detectors greater than approximately 60 degrees off the Sun are no longer included in the search process. By reducing the systematic errors with the large off-axis detectors we can conduct more rigorous statistical tests of a candidate event to ascertain whether it originated from the solar direction. We have reprocessed the period in the early mission that covers solar maximum and constructed the microflare distribution shown in the figure. The results of the automated search start to deviate from the manual search results below about 1000/s. Not only do we now have this distribution but we have a database of solar microflares that was used to construct the distribution. This database contains the signal at higher energy channels as well as that in channel zero (and below). From this one can, using software at GSFC, construct a photon spectrum for some of the larger microflares. It can also be used in other solar studies, especially those that correlate the X-ray flux with emission at other wavelengths. With some additional effort we hope to integrate this database into the corresponding one residing at the Solar Data Analysis Center at GSFC. The entire CGRO mission's data can now be reprocessed to obtain the microflare distribution at all phases of the solar cycle. This work is in progress. The results of this work will be presented in forthcoming scientific workshops and conferences.

Ryan, James M.↗

Bayesian Analysis of the Cosmic Microwave Background

There is a wealth of cosmological information encoded in the spatial power spectrum of temperature anisotropies of the cosmic microwave background! Experiments designed to map the microwave sky are returning a flood of data (time streams of instrument response as a beam is swept over the sky) at several different frequencies (from 30 to 900 GHz), all with different resolutions and noise properties. The resulting analysis challenge is to estimate, and quantify our uncertainty in, the spatial power spectrum of the cosmic microwave background given the complexities of "missing data", foreground emission, and complicated instrumental noise. Bayesian formulation of this problem allows consistent treatment of many complexities including complicated instrumental noise and foregrounds, and can be numerically implemented with Gibbs sampling. Gibbs sampling has now been validated as an efficient, statistically exact, and practically useful method for low-resolution (as demonstrated on WMAP 1 and 3 year temperature and polarization data). Continuing development for Planck - the goal is to exploit the unique capabilities of Gibbs sampling to directly propagate uncertainties in both foreground and instrument models to total uncertainty in cosmological parameters.

methods - statistical↗

Atmospheric Disturbance Environment Definition

Traditionally, the application of atmospheric disturbance data to airplane design problems has been the domain of the structures engineer. The primary concern in this case is the design of structural components sufficient to handle transient loads induced by the most severe atmospheric "gusts" that might be encountered. The concern has resulted in a considerable body of high altitude gust acceleration data obtained with VGH recorders (airplane velocity, V, vertical acceleration, G, altitude, H) on high-flying airplanes like the U-2 (Ehernberger and Love, 1975). However, the propulsion system designer is less concerned with the accelerations of the airplane than he is with the airflow entering the system's inlet. When the airplane encounters atmospheric turbulence it responds with transient fluctuations in pitch, yaw, and roll angles. These transients, together with fluctuations in the free-stream temperature and pressure will disrupt the total pressure, temperature, Mach number and angularity of the inlet flow. For the mixed compression inlet, the result is a disturbed throat Mach number and/or shock position, and in extreme cases an inlet unstart can occur (cf. Section 2.1). Interest in the effects of inlet unstart on the vehicle dynamics of large, supersonic airplanes is not new. Results published by NASA in 1962 of wind tunnel studies of the problem were used in support of the United States Supersonic Transport program (SST) (White, at aI, 1963). Such studies continued into the late 1970's. However, in spite of such interest, there never was developed an atmospheric disturbance database for inlet unstart analysis to compare with that available for the structures load analysis. Missing were data for the free-stream temperature and pressure disturbances that also contribute to the unStart problem.

Tank, William G.↗

Statistical Performance of Forced Oscillation Detectors in the Presence of Missing Measurements

In bulk power systems, measurement-based monitoring for large oscillations can help maintain system reliability. One of the challenges encountered in a recent field demonstration was the unavailability of measurements due to underlying measurement quality or communication problems. During the demonstration, the oscillation detector ignored a measurement location if even 10 seconds of data was missing. To extend the detector's ability to operate in these conditions, this paper evaluates the impact of three methods for addressing missing data. The strengths and weaknesses of each approach are evaluated using theoretical expressions for the probability of detection along with results from simulated data and publicly available field measurements. Based on these results, a suitable approach is identified that can extend the oscillation detector's performance when large segments of data are missing.

Follum, James D.↗

Explaining Missing Data in Graphs: A Constraint-based Approach

Abstract: This paper introduces a constraint-based approach to clarify missing values in graphs. Our method capitalizes on a set S of graph data constraints. An explanation is a sequence of operational enforcement of S towards the recovery of interested yet missing data (e.g., attribute values, edges). We show that constraint-based approach helps us to understand not only why a value is missing, but also how to recover the missing value. We study S-explanation problem, which is to compute the optimal explanations with guarantees on the informativeness and conciseness. We show the problem is in ?P^2 for established graph data constraints such as graph keys and graph association rules. We develop an efficient bidirectional algorithm to compute optimal explanations, without enforcing S on the entire graph. We also show our algorithm can be easily extended to support graph refinement within limited time, and to explain missing answers. Using real-world graphs, we experimentally verify the effectiveness and efficiency of our algorithms.

Data Analytics↗