Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Missing Data Recovery”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

A Deep Learning Approach for In-Network Synchrophasor Missing Data Recovery Using Programmable Network Switches

Phasor measurement unit (PMU) networks deliver accurate and timely measurements, which is essential for managing today’s electric power systems. To ensure data quality and enhance the cyber-resilience of PMU networks against malicious attacks and data errors, this study presents an online PMU missing data recovery scheme by leveraging P4 programmable switches. The data plane incorporates a customized PMU protocol parser that abstracts the necessary payload data for recovery. Recovery processes are executed in the control plane using a pre-trained machine learning model. Both traditional and advanced ML models, such as transformer and TimeGPT, are explicitly employed for data prediction. This approach ensures rapid and precise data recovery. Performance evaluations focus on recovery speed and accuracy, using a real dataset from a campus microgrid. With 20% missing PMU data, the mean absolute percentage error for voltage magnitude is 0.0384%, and the phase angle error discrepancy is approximately 0.4064%.

Phasor Measurement Unit, Machine Learning, Program↗

A Robust Event Diagnostics Platform: Integrating Tensor Analytics and Machine Learning into Real-time Grid Monitoring

The objective of this project is to develop a robust event diagnostics (RED) platform by integrating state-of-the-art tensor analytics and machine learning into real-time grid monitoring. The proposed platform can effectively analyze and discover the information hiding within the provided PMU data for effective real-time grid monitoring. The proposed RED platform provides a set of robust diagnostics tools for grid operation and management, including 1) data quality assessment, 2) data completion, 3) event detection, and 4) robust event classification. All the functionalities of the RED platform can help the operator to make informed decisions and respond in a timely manner. The developed RED platform will serve as an innovative advisory tool to reliably identify key events and discover new insights about the events and grid characteristics in the PMU data, and contribute to the efficient, safe, reliable operation and design of the nation’s electric system.

24 POWER TRANSMISSION AND DISTRIBUTION↗

miss-SNF: a multimodal patient similarity network integration approach to handle completely missing data sources

Abstract Motivation Precision medicine leverages patient-specific multimodal data to improve prevention, diagnosis, prognosis, and treatment of diseases. Advancing precision medicine requires the non-trivial integration of complex, heterogeneous, and potentially high-dimensional data sources, such as multi-omics and clinical data. In the literature, several approaches have been proposed to manage missing data, but are usually limited to the recovery of subsets of features for a subset of patients. A largely overlooked problem is the integration of multiple sources of data when one or more of them are completely missing for a subset of patients, a relatively common condition in clinical practice. Results We propose miss-Similarity Network Fusion (miss-SNF), a novel general-purpose data integration approach designed to manage completely missing data in the context of patient similarity networks. miss-SNF integrates incomplete unimodal patient similarity networks by leveraging a non-linear message-passing strategy borrowed from the SNF algorithm. miss-SNF is able to recover missing patient similarities and is “task agnostic”, in the sense that can integrate partial data for both unsupervised and supervised prediction tasks. Experimental analyses on nine cancer datasets from The Cancer Genome Atlas (TCGA) demonstrate that miss-SNF achieves state-of-the-art results in recovering similarities and in identifying patients subgroups enriched in clinically relevant variables and having differential survival. Moreover, amputation experiments show that miss-SNF supervised prediction of cancer clinical outcomes and Alzheimer’s disease diagnosis with completely missing data achieves results comparable to those obtained when all the data are available. Availability and implementation miss-SNF code, implemented in R, is available at https://github.com/AnacletoLAB/missSNF.

Biochemistry & Molecular Biology↗

AlphaFold -assisted structure determination of a bacterial protein of unknown function using X-ray and electron crystallography

Macromolecular crystallography generally requires the recovery of missing phase information from diffraction data to reconstruct an electron-density map of the crystallized molecule. Most recent structures have been solved using molecular replacement as a phasing method, requiring an a priori structure that is closely related to the target protein to serve as a search model; when no such search model exists, molecular replacement is not possible. New advances in computational machine-learning methods, however, have resulted in major advances in protein structure predictions from sequence information. Methods that generate predicted structural models of sufficient accuracy provide a powerful approach to molecular replacement. Taking advantage of these advances, AlphaFold predictions were applied to enable structure determination of a bacterial protein of unknown function (UniProtKB Q63NT7, NCBI locus BPSS0212) based on diffraction data that had evaded phasing attempts using MIR and anomalous scattering methods. Using both X-ray and micro-electron (microED) diffraction data, it was possible to solve the structure of the main fragment of the protein using a predicted model of that domain as a starting point. The use of predicted structural models importantly expands the promise of electron diffraction, where structure determination relies critically on molecular replacement.

molecular replacement↗

Sequential Image Recovery Using Joint Hierarchical Bayesian Learning

Abstract Recovering temporal image sequences (videos) based on indirect, noisy, or incomplete data is an essential yet challenging task. We specifically consider the case where each data set is missing vital information, which prevents the accurate recovery of the individual images. Although some recent (variational) methods have demonstrated high-resolution image recovery based on jointly recovering sequential images, there remain robustness issues due to parameter tuning and restrictions on the type of sequential images. Here, we present a method based on hierarchical Bayesian learning for the joint recovery of sequential images that incorporates prior intra- and inter-image information. Our method restores the missing information in each image by “borrowing” it from the other images. More precisely, we couple sequential images by penalizing their pixel-wise difference. The corresponding penalty terms (one for each pixel and pair of subsequent images) are treated as weakly-informative random variables that favor small pixel-wise differences but allow occasional outliers. As a result, all of the individual reconstructions yield improved accuracy. Our method can be used for various data acquisitions and allows for uncertainty quantification. Some preliminary results indicate its potential use for sequential deblurring and magnetic resonance imaging.

Xiao, Yao↗

Bayesian High-Rank Hankel Matrix Completion for Nonlinear Synchrophasor Data Recovery

Phasor measurement units (PMUs) provide high temporal-resolution synchrophasor measurements for power system monitoring and control. The frequent data quality issues, such as missing and bad data, prevent the incorporation of synchrophasor data in real-time operations. Most existing data-driven data recovery methods assume the power system dynamics can be approximated by a linear dynamical system, and the recovery performance degrades significantly when the power system is experiencing nonlinear dynamics during significant events. Here, this paper proposes a data-driven Bayesian nonlinear synchrophasor data recovery method (Ba-NSDR) that can recover a consecutive time period of simultaneous data losses or errors across all channels, even when the underlying system is highly nonlinear. The idea is to lift the Hankel matrix of the spatial-temporal synchrophasor data to a higher dimension such that the lifted Hankel matrix is low-rank in that space and can be processed with the kernel trick. Our proposed Bayesian method then infers the probabilistic distributions of synchrophasor from the partial observations. Some distinctive features of Ba-NSDR include an uncertainty index to measure the accuracy of the recovery result and the robustness to parameter selections. Our method is verified on both synthetic and recorded event datasets.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Introduction to Evolve Central Appalachia

The Evolve-CAPP project is developing and implementing strategies to enable the Central Appalachian basin to realize its full economic potential for producing high-value, non-fuel, carbon-ore (CO) products, rare earth elements (REE), and critical minerals (CM). The Basinal Assessment of CORE-CM Resources is a state-of-the-art geological model and comprehensive database that will be used to identify, locate, and quantify in-situ resources in the basin. Restrictive aspects of resource recovery, such as ownership and legal restrictions, are being considered in estimating recoverable resources. A gap analysis of missing or unavailable data that could not be fully integrated in the initial assessment due to budget, access, or other constraints during the assessment effort is planned. A characterization and data acquisition plan to fully characterize the basin’s CORE-CM potential will address gaps in data. The plan will be developed with input from project team members and relevant stakeholders, including land and mineral holding companies and coal operators, who maintain sampling programs as part of their normal operations.

Bishop, Richard↗

Data about data – when, why and how metadata can support the digital plant

A structured approach for recording data quality and contextual information about how and why a signal exists – i.e. metadata – is central to interpret and use sensor data correctly. This is becoming increasingly important with the global trend with data-driven applications such as digital twins and AI-models. But a structured metadata collection and organization of sensor data is not routine in most plants, which can result in lost information and missed opportunities to make use of the investments made in the data collection. Therefore, the IWA task group on Metadata Collection and Organization in wastewater resource recovery systems (MetaCO) was initiated in 2020 and recently delivered the IWA scientific and technical report number 31. The report gives and in-depth description about metadata in water resources recovery facilities (WRRFs) and is available as open access at IWA publishing. The report is the outcome of the collaboration between more than 80 water professionals with the intention to serve WRRF data users with a guide on how to structure and make use of metadata throughout the data pipeline in order to maximize the value of sensor data.

Alferes, Janelcy [VITO, Belgium]↗

Bayesian inference of anisotropic 2D small-angle scattering from sparse measurement

Here, we present a Bayesian inference framework for reconstructing anisotropic two-dimensional small-angle scattering (2D SAS) patterns from sparse, noisy, or partially missing data. The method combines a symmetry-aware angular basis with radial Gaussian process priors to enable accurate, training-free interpolation and denoising. Computational benchmarks demonstrate reliable recovery of both isotropic and high-order anisotropic features under severe data reduction. Experimental validations on stretched polymers, sheared wormlike micelles, and carbon fibers show improved fidelity and resolution compared to raw measurements, achieving comparable accuracy with up to 50-fold fewer detected neutrons. This approach enables quantitative structural analysis under low-flux, time-limited, or single-shot conditions, extending the applicability of 2D SAS techniques to compact neutron sources and mechanically driven soft matter systems undergoing transient structural changes.

Tung, Chi-Huan [Oak Ridge National Laboratory (ORN↗

Evaluating the Hybrid Modelling Competition: A Step Towards Developing Good Modelling Practice

Hybrid modelling, a combination of mechanistic and data-driven modelling, is a promis¬ing approach to advance current mathematical models towards improved deci¬sion support tools for today's water-related challenges. Researchers have been develop¬ing guidelines or references for good modelling practices in the water field for mecha¬nistic (Rieger et al., 2012) and data-driven (Zhu et al., 2023) modelling, respectively. However, good modelling practices for hybrid modelling are currently missing (Schneider et al., 2022). Therefore, the International Water Association’s (IWA) hybrid modelling working group initiated the first competition on a data science competition platform (i.e. Kaggle) for water resource recovery modelling at the Watermatex con¬ference in September 2023 in Quebec. The main objective of this competition was to gain insights and experience to create good modelling practices. Further goals were to motivate students, researchers, and practitioners model, foster a vibrant and engaged community, and evaluate the efficacy of com¬petitions in solving modelling challenges within the water domain. Our next goal is that facilities will measure and gather relevant data for future competitions to solve their challenges from a modeller’s perspective.

Schneider, Mariane↗

Mixture Model Framework for Traumatic Brain Injury Prognosis Using Heterogeneous Clinical and Outcome Data

Prognoses of Traumatic Brain Injury (TBI) outcomes are neither easily nor accurately determined from clinical indicators. This is due in part to the heterogeneity of damage inflicted to the brain, ultimately resulting in diverse and complex outcomes. Using a data-driven approach on many distinct data elements may be necessary to describe this large set of outcomes and thereby robustly depict the nuanced differences among TBI patients’ recovery. In this work, we develop a method for modeling large heterogeneous data types relevant to TBI. Our approach is geared toward the probabilistic representation of mixed continuous and discrete variables with missing values. The model is trained on a dataset encompassing a variety of data types, including demographics, blood-based biomarkers, and imaging findings. In addition, it includes a set of clinical outcome assessments at 3, 6, and 12 months post-injury. The model is used to stratify patients into distinct groups in an unsupervised learning setting. We use the model to infer outcomes using input data, and show that the collection of input data reduces uncertainty of outcomes over a baseline approach. In addition, we quantify the performance of a likelihood scoring technique that can be used to self-evaluate the extrapolation risk of prognosis on unseen patients.

97 MATHEMATICS AND COMPUTING↗

Deciphering the Role of Total Water Storage Anomalies in Mediating Regional Flooding

Regional floods result from various flood generation mechanisms. Traditional analyses mainly link flooding to extreme rainfall, with limited input from soil moisture. Total water storage (TWS) is a holistic measure of basin wetness, including additional storage components from surface water, snow, and groundwater. Utilizing a new 5-day Gravity Recovery and Climate Experiment and its Follow On (GRACE(-FO)) data set, we investigated the linkage between short-term TWS anomaly (TWSA) and regional flooding. The 5-day TWSA solutions revealed flood signals missed by monthly TWSA solutions. Global basins exhibit distinct storage-discharge co-evolution patterns, offering new insights into flood mechanisms and propensity. Our bivariate event analyses show the annual maximum river discharges co-occur more often with the TWSA maxima than with precipitation in many basins. Further analyses revealed TWSA's time-lagged effect on river discharge, particularly in basins susceptible to floods triggered by saturation-excess runoff. The 5-day TWSA provides a new source of information for enhancing global flood preparedness.

54 ENVIRONMENTAL SCIENCES↗

National Energy Water Treatment & Speciation (NEWTS): A Water & Critical Mineral Database and Dashboard

The scarcity of water resources, the need for beneficial water reuse, and the challenges of wastewater treatment are becoming increasingly pressing in economic, social, and environmental domains. Addressing these concerns requires effective treatment strategies to manage wastewater streams and tackle environmental and economic issues. Furthermore, the recovery of critical minerals from the waste streams associated with energy production holds the promise of offsetting treatment costs and securing local sources of valuable minerals. However, relevant data on these waste streams are dispersed and challenging to locate. The process of ingesting such data into modeling software often involves multiple steps, requiring data restructuring to meet software-input requirements. The non-standardized reporting of water data makes data aggregation and reformatting a time-consuming process. Additionally, essential attributes necessary for modeling water treatment and mineral scale formation are frequently missing. Moreover, data gaps vary depending on the region of interest. Consequently, there is a pressing need for high-quality energy-water composition data that can be easily imported into water chemistry modeling software. To address this need, the National Energy Technology Laboratory has created the National Energy Water Treatment and Speciation (NEWTS) Database and Dashboard—a free online tool catering to community leaders and water researchers. NEWTS facilitates a comprehensive understanding of the composition of energy-related wastewater streams in the United States. The datasets provide detailed concentrations and speciation of major and minor aqueous compounds in energy-related wastewater streams, including power plant leachate, acid mine drainage, brackish water, and oil and gas produced water across the United States. Many of the aqueous species are critical minerals (Li, REEs) in high demand to modernize the world’s energy infrastructure. Many of the datasets also contain volumetric flow-rates needed to model the treatment and reuse scenarios in advanced aqueous chemistry software programs. The NEWTS Database and Dashboard offer public access to hitherto challenging-to-access datasets, presented in a standardized format that is tailored for easy input into aqueous chemistry modeling software. By performing the work needed to transform dispersed, disparate data sources into unified, model-ready datasets, NEWTS serves as an essential resource in advancing water treatment research and sustainable water resource management.

produced water management↗

Pennsylvania Department of Environmental Protection (PA DEP) 26r Detailed Produced Water Compositions (version 1.0)

A database of geochemical compositions of aqueous species in produced water reported to the PA DEP. Samples were collected between mid-2012 to early-2020. Data from publicly-available PA DEP 26r reports were scraped from pdf files and cumulated into tabular spreadsheet format for >1000 produced water streams from Marcellus wells in Pennsylvania. In addition to providing the original values, the NETL NEWTS team has reformatted the dataset to allow sample streams to be easily copied into OLI Studio and Geochemist WorkBench (GWB) software for modeling the geochemistry and the recovery of critical minerals, such as lithium, from these produced water streams. In addition, a version of the dataset has been included with predictions for some missing values in the original dataset using machine learning techniques within CoDaRT software, a public ML software developed by the Nation Energy Technology Laboratory. We have made the Input into CoDaRT and one example output from CoDaRT available in this dataset.

Aqueous Chemistry↗

Study on Application of Distributed Network of Sensors with List Mode for NMAC Literature Review

Nuclear material accounting and control (NMAC) for nuclear security detects, deters, and resolves questions related to unauthorized removal (i.e. theft) or misuse of nuclear material. NMAC also serves as a key insider threat mitigation measure and aids in recovery of nuclear material that is missing. Effective nuclear security depends on NMAC for timely and accurate information about nuclear material types, quantities, and locations. Bulk nuclear material processing facilities, however, present unique challenges for effective NMAC due to the presence of large quantities of material in-process and the accumulation of residual material holdup within process equipment. These holdup accumulations can obscure accurate physical inventory taking and complicate efforts to resolve NMAC irregularities at the facility level. Bulk material monitoring systems often rely on material balance calculations and indirect measurement techniques, which may mask protracted theft of smaller amounts of nuclear material. These monitoring limitations have generated increased interest in continuous monitoring technologies, including distributed non-destructive assay (NDA) sensor networks capable of providing real-time or near-real-time measurement of material movement and accumulation within bulk processing environments. Recent advancements in distributed networks of NDA radiation detectors and sensing technologies provide an opportunity to address these limitations. Although such distributed sensor networks have been implemented in select facilities for IAEA Safeguards applications, their potential for supporting NMAC functions specifically tailored to nuclear security objectives remains largely unexplored. Furthermore, emerging list-mode data acquisition technologies have reached high technology readiness levels, enabling time-correlated detection of nuclear events across multiple temporal scales. These capabilities provide enhanced opportunities for accurate holdup measurement, continuous process monitoring, and improved detection of material theft or misuse over time. The increasing global expansion of civil nuclear power and development of related bulk material processing facilities, including those supporting high-assay low-enriched uranium (HALEU) and other advanced reactor fuel fabrication, further increases the need for advanced measurement and monitoring strategies for NMAC.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗