Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data fusion”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Uncertainty based Online Ensemble on Non-Stationary Data for Fusion Science

Machine Learning (ML) is poised to play a pivotal role in the development and operation of next-generation fusion devices. Fusion data shows non-stationary behavior due to drifts in the data. The drifts can arise from both experimental evolution and machine wear-and-tear. ML models assume stationary distribution and fail to maintain performance when encountered with non-stationary data streams.Online learning can be used to continuously adapt the models with new data as it is acquired. However, traditional online learning can suffer from short-term performance degradation, as ground truth are not available before making the prediction. To address this challenge, we propose uncertainty aware ensemble approach for online learning. We use Deep Gaussian Process Approximation (DGPA) technique for calibrated uncertainty estimation and use the uncertainty values to guide a meta-algorithm that produces predictions based on ensemble of learners. Moreover, DGPA also provides uncertainty estimation along with the predictions for decision makers. This paper demonstrates that the proposed method outperforms traditional online learning approach, and a naive ensemble without uncertainty guidance by about 7% and 6%, respectively, on B-coil deflection prediction at DIII-D Fusion Facility.

Rajput, Kishansingh [Thomas Jefferson National Acc↗

Uncertainty based Online Ensemble on Non-Stationary Data for Fusion Science

Machine Learning (ML) is poised to play a pivotal role in the development and operation of next-generation fusion devices. Fusion data shows non-stationary behavior due to drifts in the data. The drifts can arise from both experimental evolution and machine wear-and-tear. ML models assume stationary distribution and fail to maintain performance when encountered with non-stationary data streams.Online learning can be used to continuously adapt the models with new data as it is acquired. However, traditional online learning can suffer from short-term performance degradation, as ground truth are not available before making the prediction. To address this challenge, we propose uncertainty aware ensemble approach for online learning. We use Deep Gaussian Process Approximation (DGPA) technique for calibrated uncertainty estimation and use the uncertainty values to guide a meta-algorithm that produces predictions based on ensemble of learners. Moreover, DGPA also provides uncertainty estimation along with the predictions for decision makers. This paper demonstrates that the proposed method outperforms traditional online learning approach, and a naive ensemble without uncertainty guidance by about 7% and 6%, respectively, on B-coil deflection prediction at DIII-D Fusion Facility.

Rajput, Kishansingh [Thomas Jefferson National Acc↗

Uncertainty guided online ensemble for non-stationary data streams in fusion science

Machine Learning (ML) is poised to play a pivotal role in the development and operation of next-generation fusion devices. Fusion data shows non-stationary behavior with distribution drifts, resulted by both experimental evolution and machine wear-and-tear. ML models assume stationary distribution and fail to maintain performance when encountered with such non-stationary data streams. Online learning techniques have been leveraged in other domains, however it has been largely unexplored for fusion applications. In this paper, we investigate online learning for continuous adaptation to drifting data streams in the prediction of Toroidal Field (TF) coils deflection at the DIII-D fusion facility. We further address the short-term performance degradation inherent to standard online learning, which arises because ground truth is unavailable at prediction time. To mitigate this issue, we propose an uncertainty-guided online ensemble framework. The method leverages the Deep Gaussian Process Approximation (DGPA) for calibrated uncertainty estimation and uses these uncertainty measures to guide a meta-algorithm that aggregates predictions from learners trained over different historical horizons. Our results show that online learning reduces prediction error by 80% compared to a static model. The online ensemble and the proposed uncertainty-guided ensemble further reduce error by approximately 6%, and 10% respectively, relative to standard single-model online learning, while also providing calibrated uncertainty estimates to support operational decision-making.

AI↗

Information Fusion and Data Analytics for Human Lunar Exploration (CIF REPORT: Detailed PI Write-up)

The Information Fusion & Data Analytics (IFDA) project commenced in FY20, continued through FY21, and its final platform development phase continues in FY22. The objective remains the fusion and rapid accessibility of large quantities of disparate sourced human spaceflight data. IFDA is a platform tailored for NA (S&MA) to develop highly advanced operational data integration and analysis techniques. IFDA leverages the JSC ER7 modeling, simulation,and data fusion capabilities to collect, warehouse, and augment data human exploration data integration and analysis techniques. The IFDA project’s integrated data visualizations have been demonstrated in two validation scenarios in FY21, and provided the architecture and platform basis for development of a full-scale data analysis suite and storage solution useful to all JSC organizations engaged in real time operations and safety tasks. Scenarioand prototypical development including the construction of a full scale data analysis suite and storage solution, useful to all JSC organizations engaged in real time operations and safety tasks, is central to IFDA Phase 3 and provides a demonstrable pathway for the Digital Transformation Program. IFDA Phase 3 is focused on data provider, data utilizer, and SME hands-on workshops that will conclude the Dem / Valphase and deliver a program-ready data integration tool as a product.

information fusion↗

Daily Ambient Air Pollution Metrics for Five Cities: Evaluation of Data Fusion-Based Estimates and Uncertainties

Spatiotemporal characterization of ambient air pollutant concentrations is increasingly relying on the combination of observations and air quality models to provide well-constrained, spatially and temporally complete pollutant concentration fields. Air quality models, in particular, are attractive, as they characterize the emissions, meteorological, and physiochemical process linkages explicitly while providing continuous spatial structure. However, such modeling is computationally intensive and has biases. The limitations of spatially sparse and temporally incomplete observations can be overcome by blending the data with estimates from a physically and chemically coherent model, driven by emissions and meteorological inputs. We recently developed a data fusion method that blends ambient ground observations and chemical transport-modeled (CTM) data to estimate daily, spatially resolved pollutant concentrations and associatedcorrelations. In this study, we assess the ability of the data fusion method to produce daily metrics (i.e., 1-hr max, 8-hr max, and 24-hr average) of ambient air pollution that capture spatiotemporal air pollution trends for 12 pollutants (CO, NO2, NOx, O3, SO2, PM (sub10), PM (sub 2.5), and five PM (sub 2.5) components) across five metropolitan areas (Atlanta, Birmingham, Dallas, Pittsburgh, and St. Louis), from 2002 to 2008. Three sets of comparisons are performed: (1) the CTM concentrations are evaluated for each pollutant and metropolitan domain, (2) the data fusion concentrations are compared with the monitor data, (3) a comprehensive cross-validation analysis against observed data evaluates the quality of the data fusion model simulations across multiple metropolitan domains. The resulting daily spatial field estimates of air pollutant concentrations and uncertainties are not only consistent with observations, emissions, andmeteorology, but substantially improve CTM-derived results for nearly all pollutants and all cities, with the exception of NO2 for Birmingham. The greatest improvements occur for O3 and PM (sub 2.5). Squared spatiotemporal correlation coefficients range between simulations and observations determined using cross-validation across all cities for air pollutants of secondary and mixed origins are R-squared equal to 0.88-0.93 (O3), 0.81-0.89 (SO4), 0.67-0.83 (PM (sub 2.5)), 0.52-0.72 (NO3), 0.43-0.80 (NH4), 0.32-0.51 (OC), and 0.14-0.71 (PM (sub 10)). Results for relatively homogeneous pollutants of secondary origin, tend to be better than those for more spatially heterogeneous (larger spatial gradients) pollutants of primary origin (NOx, CO, SO2 and EC). Generally, background concentrations and spatial concentration gradients reflect interurban airshed complexity and the effects of regional transport, whereas daily spatial pattern variability shows intra urban consistency in the fused data. With sufficiently high CTM spatial resolution, traffic-related pollutants exhibit gradual concentration gradients that peak toward the urban centers. Ambient pollutant concentration uncertainty estimates for the fused data are both more accurate and smaller than those for either the observations or the model simulations alone.

spatiotemporal fusion↗

Unsupervised multimodal fusion of in-process sensor data for advanced manufacturing process monitoring

Effective monitoring of manufacturing processes is crucial for maintaining product quality and operational efficiency. Modern manufacturing environments often generate vast amounts of complementary multimodal data, including visual imagery from various perspectives and resolutions, hyperspectral data, and machine health monitoring information such as actuator positions, accelerometer readings, and temperature measurements. However, fusing and interpreting this complex, high-dimensional data presents significant challenges, particularly when labeled datasets are unavailable or impractical to obtain. This paper presents a novel approach to multimodal sensor data fusion in manufacturing processes, inspired by the Contrastive Language-Image Pre-training (CLIP) model. We leverage contrastive learning techniques to correlate different data modalities without the need for labeled data, overcoming limitations of traditional supervised machine learning methods in manufacturing contexts. Our proposed method demonstrates the ability to handle and learn encoders for five distinct modalities: visual imagery, audio signals, laser position (x and y coordinates), and laser power measurements. By compressing these high-dimensional datasets into low-dimensional representational spaces, our approach facilitates downstream tasks such as process control, anomaly detection, and quality assurance. The unsupervised nature of our method makes it broadly applicable across various manufacturing domains, where large volumes of unlabeled sensor data are common. We evaluate the effectiveness of our approach through a series of experiments, demonstrating its potential to enhance process monitoring capabilities in advanced manufacturing systems. This research contributes to the field of smart manufacturing by providing a flexible, scalable framework for multimodal data fusion that can adapt to diverse manufacturing environments and sensor configurations. The proposed method paves the way for more robust, data-driven decision-making in complex manufacturing processes.

Contrastive Learning↗

Fusion of Surface Ceilometer Data and Satellite Cloud Retrievals in 2D Mesh Interpolating Model with Clustering

For accurate cloud ceiling information, a data fusion approach is proposed that utilizes satellite data to extend surface station information to much wider areas. Cloud base height (CBH) retrieved from satellite observations provides for much larger spatial coverage and higher resolution. The direct comparison of GOES-16 CBH with surface station ceiling yields a local bias that has to be corrected for in the initial GOES-16 cloud base information. This sparsely sampled bias correction presents an irregular 2D mesh of control points, which is then interpolated by constructing a continuous smooth field using polyharmonic splines. The influence of remote stations is restricted by grouping the control points into clusters depending on an effective distance. This cluster-based approach allows for constructing separate spline surfaces corresponding to physically different clouds. The obtained continuous bias correction function is then applied to the entire GOES-16 pixel level CBH except for areas far away from surface stations in data sparse regions such as offshore. The described method is currently being tested using daytime-only observations over the central and eastern United States. Overall, this approach has potential to provide more accurate, high spatial resolution cloud ceiling information for the aviation community.

Khlopenkov, Konstantin↗