Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “multivariate data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

FunM2C: A Filter for Uncertainty Visualization of Multivariate Data on Multi-Core Devices

Uncertainty visualization is an emerging research topic in data visualization because neglecting uncertainty in visualization can lead to inaccurate assessments. In this paper, we study the propagation of multivariate data uncertainty in visualization. Although there have been a few advancements in probabilistic uncertainty visualization of multivariate data, three critical challenges remain to be addressed. First, the state-of-the-art probabilistic uncertainty visualization framework is limited to bivariate data (two variables). Second, existing uncertainty visualization algorithms use computationally intensive techniques and lack support for cross-platform portability. Third, as a consequence of the computational expense, integration into production visualization tools is impractical. In this work, we address all three issues and make a threefold contribution. First, we take a step to generalize the state-of-the-art probabilistic framework for bivariate data to multivariate data with an arbitrary number of variables. Second, through utilization of VTK-m’s shared-memory parallelism and cross-platform compatibility features, we demonstrate acceleration of multivariate uncertainty visualization on different many-core architectures, including OpenMP and AMD GPUs. Third, we demonstrate the integration of our algorithms with the ParaView software. We demonstrate the utility of our algorithms through experiments on multivariate simulation data with three and four variables.

Hari, Gautam↗

Evaluation of COTS Electronics by Power Spectrum Analysis and Multivariate Data Analysis

Power spectrum analysis (PSA) is a fast, non-destructive, sensitive method for examining commercial off-the-shelf ( COTS ) electronic components. These features make PSA attractive for both component screening and surveillance in support of component reliability efforts. Current analysis methods limit the utility of PSA due to the need to manually examine the results of analysis to identify anomalous parts. This study demonstrates the development and application of a workflow to automate the screening of COTS electronic components. Further, this study demonstrates the use of multivariate algorithms to assess aging of Zener diodes. These workflows can be readily extended to other components, combining the benefits of PSA and multivariate analysis to screen and evaluate COTS electronic components.

42 ENGINEERING↗

A Bayesian model for multivariate discrete data using spatial and expert information with application to inferring building attributes

When modeling sparsely observed multivariate data, strong prior information elicited from experts can be used to bolster predictive accuracy and counteract sampling bias. Similarly, modeling autocorrelation in space can help make use of co-occurrence patterns present in many types of spatial data. To make use of both expert prior information and spatial structure, we propose a novel graphical model for a spatial Bayesian network developed specifically to address challenges in inferring the attributes of buildings from geographically sparse observational data. This model is implemented as the sum of a spatial multivariate Gaussian random field and a tabular conditional probability function in real-valued space prior to projection onto the probability simplex. This modeling form is especially suitable for the usage of prior information in the form of sets of atomic rules obtained from experts. To perform inference with missing data, we implement a Markov chain Monte Carlo scheme composed of alternating steps of Gibbs sampling of missing entries and Hamiltonian Monte Carlo for model parameters. A case study in building attribution is presented to highlight the advantages and limitations of this approach.

97 MATHEMATICS AND COMPUTING↗

Fiber Uncertainty Visualization for Bivariate Data With Parametric and Nonparametric Noise Models

Visualization and analysis of multivariate data and their uncertainty are top research challenges in data visualization. Constructing fiber surfaces is a popular technique for multivariate data visualization that generalizes the idea of level-set visualization for univariate data to multivariate data. Here, in this paper, we present a statistical framework to quantify positional probabilities of fibers extracted from uncertain bivariate fields. Specifically, we extend the state-of-the-art Gaussian models of uncertainty for bivariate data to other parametric distributions (e.g., uniform and Epanechnikov) and more general nonparametric probability distributions (e.g., histograms and kernel density estimation) and derive corresponding spatial probabilities of fibers. In our proposed framework, we leverage Green's theorem for closed-form computation of fiber probabilities when bivariate data are assumed to have independent parametric and nonparametric noise. Additionally, we present a nonparametric approach combined with numerical integration to study the positional probability of fibers when bivariate data are assumed to have correlated noise. For uncertainty analysis, we visualize the derived probability volumes for fibers via volume rendering and extracting level sets based on probability thresholds. We present the utility of our proposed techniques via experiments on synthetic and simulation datasets.

97 MATHEMATICS AND COMPUTING↗

Application of Partial Least Squares Approaches to Pyroprocessing ER Data

Multivariate approaches show promise for application to process monitoring for safeguards of pyroprocessing. Past MPACT work explored the application of Principal Component Analysis (PCA) to detect off-normal conditions in pyroprocessing electrorefiner (ER) data from in the Hot Fuel Examination Facility (HFEF) at Idaho National Laboratory (INL) known as the Scalable Pyrochemical Recycling testbed (SPyRe) ER. PCA, however, does not consider the output variables. In FY24, multivariate analysis was extended from PCA to Partial Least Squares (PLS) analysis. PLS maximizes the variance between both the input signals and output variables. In the case of this work, PLS was applied in two different manners: Predictive PLS and Discriminant PLS. Predictive PLS maximizes the covariance between the process variables of the ER and the measured U concentration from in-situ voltammetry. Discriminant PLS maximizes the covariance between the process variables and a set of training process “states” such as known off-normal conditions. By projecting into the latent variable space in PLS, the process variables can be regressed onto the outputs and predictions can be made for new data sets. In this work, by applying predictive PLS, a penalized non-linear PLS approach was able to make predictions of concentration based on test and training data and detect when operations were off-normal. However, the predictive PLS does not classify the signals to which off-normal operations are attributable. Discriminant PLS can be used to classify off-normal operations but is inadequate to properly classify specific off-normal classes like power supply faults when the Discriminant PLS model is only specifically trained to detect that off-normal class. When all faults are trained against the observation data, all three operational classes are accurately classified and distinguished. Thus, future application of latent variable techniques should not select any given method, but should use a mixture of PCA, Predictive PLS, and Discriminant PLS.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Scalable Volume Visualization for Big Scientific Data Modeled by Functional Approximation

Considering the challenges posed by the space and time complexities in handling extensive scientific volumetric data, various data representations have been developed for the analysis of large-scale scientific data. Multivariate functional approximation (MFA) is an innovative data model designed to tackle substantial challenges in scientific data analysis. It computes values and derivatives with high-order accuracy throughout the spatial domain, mitigating artifacts associated with zero- or first-order interpolation. However, the slow query time through MFA makes it less suitable for interactively visualizing a large MFA model. In this work, we develop the first scalable interactive volume visualization pipeline, MFA-DVV, for the MFA model encoded from large-scale datasets. Our method achieves low input latency through distributed architecture, and its performance can be further enhanced by utilizing a compressed MFA model while still maintaining a high-quality rendering result for scientific datasets. We conduct comprehensive experiments to show that MFA-DVV can decrease the input latency and achieve superior visualization results for big scientific data compared with existing approaches.

big scientific dataset↗

Quasar Identification Using Multivariate Probability Density Estimated from Nonparametric Conditional Probabilities

Nonparametric estimation for a probability density function that describes multivariate data has typically been addressed by kernel density estimation (KDE). A novel density estimator recently developed by Farmer and Jacobs offers an alternative high-throughput automated approach to univariate nonparametric density estimation based on maximum entropy and order statistics, improving accuracy over univariate KDE. This article presents an extension of the single variable case to multiple variables. The univariate estimator is used to recursively calculate a product array of one-dimensional conditional probabilities. In combination with interpolation methods, a complete joint probability density estimate is generated for multiple variables. Good accuracy and speed performance in synthetic data are demonstrated by a numerical study using known distributions over a range of sample sizes from 100 to 10 6 for two to six variables. Performance in terms of speed and accuracy is compared to KDE. The multivariate density estimate developed here tends to perform better as the number of samples and/or variables increases. As an example application, measurements are analyzed over five filters of photometric data from the Sloan Digital Sky Survey Data Release 17. The multivariate estimation is used to form the basis for a binary classifier that distinguishes quasars from galaxies and stars with up to 94% accuracy.

79 ASTRONOMY AND ASTROPHYSICS↗

Carbon fiber classification using raman spectroscopy

Carbon fiber characterization processes are described that include multi-condition Raman spectroscopy-based examination combined with multivariate data analyses. Methods are a nondestructive material characterization approach that can provide predictions as to carbon fiber bulk physical properties, as well as identification of unknown carbon fiber materials for quality control purposes. The framework of the multivariate analysis methods includes a principal component-based identification protocol including comparison of Raman spectral data from an unknown carbon fiber with a data library of multiple principal component spaces.

Houk, Amanda L.↗

Unsupervised anomaly clustering via offset alignment in multivariate grid sensing data

Modern industries increasingly rely on multi-sensor technologies to acquire complex, high-dimensional data streams, enabling advanced monitoring and control systems. One critical application is online anomaly detection in electrical smart grids, where multivariate and multimodal sensing technologies play a vital role. However, detecting anomalies in such time-series data is challenging due to their inherent temporal dependencies and stochastic behavior. Traditional approaches based on supervised and semi-supervised learning methods depend on labeled datasets, which are often unavailable in real-world scenarios. While unsupervised methods have emerged as promising alternatives, these methods are highly susceptible to noise and outliers commonly present in sensing applications. Furthermore, deep learning-based anomaly detection methods, despite their performance, are often criticized for their black-box nature, limiting their applicability in safety-critical and online environments where interpretability and explainability are paramount. In this work, we propose an unsupervised anomaly clustering method leveraging a cyclic alignment-based offset detection algorithm for multivariate time-series signals. The proposed method is applied to multivariate data collected from vibrational, voltage, and magnetic field sensors deployed in a local grid substation. Our results demonstrate the robustness of the algorithm in accurately clustering various anomalies/events across different sensing modalities. Additionally, we compare the effectiveness of the proposed approach against a simple pattern-based anomaly detection method, which performs well for univariate data but fails to generalize to multivariate and multimodal time-series data.

Mukherjee, Subrata [ORNL] (ORCID:0000000309930338)↗

Adaptive Online Multivariate Signal Extraction With Locally Weighted Robust Polynomial Regression

High-frequency, multivariate data collected in real-time and used to control or make decisions regarding a process’ operation often contain some noise and outliers. Thus, a method to extract the signal is needed in order to reduce the number and magnitude of control-based adjustments that are implemented. Such a method must be (i) online, depending only on past and current observations; (ii) fast, producing a smooth value more quickly than the measurement frequency; (iii) robust, ignoring brief bursts of erroneously measured values; (iv) multivariate, ignoring observations that are jointly unusual; (v) adaptive, adjusting to periods of rapid fluctuation in the signal versus periods of stability; and (vi) purely data-driven, not incorporating any information about the process from which the data are collected. Most existing methods are only able to address a subset of these six features. Furthermore, we also require the method to be nonlinear, providing a local nonlinear estimate of the signal. In this work, we propose a novel, real-time signal extraction method based on a local, robust polynomial fit. We demonstrate the performance of our method compared to a state-of-the-art competitor through simulation. For illustration, the methodology is applied to data collected from a reverse osmosis water treatment process.

97 MATHEMATICS AND COMPUTING↗

Intrinsic Kinetics of Polyethylene Terephthalate Pyrolysis via Micropyrolysis and Multivariate Chromatographic Analysis

This study provides an in-depth investigation of the primary decomposition of polyethylene terephthalate (PET) via pyrolysis, employing an experimental-analytic workflow that integrates design of experiments (DoE), micropyrolysis coupled with comprehensive two-dimensional gas chromatography (GC×GC), and multivariate data analysis to verify intrinsic kinetic conditions and elucidate evolving product distributions for mapping key reaction pathways. Peaks that could not be identified using commercial spectral libraries were assigned using Mass Frontier simulations, enabling the identification of divinyl terephthalate, ethyl vinyl terephthalate, and 2-(benzoyloxy)ethyl vinyl terephthalate. A polar×polar (non-orthogonal) column set tailored for the detection of carboxylic acids enhanced the quantification of benzoic acid, 4-vinylbenzoic acid, 4-ethylbenzoic acid, and methylbenzoic acid by up to 6-fold relative to an orthogonal column combination (non-polar×mid-polar). Moreover, pyrolysis variables were systematically evaluated using a Box- Behnken design (BBD), encompassing pyrolysis temperature (500−600 °C), sample weight (50−150 μg), and carrier gas flow rate (100−300 mL min −1 ). Among these, pyrolysis temperature was the only statistically significant factor influencing product yields, ranging from 58.78 to 84.26 wt %. In contrast, neither the sample weight nor the carrier gas flow rate had a significant effect on product yields within the evaluated experimental space. At 600 °C, the major pyrolysis products were benzoic acid (up to 20.20 ± 1.46 wt %) and CO 2 (up to 21.28 ± 1.46 wt %), which can be produced through decarboxylation reactions. These findings underscore the critical importance of selecting appropriate analytical columns for the accurate quantification of heteroatomcontaining products such as carboxylic acids, which may otherwise be underestimated or undetected due to their reactivity with the stationary phase of non-polar and mid-polar columns, as well as other GC components. They also highlight the importance of selecting pyrolysis conditions for investigating the primary decomposition of PET under an isothermal kinetically limited regime.

aromatic compounds↗

Node Distortion as a Tunable Mechanism for Negative Thermal Expansion in Metal–Organic Frameworks

Chemically functionalized series of metal–organic frameworks (MOFs), with subtle differences in local structure but divergent properties, provide a valuable opportunity to explore how local chemistry can be coupled to long-range structure and functionality. Using in situ synchrotron X-ray total scattering, with powder diffraction and pair distribution function (PDF) analysis, we investigate the temperature dependence of the local- and long-range structure of MOFs based on NU-1000, in which Zr 6 O 8 nodes are coordinated by different capping ligands (H 2 O/OH, Cl – ions, formate, acetylacetonate, and hexafluoroacetylacetonate). We show that the local distortion of the Zr 6 nodes depends on the lability of the ligand and contributes to a negative thermal expansion (NTE) of the extended framework. Using multivariate data analyses, involving non-negative matrix factorization (NMF), we demonstrate a new mechanism for NTE: progressive increase in the population of a smaller, distorted node state with increasing temperature leads to global contraction of the framework. The transformation between discrete node states is noncooperative and not ordered within the lattice, i.e., a solid solution of regular and distorted nodes. Density functional theory calculations show that removal of ligands from the node can lead to distortions consistent with the Zr···Zr distances observed in the experiment PDF data. Control of the node distortion imparted by the nonlinker ligand in turn controls the NTE behavior. Furthermore, these results reveal a mechanism to control the dynamic structure of MOFs based on local chemistry.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Inferring Plant Acclimation and Improving Model Generalizability With Differentiable Physics‐Informed Machine Learning of Photosynthesis

Net photosynthesis (A N ) is a key component of the global carbon cycle influencing climate feedback over decadal scales. Although plant acclimation to environmental changes can modify A N , traditional vegetation models in Earth system models (ESMs) often rely on plant functional type (PFT)-specific parameterizations or simplified acclimation assumptions limiting generalizability across time, space, and PFTs. In this study, we developed a differentiable photosynthesis model to learn the environmental dependencies of V c,max25 (maximum carboxylation rate at 25°C, representing photosynthetic capacity), as this genre of hybrid physics-informed machine learning can seamlessly train neural networks and process-based equations together. Compared to PFT-specific parameterization of V c,max25 , learning the environment dependencies of key photosynthetic parameters improved model spatiotemporal generalizability. Applying environmental acclimation to V c,max25 led to substantial variations in global mean A N indicating the need to address acclimation in ESMs. The model effectively captured multivariate observations (V c,max25 , A N , and stomatal conductance (g s )) simultaneously with multivariate constraints, improving generalization across space and PFTs. It also learned sensible acclimation relationships of V c,max25 to different environmental conditions. The model explained more than 54%, 57%, and 62% of the variance of A N , g s , and V c,max25 , respectively, presenting a first global-scale spatial test benchmark of A N and g s . These results highlight the potential for differentiable modeling to enhance process-based modules in ESMs and effectively leverage information from large, multivariate data sets.

54 ENVIRONMENTAL SCIENCES↗

Causal interaction in high frequency turbulence at the biosphere–atmosphere interface: Structural behavior

High-frequency (e.g., 10 Hz) eddy covariance measurements are typically used to estimate fluxes at the land–atmosphere interface at timescales of 15–60 min. These multivariate data contain information about the interdependency at high frequency between the interacting variables such as wind, humidity, temperature, and CO 2⁠ . We use data at 10 Hz from an eddy covariance instrument located at 25 m above agricultural land in the Midwestern US, which offers an opportunity to move beyond the traditional spectral analyses to explore causal dependency among variables. In this study, we quantify the structure of inter-dependencies of interacting variables at high frequency represented by a directed acyclic graph (DAG). We compare DAGs to investigate changes in structural differences in causal interactions. We then apply a distance-based classification and -means clustering approach to identify the evolution of the causal structure represented by a DAG. Our method selects an unbiased number of clusters of similar structures and characterizes the similarities and differences between them. We explore a range of dynamic behavior using data from a clear sky day and during a solar eclipse in 2017. Our results show well-defined clusters of similar causal dependencies as the system evolves. Furthermore, our approach provides a methodological framework to understand how causal dependence in turbulence manifests in high-frequency data when represented through a DAG.

54 ENVIRONMENTAL SCIENCES↗

Causal interaction in high frequency turbulence at the biosphere–atmosphere interface: Structure–function coupling

At the biosphere–atmosphere interface, nonlinear interdependencies among components of an ecohydrological complex system can be inferred using multivariate high frequency time series observations. Information flow among these interacting variables allows us to represent the causal dependencies in the form of a directed acyclic graph (DAG). Here, we use high frequency multivariate data at 10 Hz from an eddy covariance instrument located at 25 m above agricultural land in the Midwestern US to quantify the evolutionary dynamics of this complex system using a sequence of DAGs by examining the structural dependency of information flow and the associated functional response. We investigate whether functional differences correspond to structural differences or if there are no functional variations despite the structural differences. We base our analysis on the hypothesis that causal dependencies are instigated through information flow, and the resulting interactions sustain the dynamics and its functionality. To test our hypothesis, we build upon causal structure analysis in the companion paper to characterize the information flow in similarly clustered DAGs from 3-min non-overlapping contiguous windows in the observational data. We characterize functionality as the nature of interactions as discerned through redundant, unique, and synergistic components of information flow. Through this analysis, we find that in turbulence at the biosphere–atmosphere interface, the variables that control the dynamic character of the atmosphere as well as the thermodynamics are driven by non-local conditions, while the scalar transport associated with CO and H 2 O is mainly driven by short-term local conditions.

58 GEOSCIENCES↗

Machine-Learning Assisted Identification of Accurate Battery Lifetime Models with Uncertainty

Reduced-order battery lifetime models, which consist of algebraic expressions for various aging modes, are widely utilized for extrapolating degradation trends from accelerated aging tests to real-world aging scenarios. Identifying models with high accuracy and low uncertainty is crucial for ensuring that model extrapolations are believable, however, it is difficult to compose expressions that accurately predict multivariate data trends; a review of cycling degradation models from literature reveals a wide variety of functional relationships. Here, a machine-learning assisted model identification method is utilized to fit degradation in a stand-out LFP-Gr aging data set, with uncertainty quantified by bootstrap resampling. The model identified in this work results in approximately half the mean absolute error of a human expert model. Models are validated by converting to a state-equation form and comparing predictions against cells aging under varying loads. Parameter uncertainty is carried forward into an energy storage system simulation to estimate the impact of aging model uncertainty on system lifetime. The new model identification method used here reduces life-prediction uncertainty by more than a factor of three (86% ± 5% relative capacity at 10 years for human-expert model, 88.5% ± 1.5% for machine-learning assisted model), empowering more confident estimates of energy storage system lifetime.

25 ENERGY STORAGE↗