Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “multivariate data analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Evaluation of COTS Electronics by Power Spectrum Analysis and Multivariate Data Analysis

Power spectrum analysis (PSA) is a fast, non-destructive, sensitive method for examining commercial off-the-shelf ( COTS ) electronic components. These features make PSA attractive for both component screening and surveillance in support of component reliability efforts. Current analysis methods limit the utility of PSA due to the need to manually examine the results of analysis to identify anomalous parts. This study demonstrates the development and application of a workflow to automate the screening of COTS electronic components. Further, this study demonstrates the use of multivariate algorithms to assess aging of Zener diodes. These workflows can be readily extended to other components, combining the benefits of PSA and multivariate analysis to screen and evaluate COTS electronic components.

42 ENGINEERING↗

Data Mining of Polymer Phase Transitions upon Temperature Changes by Small and Wide-Angle X-ray Scattering Combined with Raman Spectroscopy

The complex physical transformations of polymers upon external thermodynamic changes are related to the molecular length of the polymer and its associated multifaceted energetic balance. The understanding of subtle transitions or multistep phase transformation requires real-time phenomenological studies using a multi-technique approach that covers several length-scales and chemical states. A combination of X-ray scattering techniques with Raman spectroscopy and Differential Scanning Calorimetry was conducted to correlate the structural changes from the conformational chain to the polymer crystal and mesoscale organization. Current research applications and the experimental combination of Raman spectroscopy with simultaneous SAXS/WAXS measurements coupled to a DSC is discussed. In particular, we show that in order to obtain the maximum benefit from simultaneously obtained high-quality data sets from different techniques, one should look beyond traditional analysis techniques and instead apply multivariate analysis. Data mining strategies can be applied to develop methods to control polymer processing in an industrial context. Crystallization studies of a PVDF blend with a fluoroelastomer, known to feature complex phase transitions, were used to validate the combined approach and further analyzed by MVA.

36 MATERIALS SCIENCE↗

Intrinsic Kinetics of Polyethylene Terephthalate Pyrolysis via Micropyrolysis and Multivariate Chromatographic Analysis

This study provides an in-depth investigation of the primary decomposition of polyethylene terephthalate (PET) via pyrolysis, employing an experimental-analytic workflow that integrates design of experiments (DoE), micropyrolysis coupled with comprehensive two-dimensional gas chromatography (GC×GC), and multivariate data analysis to verify intrinsic kinetic conditions and elucidate evolving product distributions for mapping key reaction pathways. Peaks that could not be identified using commercial spectral libraries were assigned using Mass Frontier simulations, enabling the identification of divinyl terephthalate, ethyl vinyl terephthalate, and 2-(benzoyloxy)ethyl vinyl terephthalate. A polar×polar (non-orthogonal) column set tailored for the detection of carboxylic acids enhanced the quantification of benzoic acid, 4-vinylbenzoic acid, 4-ethylbenzoic acid, and methylbenzoic acid by up to 6-fold relative to an orthogonal column combination (non-polar×mid-polar). Moreover, pyrolysis variables were systematically evaluated using a Box- Behnken design (BBD), encompassing pyrolysis temperature (500−600 °C), sample weight (50−150 μg), and carrier gas flow rate (100−300 mL min −1 ). Among these, pyrolysis temperature was the only statistically significant factor influencing product yields, ranging from 58.78 to 84.26 wt %. In contrast, neither the sample weight nor the carrier gas flow rate had a significant effect on product yields within the evaluated experimental space. At 600 °C, the major pyrolysis products were benzoic acid (up to 20.20 ± 1.46 wt %) and CO 2 (up to 21.28 ± 1.46 wt %), which can be produced through decarboxylation reactions. These findings underscore the critical importance of selecting appropriate analytical columns for the accurate quantification of heteroatomcontaining products such as carboxylic acids, which may otherwise be underestimated or undetected due to their reactivity with the stationary phase of non-polar and mid-polar columns, as well as other GC components. They also highlight the importance of selecting pyrolysis conditions for investigating the primary decomposition of PET under an isothermal kinetically limited regime.

aromatic compounds↗

Accurate and Rapid Forecasts for Geologic Carbon Storage via Learning-Based Inversion-Free Prediction

Carbon capture and storage (CCS) is one approach being studied by the U.S. Department of Energy to help mitigate global warming. The process involves capturing CO 2 emissions from industrial sources and permanently storing them in deep geologic formations (storage reservoirs). However, CCS projects generally target “green field sites,” where there is often little characterization data and therefore large uncertainty about the petrophysical properties and other geologic attributes of the storage reservoir. Consequently, ensemble-based approaches are often used to forecast multiple realizations prior to CO 2 injection to visualize a range of potential outcomes. In addition, monitoring data during injection operations are used to update the pre-injection forecasts and thereby improve agreement between forecasted and observed behavior. Thus, a system for generating accurate, timely forecasts of pressure buildup and CO 2 movement and distribution within the storage reservoir and for updating those forecasts via monitoring measurements becomes crucial. This study proposes a learning-based prediction method that can accurately and rapidly forecast spatial distribution of CO 2 concentration and pressure with uncertainty quantification without relying on traditional inverse modeling. The machine learning techniques include dimension reduction, multivariate data analysis, and Bayesian learning. The outcome is expected to provide CO 2 storage site operators with an effective tool for timely and informative decision making based on limited simulation and monitoring data.

58 GEOSCIENCES↗

Cloudy-sky contributions to the direct aerosol effect

The radiative forcing of the aerosol–radiation interaction can be decomposed into clear-sky and cloudy-sky portions. Two sets of multi-model simulations within Aerosol Comparisons between Observations and Models (AeroCom), combined with observational methods, and the time evolution of aerosol emissions over the industrial era show that the contribution from cloudy-sky regions is likely weak. A mean of the simulations considered is 0.01±0.1 W m-2. Multivariate data analysis of results from AeroCom Phase II shows that many factors influence the strength of the cloudy-sky contribution to the forcing of the aerosol–radiation interaction. Overall, single-scattering albedo of anthropogenic aerosols and the interaction of aerosols with the short-wave cloud radiative effects are found to be important factors. A more dedicated focus on the contribution from the cloud-free and cloud-covered sky fraction, respectively, to the aerosol–radiation interaction will benefit the quantification of the radiative forcing and its uncertainty range.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Accurate and Timely Forecasts of Geologic Carbon Storage using Machine Learning Methods

Carbon capture and storage is one strategy to reduce greenhouse gas emissions. One approach to storing the captured CO2 is to inject it into deep saline aquifers. However, dynamics of the injected CO2 plume is uncertain and the potential for leakage back to the atmosphere must be assessed. Thus, accurate and timely forecasts of CO2 storage via real-time measurements integration becomes very crucial. This study proposes a learning-based, inverse-free prediction method that can accurately and rapidly forecast CO2 movement and distribution with uncertainty quantification based on limited simulation and observation data. The machine learning techniques include dimension reduction, multivariate data analysis, and Bayesian learning. The outcome is expected to provide CO2 storage site operators with an effective tool for real-time decision making.

Lu, Dan↗

From multivariate to functional data analysis: Fundamentals, recent developments, and emerging areas

Functional data analysis (FDA), which is a branch of statistics on modeling infinite dimensional random vectors resided in functional spaces, has become a major research area for Journal of Multivariate Analysis. We review some fundamental concepts of FDA, their origins and connections from multivariate analysis, and some of its recent developments, including multi-level functional data analysis, high-dimensional functional regression, and dependent functional data analysis. Here, we also discuss the impact of these new methodology developments on genetics, plant science, wearable device data analysis, image data analysis, and business analytics. Two real data examples are provided to motivate our discussions.

97 MATHEMATICS AND COMPUTING↗

GMT: A deep learning approach to generalized multivariate translation for scientific data analysis and visualization

In scientific visualization, despite the significant advances of deep learning for data generation, researchers have not thoroughly investigated the issue of data translation. We present a new deep learning approach called generalized multivariate translation (GMT) for multivariate time-varying data analysis and visualization. Like V2V, GMT assumes a preprocessing step that selects suitable variables for translation. However, unlike V2V, which only handles one-to-one variable translation during training and inference, GMT enables one-to-many and many-to-many variable translation in the same framework. We leverage the recent StarGAN design from multi-domain image-to-image translation to achieve this generalization capability. We experiment with different loss functions and injection strategies to explore the best choices and leverage pre-training for performance improvement. We compare GMT with other state-of-the-art methods (i.e., Pix2Pix, V2V, StarGAN). Furthermore, the results demonstrate the overall advantage of GMT in translation quality and generalization ability.

97 MATHEMATICS AND COMPUTING↗

Fiber Uncertainty Visualization for Bivariate Data With Parametric and Nonparametric Noise Models

Visualization and analysis of multivariate data and their uncertainty are top research challenges in data visualization. Constructing fiber surfaces is a popular technique for multivariate data visualization that generalizes the idea of level-set visualization for univariate data to multivariate data. Here, in this paper, we present a statistical framework to quantify positional probabilities of fibers extracted from uncertain bivariate fields. Specifically, we extend the state-of-the-art Gaussian models of uncertainty for bivariate data to other parametric distributions (e.g., uniform and Epanechnikov) and more general nonparametric probability distributions (e.g., histograms and kernel density estimation) and derive corresponding spatial probabilities of fibers. In our proposed framework, we leverage Green's theorem for closed-form computation of fiber probabilities when bivariate data are assumed to have independent parametric and nonparametric noise. Additionally, we present a nonparametric approach combined with numerical integration to study the positional probability of fibers when bivariate data are assumed to have correlated noise. For uncertainty analysis, we visualize the derived probability volumes for fibers via volume rendering and extracting level sets based on probability thresholds. We present the utility of our proposed techniques via experiments on synthetic and simulation datasets.

97 MATHEMATICS AND COMPUTING↗

A Bayesian nonparametric analysis for zero-inflated multivariate count data with application to microbiome study

High-throughput sequencing technology has enabled researchers to profile microbial communities from a variety of environments, but analysis of multivariate taxon count data remains challenging. Here, we develop a Bayesian nonparametric (BNP) regression model with zero inflation to analyse multivariate count data from microbiome studies. A BNP approach flexibly models microbial associations with covariates, such as environmental factors and clinical characteristics. The model produces estimates for probability distributions which relate microbial diversity and differential abundance to covariates, and facilitates community comparisons beyond those provided by simple statistical tests. We compare the model to simpler models and popular alternatives in simulation studies, showing, in addition to these additional community-level insights, it yields superior parameter estimates and model fit in various settings. The model's utility is demonstrated by applying it to a chronic wound microbiome data set and a Human Microbiome Project data set, where it is used to compare microbial communities present in different environments.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Gaussian process analysis of electron energy loss spectroscopy data: multivariate reconstruction and kernel control

Abstract Advances in hyperspectral imaging including electron energy loss spectroscopy bring forth the challenges of exploratory and physics-based analysis of multidimensional data sets. The multivariate linear unmixing methods generally explore similarities in the energy dimension, but ignore correlations in the spatial domain. At the same time, Gaussian process (GP) explicitly incorporate spatial correlations in the form of kernel functions but is computationally intensive. Here, we implement a GP method operating on the full spatial domain and reduced representations in the energy domain. In this multivariate GP, the information between the components is shared via a common spatial kernel structure, while allowing for variability in the relative noise magnitude or image morphology. We explore the role of kernel constraints on the quality of the reconstruction, and suggest an approach for estimating them from the experimental data. We further show that spatial information contained in higher-order components can be reconstructed and spatially localized.

36 MATERIALS SCIENCE↗

Practical Guide to Chemometric Analysis of Optical Spectroscopic Data

The methodology and mathematical treatment of several classic multivariate methods for the analysis of spectroscopic data is demonstrated in a straightforward way that can be used as a basis for teaching an undergraduate introductory course on chemometric analysis. The multivariate techniques of classical least squares (CLS), principal component regression (PCR), and partial least squares (PLS), as well as the univariate Beer’s law method have been described and compared, building students’ understanding by starting with the univariate method and progressing step by step into the multivariate methods. Equations for the production of regression vectors from training set spectral data is described and their use demonstrated for the prediction of constituent concentrations on a separate validation set of spectra. Extreme care is taken to ensure consistency in variable formatting of data matrices. This provides a key foundation to understanding how spectral data are manipulated using these different mathematical approaches for building quantitative regression models. Each method is applied to a real-world data set, and the results are discussed to show students the types of information that can be gleaned from each method. A training set comprised of 20 infrared absorbance spectra containing 3 constituents (benzene, polystyrene, and gasoline) of known composition are used to demonstrate the matrix operations for each regression method. A separate set of 12 real-world napalm samples (containing benzene, polystyrene and gasoline) are used as a validation set to demonstrate the ability to utilize the regression models on an unknown dataset. A toolbox (PNNL Chemometric Toolbox) written in MATLAB language is supplied in the Supplemental Information file and can be used as a companion for understanding the development and deployment of the chemometric algorithms described in this paper. The datasets of the infrared spectra are also supplied, allowing users to build and inspect the chemometric models on their own. Finally, the Toolbox includes scripts to assist users in loading their own datasets into MATLAB and performing CLS, PCR, and PLS on their data.

Upper-Division Undergraduate, Analytical Chemistry↗

Multidimensional scaling informed by F -statistic: Visualizing grouped microbiome data with inference

Multidimensional scaling (MDS) is a widely used dimensionality reduction technique in microbial ecology data analysis that captures the multivariate structure of the data while preserving pairwise distances between samples. While improvements in MDS have enhanced the ability to reveal group-specific data patterns, these MDS-based methods require prior assumptions for inference, limiting their application in general microbiome analysis. Here, in this study, we introduce a new MDS-based ordination method, “F-informed MDS,” which configures the data distribution based on the F-statistic, the ratio of dispersion between groups sharing common and different characteristics. Using semisynthetic datasets, we demonstrate that the proposed method is robust to hyperparameter selection while maintaining statistical significance throughout the ordination process. Various quality metrics for evaluating dimensionality reduction confirm that F-informed MDS is comparable to state-of-the-art methods in preserving both local and global data structures. Its application to a diatom-associated bacterial community suggests the role of this new method in interpreting the community’s response to the host. Our approach offers a well-founded refinement of MDS that aligns with statistical test results, which can be beneficial for broader multidimensional data analyses in microbiology and ecology. This new visualization tool can be incorporated into standard microbiome data analyses.

Biological and medical sciences↗

Rapid measurement of soluble xylo-oligomers using near-infrared spectroscopy (NIRS) and multivariate statistics: calibration model development and practical approaches to model optimization

Rapid monitoring of biomass conversion processes using techniques such as near-infrared (NIR) spectroscopy can be substantially quicker and less labor-, resource-, and energy-intensive than conventional measurement techniques such as gas or liquid chromatography (GC or LC) due to the lack of solvents and preparation methods, as well as removing the need to transfer samples to an external lab for analytical evaluation. The purpose of this study was to determine the feasibility of rapid monitoring of a biomass conversion process using NIR spectroscopy combined with multivariate statistical modeling, and to examine the impact of (1) subsetting the samples in the original dataset by process location and (2) reducing the spectral range used in the calibration model on model performance. We develop multivariate calibration models for the concentrations of soluble xylo-oligosaccharides (XOS), monomeric xylose, and total solids at multiple points in a biomass conversion process which produces and then purifies XOS compounds from sugar cane bagasse. A single model using samples from multiple locations in the process stream showed acceptable performance as measured by standard statistical measures. However, compared to the single model, we show that separate models built by segregating the calibration samples according to process location show improved performance. We also show that combining an understanding of the sample spectra with simple multivariate analysis tools can result in a calibration model with a substantially smaller spectral range that provides essentially equal performance to the full-range model. We demonstrate that real-time monitoring of soluble xylo-oligosaccharides (XOS), monomeric xylose, and total solids concentration at multiple points in a process stream using NIR spectroscopy coupled with multivariate statistics is feasible. Segregation of sample populations by process location improves model performance. Models using a reduced spectral range containing the most relevant spectral signatures show very similar performance to the full-range model, reinforcing the importance of performing robust exploratory data analysis before beginning multivariate modeling.

09 BIOMASS FUELS↗

Automated Algorithms for Screening Electronic Parts for Aging using Power Spectra Analysis (PSA) Data

Understanding the age of semiconductor parts being built into devices and systems is of interest for manufacturing quality control. Power spectrum analysis (PSA) is a fast, non-destructive, sensitive method for examining semiconductor parts. This talk will cover the use of multivariate analysis on both PSA data and conventional current-voltage data generated prior to PSA analysis to create algorithms that can be automated to screen semiconductor parts for aging.

Multari, Rosalie A↗

CHMMPP: A c++ library for constrained Hidden Markov Models

SAND2024-13027O The CHMMPP: A c++ Library for Constrained Hidden Markov Models (HMM) software supports the analysis of multivariate time series data to detect patterns using HMM. Many applications involve the detection and characterization of hidden or latent states in a complex system using observable states and variables. This software supports inference of latent states integrating both an HMM and application-specific constraints that reflect known relationships in hidden states. The CHMMPP software supports application-specific and generic methods for constrained inference. This includes a framework for customized Viterbi methods, constrained inference of hidden states with A* and integer programming methods, and various constraint-informed methods for learning HMM model parameters. CHMMPP focuses on supporting generic methods that enable the agile expression of complex sets of constraints that naturally arise in many real-world applications.

Hart, William↗

mvBayesR

SAND2025-11559O The mvBayesR tool performs multivariate Bayesian analysis on generic data. It includes tools for regression modeling, diagnosis, basis decomposition, sensitivity analysis, and visualization. The tool compiles state-of-the-art methodology into one easy-to-use package. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Tucker, James [Sandia National Lab. (SNL-CA), Live↗