Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “multivariate data analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

A method of using cluster analysis to study statistical dependence in multivariate data

A technique is presented that uses both cluster analysis and a Monte Carlo significance test of clusters to discover associations between variables in multidimensional data. The method is applied to an example of a noisy function in three-dimensional space, to a sample from a mixture of three bivariate normal distributions, and to the well-known Fisher's Iris data.

Borucki, W. J.↗

A Bayesian nonparametric analysis for zero-inflated multivariate count data with application to microbiome study

High-throughput sequencing technology has enabled researchers to profile microbial communities from a variety of environments, but analysis of multivariate taxon count data remains challenging. Here, we develop a Bayesian nonparametric (BNP) regression model with zero inflation to analyse multivariate count data from microbiome studies. A BNP approach flexibly models microbial associations with covariates, such as environmental factors and clinical characteristics. The model produces estimates for probability distributions which relate microbial diversity and differential abundance to covariates, and facilitates community comparisons beyond those provided by simple statistical tests. We compare the model to simpler models and popular alternatives in simulation studies, showing, in addition to these additional community-level insights, it yields superior parameter estimates and model fit in various settings. The model's utility is demonstrated by applying it to a chronic wound microbiome data set and a Human Microbiome Project data set, where it is used to compare microbial communities present in different environments.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Gaussian process analysis of electron energy loss spectroscopy data: multivariate reconstruction and kernel control

Abstract Advances in hyperspectral imaging including electron energy loss spectroscopy bring forth the challenges of exploratory and physics-based analysis of multidimensional data sets. The multivariate linear unmixing methods generally explore similarities in the energy dimension, but ignore correlations in the spatial domain. At the same time, Gaussian process (GP) explicitly incorporate spatial correlations in the form of kernel functions but is computationally intensive. Here, we implement a GP method operating on the full spatial domain and reduced representations in the energy domain. In this multivariate GP, the information between the components is shared via a common spatial kernel structure, while allowing for variability in the relative noise magnitude or image morphology. We explore the role of kernel constraints on the quality of the reconstruction, and suggest an approach for estimating them from the experimental data. We further show that spatial information contained in higher-order components can be reconstructed and spatially localized.

36 MATERIALS SCIENCE↗

Logistic Risk Model for the Unique Effects of Inherent Aerobic Capacity on (+)G(sub z) Tolerance Before and After Simulated Weightlessness

Small sample size (n less than 1O) and inappropriate analysis of multivariate data have hindered previous attempts to describe which physiologic and demographic variables are most important in determining how long humans can tolerate acceleration. Data from previous centrifuge studies conducted at NASA/Ames Research Center, utilizing a 7-14 d bed rest protocol to simulate weightlessness, were included in the current investigation. After review, data on 25 women and 22 men were available for analysis. Study variables included gender, age, weight, height, percent body fat, resting heart rate, mean arterial pressure, Vo(sub 2)max and plasma volume. Since the dependent variable was time to greyout (failure), two contemporary biostatistical modeling procedures (proportional hazard and logistic discriminant function) were used to estimate risk, given a particular subject's profile. After adjusting for pro-bed-rest tolerance time, none of the profile variables remained in the risk equation for post-bed-rest tolerance greyout. However, prior to bed rest, risk of greyout could be predicted with 91% accuracy. All of the profile variables except weight, MAP, and those related to inherent aerobic capacity (Vo(sub 2)max, percent body fat, resting heart rate) entered the risk equation for pro-bed-rest greyout. A cross-validation using 24 new subjects indicated a very stable model for risk prediction, accurate within 5% of the original equation. The result for the inherent fitness variables is significant in that a consensus as to whether an increased aerobic capacity is beneficial or detrimental has not been satisfactorily established. We conclude that tolerance to +Gz acceleration before and after simulated weightlessness is independent of inherent aerobic fitness.

Ludwig, David A.↗

A CLIPS expert system for clinical flow cytometry data analysis

An expert system is being developed using CLIPS to assist clinicians in the analysis of multivariate flow cytometry data from cancer patients. Cluster analysis is used to find subpopulations representing various cell types in multiple datasets each consisting of four to five measurements on each of 5000 cells. CLIPS facts are derived from results of the clustering. CLIPS rules are based on the expertise of Drs. Stewart, Duque, and Braylan. The rules incorporate certainty factors based on case histories.

Salzman, G. C.↗

Practical Guide to Chemometric Analysis of Optical Spectroscopic Data

The methodology and mathematical treatment of several classic multivariate methods for the analysis of spectroscopic data is demonstrated in a straightforward way that can be used as a basis for teaching an undergraduate introductory course on chemometric analysis. The multivariate techniques of classical least squares (CLS), principal component regression (PCR), and partial least squares (PLS), as well as the univariate Beer’s law method have been described and compared, building students’ understanding by starting with the univariate method and progressing step by step into the multivariate methods. Equations for the production of regression vectors from training set spectral data is described and their use demonstrated for the prediction of constituent concentrations on a separate validation set of spectra. Extreme care is taken to ensure consistency in variable formatting of data matrices. This provides a key foundation to understanding how spectral data are manipulated using these different mathematical approaches for building quantitative regression models. Each method is applied to a real-world data set, and the results are discussed to show students the types of information that can be gleaned from each method. A training set comprised of 20 infrared absorbance spectra containing 3 constituents (benzene, polystyrene, and gasoline) of known composition are used to demonstrate the matrix operations for each regression method. A separate set of 12 real-world napalm samples (containing benzene, polystyrene and gasoline) are used as a validation set to demonstrate the ability to utilize the regression models on an unknown dataset. A toolbox (PNNL Chemometric Toolbox) written in MATLAB language is supplied in the Supplemental Information file and can be used as a companion for understanding the development and deployment of the chemometric algorithms described in this paper. The datasets of the infrared spectra are also supplied, allowing users to build and inspect the chemometric models on their own. Finally, the Toolbox includes scripts to assist users in loading their own datasets into MATLAB and performing CLS, PCR, and PLS on their data.

Upper-Division Undergraduate, Analytical Chemistry↗

Materials surface contamination analysis

The original research objective was to demonstrate the ability of optical fiber spectrometry to determine contamination levels on solid rocket motor cases in order to identify surface conditions which may result in poor bonds during production. The capability of using the spectral features to identify contaminants with other sensors which might only indicate a potential contamination level provides a real enhancement to current inspection systems such as Optical Stimulated Electron Emission (OSEE). The optical fiber probe can easily fit into the same scanning fixtures as the OSEE. The initial data obtained using the Guided Wave Model 260 spectrophotometer was primarily focused on determining spectra of potential contaminants such as HD2 grease, silicones, etc. However, once we began taking data and applying multivariate analysis techniques, using a program that can handle very large data sets, i.e., Unscrambler 2, it became apparent that the techniques also might provide a nice scientific tool for determining oxidation and chemisorption rates under controlled conditions. As the ultimate power of the technique became recognized, considering that the chemical system which was most frequently studied in this work is water + D6AC steel, we became very interested in trying the spectroscopic techniques to solve a broad range of problems. The complexity of the observed spectra for the D6AC + water system is due to overlaps between the water peaks, the resulting chemisorbed species, and products of reaction which also contain OH stretching bands. Unscrambling these spectral features, without knowledge of the specific species involved, has proven to be a formidable task.

Workman, Gary L.↗

Experiments with a three-dimensional statistical objective analysis scheme using FGGE data

A three-dimensional (3D), multivariate, statistical objective analysis scheme (referred to as optimum interpolation or OI) has been developed for use in numerical weather prediction studies with the FGGE data. Some novel aspects of the present scheme include: (1) a multivariate surface analysis over the oceans, which employs an Ekman balance instead of the usual geostrophic relationship, to model the pressure-wind error cross correlations, and (2) the capability to use an error correlation function which is geographically dependent. A series of 4-day data assimilation experiments are conducted to examine the importance of some of the key features of the OI in terms of their effects on forecast skill, as well as to compare the forecast skill using the OI with that utilizing a successive correction method (SCM) of analysis developed earlier. For the three cases examined, the forecast skill is found to be rather insensitive to varying the error correlation function geographically. However, significant differences are noted between forecasts from a two-dimensional (2D) version of the OI and those from the 3D OI, with the 3D OI forecasts exhibiting better forecast skill. The 3D OI forecasts are also more accurate than those from the SCM initial conditions. The 3D OI with the multivariate oceanic surface analysis was found to produce forecasts which were slightly more accurate, on the average, than a univariate version.

Baker, Wayman E.↗

Interfaces between statistical analysis packages and the ESRI geographic information system

Interfaces between ESRI's geographic information system (GIS) data files and real valued data files written to facilitate statistical analysis and display of spatially referenced multivariable data are described. An example of data analysis which utilized the GIS and the statistical analysis system is presented to illustrate the utility of combining the analytic capability of a statistical package with the data management and display features of the GIS.

Masuoka, E.↗

Multidimensional scaling informed by F -statistic: Visualizing grouped microbiome data with inference

Multidimensional scaling (MDS) is a widely used dimensionality reduction technique in microbial ecology data analysis that captures the multivariate structure of the data while preserving pairwise distances between samples. While improvements in MDS have enhanced the ability to reveal group-specific data patterns, these MDS-based methods require prior assumptions for inference, limiting their application in general microbiome analysis. Here, in this study, we introduce a new MDS-based ordination method, “F-informed MDS,” which configures the data distribution based on the F-statistic, the ratio of dispersion between groups sharing common and different characteristics. Using semisynthetic datasets, we demonstrate that the proposed method is robust to hyperparameter selection while maintaining statistical significance throughout the ordination process. Various quality metrics for evaluating dimensionality reduction confirm that F-informed MDS is comparable to state-of-the-art methods in preserving both local and global data structures. Its application to a diatom-associated bacterial community suggests the role of this new method in interpreting the community’s response to the host. Our approach offers a well-founded refinement of MDS that aligns with statistical test results, which can be beneficial for broader multidimensional data analyses in microbiology and ecology. This new visualization tool can be incorporated into standard microbiome data analyses.

Biological and medical sciences↗

TPSAS-NF1676L-36113-DND

Concerns about the effects of extreme heat and poor air quality are increasing in North America’s largest urban centers. In Philadelphia, environmental and public health groups are concerned about how these phenomena disproportionality affect marginalized communities and populations, which often have extensive impervious surfaces and little access to green space. In order to address these concerns, the Philadelphia Department of Public Health and the Office of Sustainability seek to effectively prioritize cooling initiatives to reduce urban heat and decrease air pollutants. We evaluated land surface temperature (LST) and the Normalized Difference Vegetation Index (NDVI), as a measure of overall greenness, obtained from NASA Earth observations Aqua and Terra Moderate Resolution Imaging Spectroradiometer (MODIS), and the Ecosystem Spaceborne Thermal Radiometer Experiment on Space Station (ECOSTRESS). These analyses we recombined with local tree inventory, air quality, and socioeconomic data through a multivariate analysis to identify areas where new trees or cooling adaptations are most needed. The results and data of this project can be used by our partners to inform both short-term heat relief planning and a long-term, multi-agency heat response.

Spencer Nelson↗

Rapid measurement of soluble xylo-oligomers using near-infrared spectroscopy (NIRS) and multivariate statistics: calibration model development and practical approaches to model optimization

Rapid monitoring of biomass conversion processes using techniques such as near-infrared (NIR) spectroscopy can be substantially quicker and less labor-, resource-, and energy-intensive than conventional measurement techniques such as gas or liquid chromatography (GC or LC) due to the lack of solvents and preparation methods, as well as removing the need to transfer samples to an external lab for analytical evaluation. The purpose of this study was to determine the feasibility of rapid monitoring of a biomass conversion process using NIR spectroscopy combined with multivariate statistical modeling, and to examine the impact of (1) subsetting the samples in the original dataset by process location and (2) reducing the spectral range used in the calibration model on model performance. We develop multivariate calibration models for the concentrations of soluble xylo-oligosaccharides (XOS), monomeric xylose, and total solids at multiple points in a biomass conversion process which produces and then purifies XOS compounds from sugar cane bagasse. A single model using samples from multiple locations in the process stream showed acceptable performance as measured by standard statistical measures. However, compared to the single model, we show that separate models built by segregating the calibration samples according to process location show improved performance. We also show that combining an understanding of the sample spectra with simple multivariate analysis tools can result in a calibration model with a substantially smaller spectral range that provides essentially equal performance to the full-range model. We demonstrate that real-time monitoring of soluble xylo-oligosaccharides (XOS), monomeric xylose, and total solids concentration at multiple points in a process stream using NIR spectroscopy coupled with multivariate statistics is feasible. Segregation of sample populations by process location improves model performance. Models using a reduced spectral range containing the most relevant spectral signatures show very similar performance to the full-range model, reinforcing the importance of performing robust exploratory data analysis before beginning multivariate modeling.

09 BIOMASS FUELS↗

Hierarchical analysis of spatial pattern and processes of Douglas-fir forests

There has been an increased interest in the quantification of pattern in ecological systems over the past years. This interest is motivated by the desire to construct valid models which extend across many scales. Spatial methods must quantify pattern, discriminate types of pattern, and relate hierarchical phenomena across scales. Wavelet analysis is introduced as a method to identify spatial structure in ecological transect data. The main advantage of the wavelet transform over other methods is its ability to preserve and display hierarchical information while allowing for pattern decomposition. Two applications of wavelet analysis are illustrated, as a means to: (1) quantify known spatial patterns in Douglas-fir forests at several scales, and (2) construct spatially-explicit hypotheses regarding pattern generating mechanisms. Application of the wavelet variance, derived from the wavelet transform, is developed for forest ecosystem analysis to obtain additional insight into spatially-explicit data. Specifically, the resolution capabilities of the wavelet variance are compared to the semi-variogram and Fourier power spectra for the description of spatial data using a set of one-dimensional stationary and non-stationary processes. The wavelet cross-covariance function is derived from the wavelet transform and introduced as a alternative method for the analysis of multivariate spatial data of understory vegetation and canopy in Douglas-fir forests of the western Cascades of Oregon.

Bradshaw, G. A.↗

Digital preprocessing and classification of multispectral earth observation data

The development of airborne and satellite multispectral image scanning sensors has generated wide-spread interest in application of these sensors to earth resource mapping. These point scanning sensors permit scenes to be imaged in a large number of electromagnetic energy bands between .3 and 15 micrometers. The energy sensed in each band can be used as a feature in a computer based multi-dimensional pattern recognition process to aid in interpreting the nature of elements in the scene. Images from each band can also be interpreted visually. Visual interpretation of five or ten multispectral images simultaneously becomes impractical especially as area studied increases; hence, great emphasis has been placed on machine (computer) techniques for aiding in the interpretation process. This paper describes a computer software system concept called LARSYS for analysis of multivariate image data and presents some examples of its application.

Anuta, P. E.↗

Studies in Astronomical Time Series Analysis. VI. Bayesian Block Representations

This paper addresses the problem of detecting and characterizing local variability in time series and other forms of sequential data. The goal is to identify and characterize statistically significant variations, at the same time suppressing the inevitable corrupting observational errors. We present a simple nonparametric modeling technique and an algorithm implementing it-an improved and generalized version of Bayesian Blocks [Scargle 1998]-that finds the optimal segmentation of the data in the observation interval. The structure of the algorithm allows it to be used in either a real-time trigger mode, or a retrospective mode. Maximum likelihood or marginal posterior functions to measure model fitness are presented for events, binned counts, and measurements at arbitrary times with known error distributions. Problems addressed include those connected with data gaps, variable exposure, extension to piece- wise linear and piecewise exponential representations, multivariate time series data, analysis of variance, data on the circle, other data modes, and dispersed data. Simulations provide evidence that the detection efficiency for weak signals is close to a theoretical asymptotic limit derived by [Arias-Castro, Donoho and Huo 2003]. In the spirit of Reproducible Research [Donoho et al. (2008)] all of the code and data necessary to reproduce all of the figures in this paper are included as auxiliary material.

signal detection↗

A multiparametric analysis of the Einstein sample of early-type galaxies. 1: Luminosity and ISM parameters

We have conducted bivariate and multivariate statistical analysis of data measuring the luminosity and interstellar medium of the Einstein sample of early-type galaxies (presented by Fabbiano, Kim, & Trinchieri 1992). We find a strong nonlinear correlation between L(sub B) and L(sub X), with a power-law slope of 1.8 +/- 0.1, steepening to 2.0 +/- if we do not consider the Local Group dwarf galaxies M32 and NGC 205. Considering only galaxies with log L(sub X) less than or equal to 40.5, we instead find a slope of 1.0 +/- 0.2 (with or without the Local Group dwarfs). Although E and S0 galaxies have consistent slopes for their L(sub B)-L(sub X) relationships, the mean values of the distribution functions of both L(sub X) and L(sub X)/L(sub B) for the S0 galaxies are lower than those for the E galaxies at the 2.8 sigma and 3.5 sigma levels, respectively. We find clear evidence for a correlation between L(sub X) and the X-ray color C(sub 21), defined by Kim, Fabbiano, & Trinchieri (1992b), which indicates that X-ray luminosity is correlated with the spectral shape below 1 keV in the sense that low-L(sub X) systems have relatively large contributions from a soft component compared with high-L(sub X) systems. We find evidence from our analysis of the 12 micron IRAS data for our sample that our S0 sample has excess 12 micron emission compared with the E sample, scaled by their optical luminosities. This may be due to emission from dust heated in star-forming regions in S0 disks. This interpretation is reinforced by the existence of a strong L(sub 12)-L(sub 100) correlation for our S0 sample that is not found for the E galaxies, and by an analysis of optical-IR colors. We find steep slopes for power-law relationships between radio luminosity and optical, X-ray, and far-IR (FIR) properties. This last point argues that the presence of an FIR-emitting interstellar medium (ISM) in early-type galaxies is coupled to their ability to generate nonthermal radio continuum, as previously argued by, e.g., Walsh et al. (1989). We also find that, for a given L(sub 100), galaxies with larger L(sub X)/L(sub B) tend to be stronger nonthermal radio sources, as originally suggested by Kim & Fabbiano (1990). We note that, while L(sub B) is most strongly correlated with L(sub 6), the total radio luminosity, both L(sub X) and L(sub X)/L(sub B) are more strongly correlated with L(sub 6 CO), the core radio luminosity. These points support the argument (proposed by Fabbiano, Gioia, & Trinchieri 1989) that radio cores in early-type galaxies are fueled by the hot ISM.

Eskridge, Paul B.↗

A multiparametric analysis of the Einstein sample of early-type galaxies. 2: Galaxy formation history and properties of the interstellar medium

We have conducted bivariate and multivariate statistical analysis of data measuring the integrated luminosity, shape, and potential depth of the Einstein sample of early-type galaxies (presented by Fabbiano et al. 1992). We find significant correlations between the X-ray properties and the axial ratios (a/b) of our sample, such that the roundest systems tend to have the highest L(sub x) and L(sub x)/L(sub B). The most radio-loud objects are also the roundest. We confirm the assertion of Bender et al. (1989) that galaxies with high L(sub x) are boxy (have negative a(sub 4)). Both a/b and a(sub 4) are correlated with L(sub B), but not with IRAS 12 um and 100 um luminosities. There are strong correlations between L(sub x), Mg(sub 2), and sigma(sub nu) in the sense that those systems with the deepest potential wells have the highest L(sub x) and Mg(sub 2). Thus the depth of the potential well appears to govern both the ability to reatin an ISM at the present epoch and to retain the enriched ejecta of early star formation bursts. Both L(sub x)/L(sub B) and L(sub 6) (the 6 cm radio luminosity) show threshold effects with sigma(sub nu) exhibiting sharp increases at log sigma(sub nu) approximately = 2.2. Finally, there is clearly an interrelationship between the various stellar and structural parameters: The scatter in the bivariate relationships between the shape parameters (a/b and a(sub 4)) and the depth parameter sigma(sub nu) is a function of abundance in the sense that, for a given a(sub 4) or a/b, the systems with the highest sigma(sub nu) also have the highest Mg(sub 2). Furthermore, for a constant sigma(sun nu), disky galaxies tend to have higher Mg(sub 2) than boxy ones. Alternatively, for a given abundance, boxy ellipticals tend to be more massive than disky ellipticals. One possibility is that early-type galaxies of a given mass, originating from mergers (boxy ellipticals), have lower abundances than 'primordial' (disky) early-type galaxies. Another is that disky inner isophotes are due not to primordial dissipation collapse, but to either the self-gravitating inner disks of captured spirals or the dissipational collapse of new disk structures from the premerger ISM. The high measured nuclear Mg(sub 2) values would thus be due to enrichment from secondary bursts of star formation triggered by the merging event.

Eskridge, Paul B.↗