Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “multivariate data analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Evaluation of COTS Electronics by Power Spectrum Analysis and Multivariate Data Analysis

Power spectrum analysis (PSA) is a fast, non-destructive, sensitive method for examining commercial off-the-shelf ( COTS ) electronic components. These features make PSA attractive for both component screening and surveillance in support of component reliability efforts. Current analysis methods limit the utility of PSA due to the need to manually examine the results of analysis to identify anomalous parts. This study demonstrates the development and application of a workflow to automate the screening of COTS electronic components. Further, this study demonstrates the use of multivariate algorithms to assess aging of Zener diodes. These workflows can be readily extended to other components, combining the benefits of PSA and multivariate analysis to screen and evaluate COTS electronic components.

42 ENGINEERING↗

Evaluation of Meterorite Amono Acid Analysis Data Using Multivariate Techniques

The amino acid distributions in the Murchison carbonaceous chondrite, Mars meteorite ALH84001, and ice from the Allan Hills region of Antarctica are shown, using a multivariate technique known as Principal Component Analysis (PCA), to be statistically distinct from the average amino acid compostion of 101 terrestrial protein superfamilies.

meteorites Amino acids Multivariate Analysis↗

Data Mining of Polymer Phase Transitions upon Temperature Changes by Small and Wide-Angle X-ray Scattering Combined with Raman Spectroscopy

The complex physical transformations of polymers upon external thermodynamic changes are related to the molecular length of the polymer and its associated multifaceted energetic balance. The understanding of subtle transitions or multistep phase transformation requires real-time phenomenological studies using a multi-technique approach that covers several length-scales and chemical states. A combination of X-ray scattering techniques with Raman spectroscopy and Differential Scanning Calorimetry was conducted to correlate the structural changes from the conformational chain to the polymer crystal and mesoscale organization. Current research applications and the experimental combination of Raman spectroscopy with simultaneous SAXS/WAXS measurements coupled to a DSC is discussed. In particular, we show that in order to obtain the maximum benefit from simultaneously obtained high-quality data sets from different techniques, one should look beyond traditional analysis techniques and instead apply multivariate analysis. Data mining strategies can be applied to develop methods to control polymer processing in an industrial context. Crystallization studies of a PVDF blend with a fluoroelastomer, known to feature complex phase transitions, were used to validate the combined approach and further analyzed by MVA.

36 MATERIALS SCIENCE↗

Salvaging Data Records with Missing Data: Data Imputation using the Multivariate t Distribution

When doing multivariate data analysis, one commonobstacle is the presence of incomplete observations, i.e., observationsfor which one or more key fields are blank. Missing datais often countered by deleting entire observations that containmissing data. The negative effects of deleting entire observationsare multiple: deleting observations reduces sample size andcan also result in biased inferences even if data is missing atrandom. In addition, knowledge contained within incompleteobservations is knowledge lost when they are deleted– and theeffort spent collecting that knowledge is effort wasted. Data imputationmethods, or methods of statistically “filling-in” missingdata, can help combat small sample sizes by using the existinginformation in partially complete observations with the end goalof producing less biased and higher confidence inferences. Whena sample from a multivariate normal population is only partiallycomplete, and the missing data meets appropriate assumptions(missing at random), robust data imputation of the missing datacan be implemented with monotone data augmentation (MDA)using the multivariate t distribution.Missing data imputation is applied to data from the NASA InstrumentCost Model (NICM) using the MDA algorithm underthe assumption of having a multivariate t distribution with fixeddegrees of freedom. A sensitivity analysis to the degrees offreedom parameter is presented to demonstrate robustness ofthe multivariate t distribution when dealing with small samplesas compared to the multivariate normal distribution.

DiNicola, Michael↗

Intrinsic Kinetics of Polyethylene Terephthalate Pyrolysis via Micropyrolysis and Multivariate Chromatographic Analysis

This study provides an in-depth investigation of the primary decomposition of polyethylene terephthalate (PET) via pyrolysis, employing an experimental-analytic workflow that integrates design of experiments (DoE), micropyrolysis coupled with comprehensive two-dimensional gas chromatography (GC×GC), and multivariate data analysis to verify intrinsic kinetic conditions and elucidate evolving product distributions for mapping key reaction pathways. Peaks that could not be identified using commercial spectral libraries were assigned using Mass Frontier simulations, enabling the identification of divinyl terephthalate, ethyl vinyl terephthalate, and 2-(benzoyloxy)ethyl vinyl terephthalate. A polar×polar (non-orthogonal) column set tailored for the detection of carboxylic acids enhanced the quantification of benzoic acid, 4-vinylbenzoic acid, 4-ethylbenzoic acid, and methylbenzoic acid by up to 6-fold relative to an orthogonal column combination (non-polar×mid-polar). Moreover, pyrolysis variables were systematically evaluated using a Box- Behnken design (BBD), encompassing pyrolysis temperature (500−600 °C), sample weight (50−150 μg), and carrier gas flow rate (100−300 mL min −1 ). Among these, pyrolysis temperature was the only statistically significant factor influencing product yields, ranging from 58.78 to 84.26 wt %. In contrast, neither the sample weight nor the carrier gas flow rate had a significant effect on product yields within the evaluated experimental space. At 600 °C, the major pyrolysis products were benzoic acid (up to 20.20 ± 1.46 wt %) and CO 2 (up to 21.28 ± 1.46 wt %), which can be produced through decarboxylation reactions. These findings underscore the critical importance of selecting appropriate analytical columns for the accurate quantification of heteroatomcontaining products such as carboxylic acids, which may otherwise be underestimated or undetected due to their reactivity with the stationary phase of non-polar and mid-polar columns, as well as other GC components. They also highlight the importance of selecting pyrolysis conditions for investigating the primary decomposition of PET under an isothermal kinetically limited regime.

aromatic compounds↗

Accurate and Rapid Forecasts for Geologic Carbon Storage via Learning-Based Inversion-Free Prediction

Carbon capture and storage (CCS) is one approach being studied by the U.S. Department of Energy to help mitigate global warming. The process involves capturing CO 2 emissions from industrial sources and permanently storing them in deep geologic formations (storage reservoirs). However, CCS projects generally target “green field sites,” where there is often little characterization data and therefore large uncertainty about the petrophysical properties and other geologic attributes of the storage reservoir. Consequently, ensemble-based approaches are often used to forecast multiple realizations prior to CO 2 injection to visualize a range of potential outcomes. In addition, monitoring data during injection operations are used to update the pre-injection forecasts and thereby improve agreement between forecasted and observed behavior. Thus, a system for generating accurate, timely forecasts of pressure buildup and CO 2 movement and distribution within the storage reservoir and for updating those forecasts via monitoring measurements becomes crucial. This study proposes a learning-based prediction method that can accurately and rapidly forecast spatial distribution of CO 2 concentration and pressure with uncertainty quantification without relying on traditional inverse modeling. The machine learning techniques include dimension reduction, multivariate data analysis, and Bayesian learning. The outcome is expected to provide CO 2 storage site operators with an effective tool for timely and informative decision making based on limited simulation and monitoring data.

58 GEOSCIENCES↗

Cloudy-sky contributions to the direct aerosol effect

The radiative forcing of the aerosol–radiation interaction can be decomposed into clear-sky and cloudy-sky portions. Two sets of multi-model simulations within Aerosol Comparisons between Observations and Models (AeroCom), combined with observational methods, and the time evolution of aerosol emissions over the industrial era show that the contribution from cloudy-sky regions is likely weak. A mean of the simulations considered is 0.01±0.1 W m-2. Multivariate data analysis of results from AeroCom Phase II shows that many factors influence the strength of the cloudy-sky contribution to the forcing of the aerosol–radiation interaction. Overall, single-scattering albedo of anthropogenic aerosols and the interaction of aerosols with the short-wave cloud radiative effects are found to be important factors. A more dedicated focus on the contribution from the cloud-free and cloud-covered sky fraction, respectively, to the aerosol–radiation interaction will benefit the quantification of the radiative forcing and its uncertainty range.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Computer program documentation for the patch subsampling processor

The programs presented are intended to provide a way to extract a sample from a full-frame scene and summarize it in a useful way. The sample in each case was chosen to fill a 512-by-512 pixel (sample-by-line) image since this is the largest image that can be displayed on the Integrated Multivariant Data Analysis and Classification System. This sample size provides one megabyte of data for manipulation and storage and contains about 3% of the full-frame data. A patch image processor computes means for 256 32-by-32 pixel squares which constitute the 512-by-512 pixel image. Thus, 256 measurements are available for 8 vegetation indexes over a 100-mile square.

Nieves, M. J.↗

SIMBAD quality-control

The astronomical database SIMBAD developed at the Centre de donnees astronomiques de Strasbourg presently contains 760,000 objects (stellar and non-stellar). It has the unique characteristic of being structured specifically for astronomical objects. All types of heterogeneous data (bibliographic references, measurements, and sets of identification) are connected with each object. The attributes that define quality of the database include the following. Reliability: cross-identification should not rely upon just exact values object coordinates. It also means that information attached to one simple object should be consistent. The existing data must be controlled in order to start with a reliable base and to cross-identify new data assuring the quality as data grows. Exhaustivity: delays between publication of new informations and their inclusion in the database should be as short as possible. The integrity of the database has to be maintained as data accumulates. Taking the amount of data into consideration and the rate of new data production, it is necessary to use automatic methods. One of the possibilities is to use multivariate data analysis. The factor-space is a n-dimensional relevancy space which is described by the n-axes representing a set of n subject matter headings; the words and phrases can be used to scale the axes and the documents are then a vector average of the terms within them. The application reported herein is based on the NASA-STI bibliographical database. The selected data concern astronomy, astrophysics, and space radiation (102,963 references from 1975 to 1991 included 8070 keywords). The F-space is built from this bibliographical data. By comparing the F-space position obtained from the NASA-STI keywords with the F-space position obtained from the SIMBAD references, the authors will be able to show whether it is possible to retrieve information with a restricted set of words only. If the comparison is valid, this will be a way to enter bibliographic information in the SIMBAD quality control process. Furthermore, it is possible to connect the physical measurements of stars from SIMBAD to literature concerning these stars from the NASA-STI abstracts. The physical properties of stars (e.g. UBV colors) are not randomly distributed. Stars are distributed among different clusters in a physical parameter space. The authors will show that there are some relations between this classification and the literature concerning these objects clusters in a factor space. They will investigate the nature of the relationship between the SIMBAD measurements and the bibliography. These would be new relationships that are not pre-established by an astronomer. In addition, the bibliography could be neutral information that can be used in combination with the measured parameters.

Lesteven, Soizick↗

Analysis of Multivariate Experimental Data Using A Simplified Regression Model Search Algorithm

A new regression model search algorithm was developed in 2011 that may be used to analyze both general multivariate experimental data sets and wind tunnel strain-gage balance calibration data. The new algorithm is a simplified version of a more complex search algorithm that was originally developed at the NASA Ames Balance Calibration Laboratory. The new algorithm has the advantage that it needs only about one tenth of the original algorithm's CPU time for the completion of a search. In addition, extensive testing showed that the prediction accuracy of math models obtained from the simplified algorithm is similar to the prediction accuracy of math models obtained from the original algorithm. The simplified algorithm, however, cannot guarantee that search constraints related to a set of statistical quality requirements are always satisfied in the optimized regression models. Therefore, the simplified search algorithm is not intended to replace the original search algorithm. Instead, it may be used to generate an alternate optimized regression model of experimental data whenever the application of the original search algorithm either fails or requires too much CPU time. Data from a machine calibration of NASA's MK40 force balance is used to illustrate the application of the new regression model search algorithm.

multivariate experimental data↗

Accurate and Timely Forecasts of Geologic Carbon Storage using Machine Learning Methods

Carbon capture and storage is one strategy to reduce greenhouse gas emissions. One approach to storing the captured CO2 is to inject it into deep saline aquifers. However, dynamics of the injected CO2 plume is uncertain and the potential for leakage back to the atmosphere must be assessed. Thus, accurate and timely forecasts of CO2 storage via real-time measurements integration becomes very crucial. This study proposes a learning-based, inverse-free prediction method that can accurately and rapidly forecast CO2 movement and distribution with uncertainty quantification based on limited simulation and observation data. The machine learning techniques include dimension reduction, multivariate data analysis, and Bayesian learning. The outcome is expected to provide CO2 storage site operators with an effective tool for real-time decision making.

Lu, Dan↗

Analysis/forecast experiments with a multivariate statistical analysis scheme using FGGE data

A three-dimensional, multivariate, statistical analysis method, optimal interpolation (OI) is described for modeling meteorological data from widely dispersed sites. The model was developed to analyze FGGE data at the NASA-Goddard Laboratory of Atmospherics. The model features a multivariate surface analysis over the oceans, including maintenance of the Ekman balance and a geographically dependent correlation function. Preliminary comparisons are made between the OI model and similar schemes employed at the European Center for Medium Range Weather Forecasts and the National Meteorological Center. The OI scheme is used to provide input to a GCM, and model error correlations are calculated for forecasts of 500 mb vertical water mixing ratios and the wind profiles. Comparisons are made between the predictions and measured data. The model is shown to be as accurate as a successive corrections model out to 4.5 days.

Baker, W. E.↗

From multivariate to functional data analysis: Fundamentals, recent developments, and emerging areas

Functional data analysis (FDA), which is a branch of statistics on modeling infinite dimensional random vectors resided in functional spaces, has become a major research area for Journal of Multivariate Analysis. We review some fundamental concepts of FDA, their origins and connections from multivariate analysis, and some of its recent developments, including multi-level functional data analysis, high-dimensional functional regression, and dependent functional data analysis. Here, we also discuss the impact of these new methodology developments on genetics, plant science, wearable device data analysis, image data analysis, and business analytics. Two real data examples are provided to motivate our discussions.

97 MATHEMATICS AND COMPUTING↗

Analysis of Multivariate Experimental Data Using A Simplified Regression Model Search Algorithm

A new regression model search algorithm was developed that may be applied to both general multivariate experimental data sets and wind tunnel strain-gage balance calibration data. The algorithm is a simplified version of a more complex algorithm that was originally developed for the NASA Ames Balance Calibration Laboratory. The new algorithm performs regression model term reduction to prevent overfitting of data. It has the advantage that it needs only about one tenth of the original algorithm's CPU time for the completion of a regression model search. In addition, extensive testing showed that the prediction accuracy of math models obtained from the simplified algorithm is similar to the prediction accuracy of math models obtained from the original algorithm. The simplified algorithm, however, cannot guarantee that search constraints related to a set of statistical quality requirements are always satisfied in the optimized regression model. Therefore, the simplified algorithm is not intended to replace the original algorithm. Instead, it may be used to generate an alternate optimized regression model of experimental data whenever the application of the original search algorithm fails or requires too much CPU time. Data from a machine calibration of NASA's MK40 force balance is used to illustrate the application of the new search algorithm.

Ulbrich, Norbert M.↗

GMT: A deep learning approach to generalized multivariate translation for scientific data analysis and visualization

In scientific visualization, despite the significant advances of deep learning for data generation, researchers have not thoroughly investigated the issue of data translation. We present a new deep learning approach called generalized multivariate translation (GMT) for multivariate time-varying data analysis and visualization. Like V2V, GMT assumes a preprocessing step that selects suitable variables for translation. However, unlike V2V, which only handles one-to-one variable translation during training and inference, GMT enables one-to-many and many-to-many variable translation in the same framework. We leverage the recent StarGAN design from multi-domain image-to-image translation to achieve this generalization capability. We experiment with different loss functions and injection strategies to explore the best choices and leverage pre-training for performance improvement. We compare GMT with other state-of-the-art methods (i.e., Pix2Pix, V2V, StarGAN). Furthermore, the results demonstrate the overall advantage of GMT in translation quality and generalization ability.

97 MATHEMATICS AND COMPUTING↗

Fiber Uncertainty Visualization for Bivariate Data With Parametric and Nonparametric Noise Models

Visualization and analysis of multivariate data and their uncertainty are top research challenges in data visualization. Constructing fiber surfaces is a popular technique for multivariate data visualization that generalizes the idea of level-set visualization for univariate data to multivariate data. Here, in this paper, we present a statistical framework to quantify positional probabilities of fibers extracted from uncertain bivariate fields. Specifically, we extend the state-of-the-art Gaussian models of uncertainty for bivariate data to other parametric distributions (e.g., uniform and Epanechnikov) and more general nonparametric probability distributions (e.g., histograms and kernel density estimation) and derive corresponding spatial probabilities of fibers. In our proposed framework, we leverage Green's theorem for closed-form computation of fiber probabilities when bivariate data are assumed to have independent parametric and nonparametric noise. Additionally, we present a nonparametric approach combined with numerical integration to study the positional probability of fibers when bivariate data are assumed to have correlated noise. For uncertainty analysis, we visualize the derived probability volumes for fibers via volume rendering and extracting level sets based on probability thresholds. We present the utility of our proposed techniques via experiments on synthetic and simulation datasets.

97 MATHEMATICS AND COMPUTING↗

A method of using cluster analysis to study statistical dependence in multivariate data

A technique is presented that uses both cluster analysis and a Monte Carlo significance test of clusters to discover associations between variables in multidimensional data. The method is applied to an example of a noisy function in three-dimensional space, to a sample from a mixture of three bivariate normal distributions, and to the well-known Fisher's Iris data.

Borucki, W. J.↗