Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “multivariate data analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Evaluation of Meterorite Amono Acid Analysis Data Using Multivariate Techniques

The amino acid distributions in the Murchison carbonaceous chondrite, Mars meteorite ALH84001, and ice from the Allan Hills region of Antarctica are shown, using a multivariate technique known as Principal Component Analysis (PCA), to be statistically distinct from the average amino acid compostion of 101 terrestrial protein superfamilies.

meteorites Amino acids Multivariate Analysis↗

Salvaging Data Records with Missing Data: Data Imputation using the Multivariate t Distribution

When doing multivariate data analysis, one commonobstacle is the presence of incomplete observations, i.e., observationsfor which one or more key fields are blank. Missing datais often countered by deleting entire observations that containmissing data. The negative effects of deleting entire observationsare multiple: deleting observations reduces sample size andcan also result in biased inferences even if data is missing atrandom. In addition, knowledge contained within incompleteobservations is knowledge lost when they are deleted– and theeffort spent collecting that knowledge is effort wasted. Data imputationmethods, or methods of statistically “filling-in” missingdata, can help combat small sample sizes by using the existinginformation in partially complete observations with the end goalof producing less biased and higher confidence inferences. Whena sample from a multivariate normal population is only partiallycomplete, and the missing data meets appropriate assumptions(missing at random), robust data imputation of the missing datacan be implemented with monotone data augmentation (MDA)using the multivariate t distribution.Missing data imputation is applied to data from the NASA InstrumentCost Model (NICM) using the MDA algorithm underthe assumption of having a multivariate t distribution with fixeddegrees of freedom. A sensitivity analysis to the degrees offreedom parameter is presented to demonstrate robustness ofthe multivariate t distribution when dealing with small samplesas compared to the multivariate normal distribution.

DiNicola, Michael↗

Computer program documentation for the patch subsampling processor

The programs presented are intended to provide a way to extract a sample from a full-frame scene and summarize it in a useful way. The sample in each case was chosen to fill a 512-by-512 pixel (sample-by-line) image since this is the largest image that can be displayed on the Integrated Multivariant Data Analysis and Classification System. This sample size provides one megabyte of data for manipulation and storage and contains about 3% of the full-frame data. A patch image processor computes means for 256 32-by-32 pixel squares which constitute the 512-by-512 pixel image. Thus, 256 measurements are available for 8 vegetation indexes over a 100-mile square.

Nieves, M. J.↗

SIMBAD quality-control

The astronomical database SIMBAD developed at the Centre de donnees astronomiques de Strasbourg presently contains 760,000 objects (stellar and non-stellar). It has the unique characteristic of being structured specifically for astronomical objects. All types of heterogeneous data (bibliographic references, measurements, and sets of identification) are connected with each object. The attributes that define quality of the database include the following. Reliability: cross-identification should not rely upon just exact values object coordinates. It also means that information attached to one simple object should be consistent. The existing data must be controlled in order to start with a reliable base and to cross-identify new data assuring the quality as data grows. Exhaustivity: delays between publication of new informations and their inclusion in the database should be as short as possible. The integrity of the database has to be maintained as data accumulates. Taking the amount of data into consideration and the rate of new data production, it is necessary to use automatic methods. One of the possibilities is to use multivariate data analysis. The factor-space is a n-dimensional relevancy space which is described by the n-axes representing a set of n subject matter headings; the words and phrases can be used to scale the axes and the documents are then a vector average of the terms within them. The application reported herein is based on the NASA-STI bibliographical database. The selected data concern astronomy, astrophysics, and space radiation (102,963 references from 1975 to 1991 included 8070 keywords). The F-space is built from this bibliographical data. By comparing the F-space position obtained from the NASA-STI keywords with the F-space position obtained from the SIMBAD references, the authors will be able to show whether it is possible to retrieve information with a restricted set of words only. If the comparison is valid, this will be a way to enter bibliographic information in the SIMBAD quality control process. Furthermore, it is possible to connect the physical measurements of stars from SIMBAD to literature concerning these stars from the NASA-STI abstracts. The physical properties of stars (e.g. UBV colors) are not randomly distributed. Stars are distributed among different clusters in a physical parameter space. The authors will show that there are some relations between this classification and the literature concerning these objects clusters in a factor space. They will investigate the nature of the relationship between the SIMBAD measurements and the bibliography. These would be new relationships that are not pre-established by an astronomer. In addition, the bibliography could be neutral information that can be used in combination with the measured parameters.

Lesteven, Soizick↗

Analysis of Multivariate Experimental Data Using A Simplified Regression Model Search Algorithm

A new regression model search algorithm was developed in 2011 that may be used to analyze both general multivariate experimental data sets and wind tunnel strain-gage balance calibration data. The new algorithm is a simplified version of a more complex search algorithm that was originally developed at the NASA Ames Balance Calibration Laboratory. The new algorithm has the advantage that it needs only about one tenth of the original algorithm's CPU time for the completion of a search. In addition, extensive testing showed that the prediction accuracy of math models obtained from the simplified algorithm is similar to the prediction accuracy of math models obtained from the original algorithm. The simplified algorithm, however, cannot guarantee that search constraints related to a set of statistical quality requirements are always satisfied in the optimized regression models. Therefore, the simplified search algorithm is not intended to replace the original search algorithm. Instead, it may be used to generate an alternate optimized regression model of experimental data whenever the application of the original search algorithm either fails or requires too much CPU time. Data from a machine calibration of NASA's MK40 force balance is used to illustrate the application of the new regression model search algorithm.

multivariate experimental data↗

Analysis/forecast experiments with a multivariate statistical analysis scheme using FGGE data

A three-dimensional, multivariate, statistical analysis method, optimal interpolation (OI) is described for modeling meteorological data from widely dispersed sites. The model was developed to analyze FGGE data at the NASA-Goddard Laboratory of Atmospherics. The model features a multivariate surface analysis over the oceans, including maintenance of the Ekman balance and a geographically dependent correlation function. Preliminary comparisons are made between the OI model and similar schemes employed at the European Center for Medium Range Weather Forecasts and the National Meteorological Center. The OI scheme is used to provide input to a GCM, and model error correlations are calculated for forecasts of 500 mb vertical water mixing ratios and the wind profiles. Comparisons are made between the predictions and measured data. The model is shown to be as accurate as a successive corrections model out to 4.5 days.

Baker, W. E.↗

Analysis of Multivariate Experimental Data Using A Simplified Regression Model Search Algorithm

A new regression model search algorithm was developed that may be applied to both general multivariate experimental data sets and wind tunnel strain-gage balance calibration data. The algorithm is a simplified version of a more complex algorithm that was originally developed for the NASA Ames Balance Calibration Laboratory. The new algorithm performs regression model term reduction to prevent overfitting of data. It has the advantage that it needs only about one tenth of the original algorithm's CPU time for the completion of a regression model search. In addition, extensive testing showed that the prediction accuracy of math models obtained from the simplified algorithm is similar to the prediction accuracy of math models obtained from the original algorithm. The simplified algorithm, however, cannot guarantee that search constraints related to a set of statistical quality requirements are always satisfied in the optimized regression model. Therefore, the simplified algorithm is not intended to replace the original algorithm. Instead, it may be used to generate an alternate optimized regression model of experimental data whenever the application of the original search algorithm fails or requires too much CPU time. Data from a machine calibration of NASA's MK40 force balance is used to illustrate the application of the new search algorithm.

Ulbrich, Norbert M.↗

A method of using cluster analysis to study statistical dependence in multivariate data

A technique is presented that uses both cluster analysis and a Monte Carlo significance test of clusters to discover associations between variables in multidimensional data. The method is applied to an example of a noisy function in three-dimensional space, to a sample from a mixture of three bivariate normal distributions, and to the well-known Fisher's Iris data.

Borucki, W. J.↗

Logistic Risk Model for the Unique Effects of Inherent Aerobic Capacity on (+)G(sub z) Tolerance Before and After Simulated Weightlessness

Small sample size (n less than 1O) and inappropriate analysis of multivariate data have hindered previous attempts to describe which physiologic and demographic variables are most important in determining how long humans can tolerate acceleration. Data from previous centrifuge studies conducted at NASA/Ames Research Center, utilizing a 7-14 d bed rest protocol to simulate weightlessness, were included in the current investigation. After review, data on 25 women and 22 men were available for analysis. Study variables included gender, age, weight, height, percent body fat, resting heart rate, mean arterial pressure, Vo(sub 2)max and plasma volume. Since the dependent variable was time to greyout (failure), two contemporary biostatistical modeling procedures (proportional hazard and logistic discriminant function) were used to estimate risk, given a particular subject's profile. After adjusting for pro-bed-rest tolerance time, none of the profile variables remained in the risk equation for post-bed-rest tolerance greyout. However, prior to bed rest, risk of greyout could be predicted with 91% accuracy. All of the profile variables except weight, MAP, and those related to inherent aerobic capacity (Vo(sub 2)max, percent body fat, resting heart rate) entered the risk equation for pro-bed-rest greyout. A cross-validation using 24 new subjects indicated a very stable model for risk prediction, accurate within 5% of the original equation. The result for the inherent fitness variables is significant in that a consensus as to whether an increased aerobic capacity is beneficial or detrimental has not been satisfactorily established. We conclude that tolerance to +Gz acceleration before and after simulated weightlessness is independent of inherent aerobic fitness.

Ludwig, David A.↗

A CLIPS expert system for clinical flow cytometry data analysis

An expert system is being developed using CLIPS to assist clinicians in the analysis of multivariate flow cytometry data from cancer patients. Cluster analysis is used to find subpopulations representing various cell types in multiple datasets each consisting of four to five measurements on each of 5000 cells. CLIPS facts are derived from results of the clustering. CLIPS rules are based on the expertise of Drs. Stewart, Duque, and Braylan. The rules incorporate certainty factors based on case histories.

Salzman, G. C.↗

Materials surface contamination analysis

The original research objective was to demonstrate the ability of optical fiber spectrometry to determine contamination levels on solid rocket motor cases in order to identify surface conditions which may result in poor bonds during production. The capability of using the spectral features to identify contaminants with other sensors which might only indicate a potential contamination level provides a real enhancement to current inspection systems such as Optical Stimulated Electron Emission (OSEE). The optical fiber probe can easily fit into the same scanning fixtures as the OSEE. The initial data obtained using the Guided Wave Model 260 spectrophotometer was primarily focused on determining spectra of potential contaminants such as HD2 grease, silicones, etc. However, once we began taking data and applying multivariate analysis techniques, using a program that can handle very large data sets, i.e., Unscrambler 2, it became apparent that the techniques also might provide a nice scientific tool for determining oxidation and chemisorption rates under controlled conditions. As the ultimate power of the technique became recognized, considering that the chemical system which was most frequently studied in this work is water + D6AC steel, we became very interested in trying the spectroscopic techniques to solve a broad range of problems. The complexity of the observed spectra for the D6AC + water system is due to overlaps between the water peaks, the resulting chemisorbed species, and products of reaction which also contain OH stretching bands. Unscrambling these spectral features, without knowledge of the specific species involved, has proven to be a formidable task.

Workman, Gary L.↗

Experiments with a three-dimensional statistical objective analysis scheme using FGGE data

A three-dimensional (3D), multivariate, statistical objective analysis scheme (referred to as optimum interpolation or OI) has been developed for use in numerical weather prediction studies with the FGGE data. Some novel aspects of the present scheme include: (1) a multivariate surface analysis over the oceans, which employs an Ekman balance instead of the usual geostrophic relationship, to model the pressure-wind error cross correlations, and (2) the capability to use an error correlation function which is geographically dependent. A series of 4-day data assimilation experiments are conducted to examine the importance of some of the key features of the OI in terms of their effects on forecast skill, as well as to compare the forecast skill using the OI with that utilizing a successive correction method (SCM) of analysis developed earlier. For the three cases examined, the forecast skill is found to be rather insensitive to varying the error correlation function geographically. However, significant differences are noted between forecasts from a two-dimensional (2D) version of the OI and those from the 3D OI, with the 3D OI forecasts exhibiting better forecast skill. The 3D OI forecasts are also more accurate than those from the SCM initial conditions. The 3D OI with the multivariate oceanic surface analysis was found to produce forecasts which were slightly more accurate, on the average, than a univariate version.

Baker, Wayman E.↗

Interfaces between statistical analysis packages and the ESRI geographic information system

Interfaces between ESRI's geographic information system (GIS) data files and real valued data files written to facilitate statistical analysis and display of spatially referenced multivariable data are described. An example of data analysis which utilized the GIS and the statistical analysis system is presented to illustrate the utility of combining the analytic capability of a statistical package with the data management and display features of the GIS.

Masuoka, E.↗

TPSAS-NF1676L-36113-DND

Concerns about the effects of extreme heat and poor air quality are increasing in North America’s largest urban centers. In Philadelphia, environmental and public health groups are concerned about how these phenomena disproportionality affect marginalized communities and populations, which often have extensive impervious surfaces and little access to green space. In order to address these concerns, the Philadelphia Department of Public Health and the Office of Sustainability seek to effectively prioritize cooling initiatives to reduce urban heat and decrease air pollutants. We evaluated land surface temperature (LST) and the Normalized Difference Vegetation Index (NDVI), as a measure of overall greenness, obtained from NASA Earth observations Aqua and Terra Moderate Resolution Imaging Spectroradiometer (MODIS), and the Ecosystem Spaceborne Thermal Radiometer Experiment on Space Station (ECOSTRESS). These analyses we recombined with local tree inventory, air quality, and socioeconomic data through a multivariate analysis to identify areas where new trees or cooling adaptations are most needed. The results and data of this project can be used by our partners to inform both short-term heat relief planning and a long-term, multi-agency heat response.

Spencer Nelson↗

Hierarchical analysis of spatial pattern and processes of Douglas-fir forests

There has been an increased interest in the quantification of pattern in ecological systems over the past years. This interest is motivated by the desire to construct valid models which extend across many scales. Spatial methods must quantify pattern, discriminate types of pattern, and relate hierarchical phenomena across scales. Wavelet analysis is introduced as a method to identify spatial structure in ecological transect data. The main advantage of the wavelet transform over other methods is its ability to preserve and display hierarchical information while allowing for pattern decomposition. Two applications of wavelet analysis are illustrated, as a means to: (1) quantify known spatial patterns in Douglas-fir forests at several scales, and (2) construct spatially-explicit hypotheses regarding pattern generating mechanisms. Application of the wavelet variance, derived from the wavelet transform, is developed for forest ecosystem analysis to obtain additional insight into spatially-explicit data. Specifically, the resolution capabilities of the wavelet variance are compared to the semi-variogram and Fourier power spectra for the description of spatial data using a set of one-dimensional stationary and non-stationary processes. The wavelet cross-covariance function is derived from the wavelet transform and introduced as a alternative method for the analysis of multivariate spatial data of understory vegetation and canopy in Douglas-fir forests of the western Cascades of Oregon.

Bradshaw, G. A.↗

Digital preprocessing and classification of multispectral earth observation data

The development of airborne and satellite multispectral image scanning sensors has generated wide-spread interest in application of these sensors to earth resource mapping. These point scanning sensors permit scenes to be imaged in a large number of electromagnetic energy bands between .3 and 15 micrometers. The energy sensed in each band can be used as a feature in a computer based multi-dimensional pattern recognition process to aid in interpreting the nature of elements in the scene. Images from each band can also be interpreted visually. Visual interpretation of five or ten multispectral images simultaneously becomes impractical especially as area studied increases; hence, great emphasis has been placed on machine (computer) techniques for aiding in the interpretation process. This paper describes a computer software system concept called LARSYS for analysis of multivariate image data and presents some examples of its application.

Anuta, P. E.↗

Studies in Astronomical Time Series Analysis. VI. Bayesian Block Representations

This paper addresses the problem of detecting and characterizing local variability in time series and other forms of sequential data. The goal is to identify and characterize statistically significant variations, at the same time suppressing the inevitable corrupting observational errors. We present a simple nonparametric modeling technique and an algorithm implementing it-an improved and generalized version of Bayesian Blocks [Scargle 1998]-that finds the optimal segmentation of the data in the observation interval. The structure of the algorithm allows it to be used in either a real-time trigger mode, or a retrospective mode. Maximum likelihood or marginal posterior functions to measure model fitness are presented for events, binned counts, and measurements at arbitrary times with known error distributions. Problems addressed include those connected with data gaps, variable exposure, extension to piece- wise linear and piecewise exponential representations, multivariate time series data, analysis of variance, data on the circle, other data modes, and dispersed data. Simulations provide evidence that the detection efficiency for weak signals is close to a theoretical asymptotic limit derived by [Arias-Castro, Donoho and Huo 2003]. In the spirit of Reproducible Research [Donoho et al. (2008)] all of the code and data necessary to reproduce all of the figures in this paper are included as auxiliary material.

signal detection↗