Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “multivariate data analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Nonparametric analysis of Minnesota spruce and aspen tree data and LANDSAT data

The application of nonparametric methods in data-intensive problems faced by NASA is described. The theoretical development of efficient multivariate density estimators and the novel use of color graphics workstations are reviewed. The use of nonparametric density estimates for data representation and for Bayesian classification are described and illustrated. Progress in building a data analysis system in a workstation environment is reviewed and preliminary runs presented.

Scott, D. W.↗

A computer analysis of ERTS data of the Lake Gregory area of South Australia with particular emphasis on its role in terrain classification for engineering

A digital computer and multivariate statistical techniques were used to analyze 4-band multispectral data. A representation of the original data for each of the four bands allows a certain degree of terrain interpretation; however, variations in appearance of sites within and between bands, without additional criteria for deciding which representation should be preferred, create difficulties for classification. Investigation of the video data groups produced by principal components analysis and cluster analysis techniques shows that effective correlations with classifications of terrain produced by conventional methods could be carried out. The analyses also highlighted underlying relationships between the various elements. The approach used allows large areas (185 cm by 185 cm) to be classified into fundamental units within a matter of hours and can be applied to those parts of the Earth where facilities for conventional studies are poor or lacking.

Lodwick, G. D.↗

Probabilistic Machine Learning Estimation of Ocean Mixed Layer Depth from Dense Satellite and Sparse In-Situ Observations

The ocean mixed layer plays an important role in the coupling between the upper ocean and atmosphere across a wide range of time scales. Estimation of the variability of the ocean mixed layer is therefore important for atmosphere-ocean prediction and analysis. The increasing coverage of in situ Argo profile data allows for an increasingly accurate analysis of the mixed layer depth (MLD) variability associated with deviations from the seasonal climatology. However, sampling rates are not sufficient to fully resolve subseasonal (<90 day) MLD variability. Yet, many multivariate observations-based analyses include implicit modeled subseasonal MLD variability. One analysis method is optimal interpolation of in situ data, but the interior analysis can be improved by leveraging surface data with regression or variational approaches. Here, we demonstrate how machine learning methods and satellite sea surface temperature, salinity, and height facilitate MLD estimation in a pilot study of two regions: the mid-latitude southern Indian and the eastern equatorial Pacific Oceans. We construct multiple machine learning architectures to produce weekly 1/2° gridded MLD anomaly fields (relative to a monthly climatology) with uncertainty estimates. We test multiple traditional and probabilistic machine learning techniques to compare both accuracy and probabilistic calibration. We validate our methodology by applying it to ocean model simulations. We find that incorporating sea surface data through a machine learning model improves the performance of spatiotemporal MLD variability estimation compared to optimal interpolation of Argo observations alone. These preliminary results are a promising first step for the application of machine learning to MLD prediction.

Machine Learning↗

New Constraints on Titan's Stratospheric n-Butane Abundance

Curiously, n-butane has yet to be detected at Titan, though it is predicted to be present in a wide range of abundances that span over 2.5 orders of magnitude. We have searched infrared spectroscopic observations of Titan for signals from n-butane (n-C4H10) in Titan's stratosphere. Three sets of Cassini Composite Infrared Spectrometer Focal Plane 4 (1050–1500 cm−1) observations were selected for modeling, having been collected from different flybys and pointing latitudes. We modeled the observations with the Nonlinear Optimal Estimator for MultivariatE Spectral AnalySIS radiative transfer tool. Temperature profiles were retrieved for each of the data sets by modeling the ν4 emission from methane near 1305 cm−1. Then, incorporating the temperature profiles, we retrieved abundances of all of Titan's known trace gases that are active in this spectral region, reliably reproducing the observations. We then systematically tested a set of models with varying abundances of n-butane, investigating how the addition of this gas affected the fits. We did this for several different photochemically predicted abundance profiles from the literature, as well as for a constant-with-altitude profile. Ultimately, though we did not produce any firm detection of n-butane, we derived new upper limits on its abundance specific to the use of each profile and to multiple different ranges of stratospheric altitudes. These results will tightly constrain the C4 chemistry of future photochemical modeling of Titan's atmosphere and also motivate the continued search for n-butane and its isomer, isobutane.

Brendan L. Steffens↗

Chemical studies of H chondrites. 4: New data and comparison of Antarctic suites

We report data for the trace elements Au, Co, Sb, Ga, Rb, Ag, Se, Cs, Te, Zn, Cd, Bi, Ti, and In (ordered by putative volatility during nebular condensation and accretion) determined by neutron activation analysis in 13 H5 chondrites from Victoria Land and 20 H4-6 chondrites from Queen Maud Land, Antarctica. These and earlier results provide Antarctic sample suites of 34 chondrites from Victoria Land and 25 from Queen Maud Land. Treatment of data for the most volatile 10 elements (Rb to In) in these studies by multivariate statistical techniques more robust, as well as more conservative, than conventional linear discriminant analysis and logistic regression demonstrates that compositions differ at marginally significant levels. This difference cannot be explained by trivial (terrestrial) causes and becomes more significant, despite the smaller size of the database, when comparisons are limited to data from a single analyst and when all upper limits are eliminated from consideration. The Victoria Land and Queen Maud Land suites have different mean terrestrial ages (approximately 300 kyr and approximately 100 kyr, respectively) and age distributions, suggesting that a time-dependent variation of chondritic sources with different thermal histories is responsible. As a result, these two Antarctic suites are, on average, chemically distinguishable from each other. Since H chondrites serve as a paradigm for other meteorite classes, these results indicate that the near-Earth populations of planetary materials varied with time on the 10(exp 5)-year timescale.

Wolf, Stephen F.↗

Reflectance of vegetation, soil, and water

There are no author-identified significant results in this report. This report deals with the selection of the best channels from the 24-channel aircraft data to represent crop and soil conditions. A three-step procedure has been developed that involves using univariate statistics and an F-ratio test to indicate the best 14 channels. From the 14, the 10 best channels are selected by a multivariate stochastic process. The third step involves the pattern recognition procedures developed in the data analysis plan. Indications are that the procedures in use are satsifactory and will extract the desired information from the data.

Wiegand, C. L.↗

The second-moment climatology of the GATE rain rate data

The first part of this paper presents the description of the GARP (Global Atmospheric Research Program) Atlantic Tropical Experiment (GATE) 1 rain-rate data and its two-dimensional spectral and correlation characteristics, which has made it possible to accomplish the following: to show the concentration of a significant power along the frequency axis in the spatiotemporal spectra; to detect a diurnal cycle (which has a range of variation of about 3.4-5.4 mm/n) as one of the sources of bias in the rain statistics of satellite data; to study the distinction between the north-south and east-west transport of spatial rain-rate field and character of its anisotropy; to evaluate the scales of the distinction between second-moment estimates associated with ground and satellite samples; and to determine the appropriate spatial and temporal scales of simple linear stochastic models fitted to averaged rain-rate fields. The second part of this paper is devoted to an analysis of the diffusion of the rain rate by establishing a relationship between the parameters of the multivariate autoregressive model and the coefficients of a diffusion equation. This analysis led to the use of rain data to estimate the rain advection velocity as well as other coefficients of the diffusion equation of the corresponding field. The results obtained can be used for comparison with corresponding estimates of other sources of data (satellite, Tropical Oceans Global Atmosphere Coupled Ocean - Atmosphere Response Experiment (TOGA, COARE) or simulated by physical models), for generating multiple samples of any size, for solving the inverse problems of some of the hydrodynamic equations, and in some other areas of rain data analysis and modeling.

Polyak, Ilya↗

Statistical Analysis of Factors Riving Surface Ozone Variability over Continental South Africa

Statistical relationships between surface ozone (O3) concentration, precursor species and meteorological conditions in continental South Africa were examined from data obtained from measurement stations in north-eastern South Africa. Three multivariate statistical methods were applied in the investigation, i.e. multiple linear regression (MLR), principal component analysis (PCA) and –regression (PCR), and generalised additive model (GAM) analysis. The daily maximum 8-h moving average O3 concentrations were considered in these statistical models (dependent variable). MLR models indicated that meteorology and precursor species concentrations are able to explain ~50% of the variability in daily maximum O3 levels. MLR analysis revealed that atmospheric carbon monoxide (CO), temperature and relative humidity were the strongest factors affecting the daily O3 variability. In summer, daily O3 variances were mostly associated with relative humidity, while winter O3 levels were mostly linked to temperature and CO. PCA indicated that CO, temperature and relative humidity were not strongly collinear. GAM also identified CO, temperature and relative humidity as the strongest factors affecting the daily variation of O3. Partial residual plots found that temperature, radiation and nitrogen oxides most likely have a non-linear relationship with O3,while the relationship with relative humidity and CO is probably linear. An inter-comparison between O3 levels modelled with the three statistical models compared to measured O3 concentrations showed that the GAM model offered a slight improvement over the MLR model. These findings emphasise the critical role of regional-scale O3 precursors coupled with meteorological conditions in daily variances of O3 levels in continental South Africa.

multiple linear regression (MLR)↗

An Integrated Analysis of the Physiological Effects of Space Flight: Executive Summary

A large array of models were applied in a unified manner to solve problems in space flight physiology. Mathematical simulation was used as an alternative way of looking at physiological systems and maximizing the yield from previous space flight experiments. A medical data analysis system was created which consist of an automated data base, a computerized biostatistical and data analysis system, and a set of simulation models of physiological systems. Five basic models were employed: (1) a pulsatile cardiovascular model; (2) a respiratory model; (3) a thermoregulatory model; (4) a circulatory, fluid, and electrolyte balance model; and (5) an erythropoiesis regulatory model. Algorithms were provided to perform routine statistical tests, multivariate analysis, nonlinear regression analysis, and autocorrelation analysis. Special purpose programs were prepared for rank correlation, factor analysis, and the integration of the metabolic balance data.

Leonard, J. I.↗

Multivariate space - time analysis of PRE-STORM precipitation

This paper presents the methodologies and results of the multivariate modeling and two-dimensional spectral and correlation analysis of PRE-STORM rainfall gauge data. Estimated parameters of the models for the specific spatial averages clearly indicate the eastward and southeastward wave propagation of rainfall fluctuations. A relationship between the coefficients of the diffusion equation and the parameters of the stochastic model of rainfall fluctuations is derived that leads directly to the exclusive use of rainfall data to estimate advection speed (about 12 m/s) as well as other coefficients of the diffusion equation of the corresponding fields. The statistical methodology developed here can be used for confirmation of physical models by comparison of the corresponding second-moment statistics of the observed and simulated data, for generating multiple samples of any size, for solving the inverse problem of the hydrodynamic equations, and for application in some other areas of meteorological and climatological data analysis and modeling.

Polyak, Ilya↗

M-DAS: System for multispectral data analysis

M-DAS is a ground data processing system designed for analysis of multispectral data. M-DAS operates on multispectral data from LANDSAT, S-192, M2S and other sources in CCT form. Interactive training by operator-investigators using a variable cursor on a color display was used to derive optimum processing coefficients and data on cluster separability. An advanced multivariate normal-maximum likelihood processing algorithm was used to produce output in various formats: color-coded film images, geometrically corrected map overlays, moving displays of scene sections, coverage tabulations and categorized CCTs. The analysis procedure for M-DAS involves three phases: (1) screening and training, (2) analysis of training data to compute performance predictions and processing coefficients, and (3) processing of multichannel input data into categorized results. Typical M-DAS applications involve iteration between each of these phases. A series of photographs of the M-DAS display are used to illustrate M-DAS operation.

Johnson, R. H.↗

Noise and drift analysis of non-equally spaced timing data

Generally, it is possible to obtain equally spaced timing data from oscillators. The measurement of the drifts and noises affecting oscillators is then performed by using a variance (Allan variance, modified Allan variance, or time variance) or a system of several variances (multivariance method). However, in some cases, several samples, or even several sets of samples, are missing. In the case of millisecond pulsar timing data, for instance, observations are quite irregularly spaced in time. Nevertheless, since some observations are very close together (one minute) and since the timing data sequence is very long (more than ten years), information on both short-term and long-term stability is available. Unfortunately, a direct variance analysis is not possible without interpolating missing data. Different interpolation algorithms (linear interpolation, cubic spline) are used to calculate variances in order to verify that they neither lose information nor add erroneous information. A comparison of the results of the different algorithms is given. Finally, the multivariance method was adapted to the measurement sequence of the millisecond pulsar timing data: the responses of each variance of the system are calculated for each type of noise and drift, with the same missing samples as in the pulsar timing sequence. An estimation of precision, dynamics, and separability of this method is given.

Vernotte, F.↗

The MIDAS processor

The MIDAS (Multivariate Interactive Digital Analysis System) processor is a high-speed processor designed to process multispectral scanner data (from Landsat, EOS, aircraft, etc.) quickly and cost-effectively to meet the requirements of users of remote sensor data, especially from very large areas. MIDAS consists of a fast multipipeline preprocessor and classifier, an interactive color display and color printer, and a medium scale computer system for analysis and control. The system is designed to process data having as many as 16 spectral bands per picture element at rates of 200,000 picture elements per second into as many as 17 classes using a maximum likelihood decision rule.

Kriegler, F. J.↗

Early experiences building a software quality prediction model

Early experiences building a software quality prediction model are discussed. The overall research objective is to establish a capability to project a software system's quality from an analysis of its design. The technical approach is to build multivariate models for estimating reliability and maintainability. Data from 21 Ada subsystems were analyzed to test hypotheses about various design structures leading to failure-prone or unmaintainable systems. Current design variables highlight the interconnectivity and visibility of compilation units. Other model variables provide for the effects of reusability and software changes. Reported results are preliminary because additional project data is being obtained and new hypotheses are being developed and tested. Current multivariate regression models are encouraging, explaining 60 to 80 percent of the variation in error density of the subsystems.

Agresti, W. W.↗

Statistical analysis of Thematic Mapper Simulator data for the geobotanical discrimination of rock types in southwest Oregon

An evaluation of Thematic Mapper Simulator (TMS) data for the geobotanical discrimination of rock types based on vegetative cover characteristics is addressed in this research. A methodology for accomplishing this evaluation utilizing univariate and multivariate techniques is presented. TMS data acquired with a Daedalus DEI-1260 multispectral scanner were integrated with vegetation and geologic information for subsequent statistical analyses, which included a chi-square test, an analysis of variance, stepwise discriminant analysis, and Duncan's multiple range test. Results indicate that ultramafic rock types are spectrally separable from nonultramafics based on vegetative cover through the use of statistical analyses.

Morrissey, L. A.↗

MIDAS, prototype Multivariate Interactive Digital Analysis System, phase 1. Volume 3: Wiring diagrams

The Midas System is a third-generation, fast, multispectral recognition system able to keep pace with the large quantity and high rates of data acquisition from present and projected sensors. A principal objective of the MIDAS Program is to provide a system well interfaced with the human operator and thus to obtain large overall reductions in turn-around time and significant gains in throughput. The hardware and software generated in Phase I of the overall program are described. The system contains a mini-computer to control the various high-speed processing elements in the data path and a classifier which implements an all-digital prototype multivariate-Gaussian maximum likelihood decision algorithm operating at 2 x 100,000 pixels/sec. Sufficient hardware was developed to perform signature extraction from computer-compatible tapes, compute classifier coefficients, control the classifier operation, and diagnose operation. The MIDAS construction and wiring diagrams are given.

Kriegler, F. J.↗

MIDAS, prototype Multivariate Interactive Digital Analysis System, Phase 1. Volume 2: Diagnostic system

The MIDAS System is a third-generation, fast, multispectral recognition system able to keep pace with the large quantity and high rates of data acquisition from present and projected sensors. A principal objective of the MIDAS Program is to provide a system well interfaced with the human operator and thus to obtain large overall reductions in turn-around time and significant gains in throughout. The hardware and software generated in Phase I of the over-all program are described. The system contains a mini-computer to control the various high-speed processing elements in the data path and a classifier which implements an all-digital prototype multivariate-Gaussian maximum likelihood decision algorithm operating 2 x 105 pixels/sec. Sufficient hardware was developed to perform signature extraction from computer-compatible tapes, compute classifier coefficients, control the classifier operation, and diagnose operation. Diagnostic programs used to test MIDAS' operations are presented.

Kriegler, F. J.↗

MIDAS, prototype Multivariate Interactive Digital Analysis System, phase 1. Volume 1: System description

The MIDAS System is described as a third-generation fast multispectral recognition system able to keep pace with the large quantity and high rates of data acquisition from present and projected sensors. A principal objective of the MIDAS program is to provide a system well interfaced with the human operator and thus to obtain large overall reductions in turnaround time and significant gains in throughput. The hardware and software are described. The system contains a mini-computer to control the various high-speed processing elements in the data path, and a classifier which implements an all-digital prototype multivariate-Gaussian maximum likelihood decision algorithm operating at 200,000 pixels/sec. Sufficient hardware was developed to perform signature extraction from computer-compatible tapes, compute classifier coefficients, control the classifier operation, and diagnose operation.

Kriegler, F. J.↗