Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “multivariate data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Alteration mapping at Goldfield, Nevada, by cluster and discriminant analysis of Landsat digital data

The ability of Landsat multispectral digital data to differentiate among 62 combinations of rock and alteration types at the Goldfield mining district of Western Nevada was investigated by using statistical techniques of cluster and discriminant analysis. Multivariate discriminant analysis was not effective in classifying each of the 62 groups, with classification results essentially the same whether data of four channels alone or combined with six ratios of channels were used. Bivariate plots of group means revealed a cluster of three groups including mill tailings, basalt and all other rock and alteration types. Automatic hierarchical clustering based on the fourth dimensional Mahalanobis distance between group means of 30 groups having five or more samples was performed using Johnson's HICLUS program. The results of the cluster analysis revealed hierarchies of mill tailings vs. natural materials, basalt vs. non-basalt, highly reflectant rocks vs. other rocks and exclusively unaltered rocks vs. predominantly altered rocks. The hierarchies were used to determine the order in which sets of multiple discriminant analyses were to be performed and the resulting discriminant functions were used to produce a map of geology and alteration which has an overall accuracy of 70 percent for discriminating exclusively altered rocks from predominantly altered rocks.

Ballew, G.↗

Multivariate optimum interpolation of surface pressure and winds over oceans

The observations of surface pressure are quite sparse over oceanic areas. An effort to improve the analysis of surface pressure over oceans through the development of a multivariate surface analysis scheme which makes use of surface pressure and wind data is discussed. Although the present research used ship winds, future versions of this analysis scheme could utilize winds from additional sources, such as satellite scatterometer data.

Bloom, S. C.↗

Seasonal variability of bio-optical and physical properties in the Sargasso Sea

The seasonal variability of bio-optical and physical properties within the upper ocean at a site in the Sargasso Sea has been observed in multivariable moored systems during a 9-month period. In addition, complementary meteorological data, sea surface height and sea surface temperature maps, and expendable bathythermograph and shipboard profile data have been utilized for interpretation. The observations during March are characteristic of late wintertime conditions of a deep isothermal layer, but with intervening periods of warming due to the advection of warm outbreak waters associated with Gulf Stream meanders. The mixed layer depth shoals from greater than 160 m to about 25 m in late March (spring transition). Phytoplankton blooms follow the mixed layer shoaling. A succession of phytoplankton populations occurs during this transitional interval. The mixed layer remains near 25 m for the summer and deepens in mid-September. A relatively intense subsurface maximum in chlorophyll develops at about 75 m following the spring transition. The maximum persists, but weakens in mid-summer.

Dickey, T.↗

Fast and Flexible Multivariate Time Series Subsequence Search

Multivariate Time-Series (MTS) are ubiquitous, and are generated in areas as disparate as sensor recordings in aerospace systems, music and video streams, medical monitoring, and financial systems. Domain experts are often interested in searching for interesting multivariate patterns from these MTS databases which often contain several gigabytes of data. Surprisingly, research on MTS search is very limited. Most of the existing work only supports queries with the same length of data, or queries on a fixed set of variables. In this paper, we propose an efficient and flexible subsequence search framework for massive MTS databases, that, for the first time, enables querying on any subset of variables with arbitrary time delays between them. We propose two algorithms to solve this problem (1) a List Based Search (LBS) algorithm which uses sorted lists for indexing, and (2) a R*-tree Based Search (RBS) which uses Minimum Bounding Rectangles (MBR) to organize the subsequences. Both algorithms guarantee that all matching patterns within the specified thresholds will be returned (no false dismissals). The very few false alarms can be removed by a post-processing step. Since our framework is also capable of Univariate Time-Series (UTS) subsequence search, we first demonstrate the efficiency of our algorithms on several UTS datasets previously used in the literature. We follow this up with experiments using two large MTS databases from the aviation domain, each containing several millions of observations. Both these tests show that our algorithms have very high prune rates (>99%) thus needing actual disk access for only less than 1% of the observations. To the best of our knowledge, MTS subsequence search has never been attempted on datasets of the size we have used in this paper.

Bhaduri, Kanishka↗

The Challenge Posed by Geomagnetic Activity to Electric Power Reliability: Evidence from England and Wales

This paper addresses whether geomagnetic activity challenged the reliability of the electric power system during part of the declining phase of solar cycle 23. Operations by National Grid in England and Wales are examined over the period of 11 March 2003 through 31 March 2005. This paper examines the relationship between measures of geomagnetic activity and a metric of challenged electric power reliability known as the net imbalance volume (NIV). Measured in megawatt hours, NIV represents the sum of all energy deployments initiated by the system operator to balance the electric power system. The relationship between geomagnetic activity and NIV is assessed using a multivariate econometric model. The model was estimated using half-hour settlement data over the period of 11 March 2003 through 31 December 2004. The results indicate that geomagnetic activity had a demonstrable effect on NIV over the sample period. Based on the parameter estimates, out-of-sample predictions of NIV were generated for each half hour over the period of 1 January to 31 March 2005. Consistent with the existence of a causal relationship between geomagnetic activity and the electricity market imbalance, the root-mean-square error of the out-of-sample predictions of NIV is smaller; that is, the predictions are more accurate, when the statistically significant estimated effects of geomagnetic activity are included as drivers in the predictions.

Forbes, Kevin F.↗

Statistical Analysis of Factors Riving Surface Ozone Variability over Continental South Africa

Statistical relationships between surface ozone (O3) concentration, precursor species and meteorological conditions in continental South Africa were examined from data obtained from measurement stations in north-eastern South Africa. Three multivariate statistical methods were applied in the investigation, i.e. multiple linear regression (MLR), principal component analysis (PCA) and –regression (PCR), and generalised additive model (GAM) analysis. The daily maximum 8-h moving average O3 concentrations were considered in these statistical models (dependent variable). MLR models indicated that meteorology and precursor species concentrations are able to explain ~50% of the variability in daily maximum O3 levels. MLR analysis revealed that atmospheric carbon monoxide (CO), temperature and relative humidity were the strongest factors affecting the daily O3 variability. In summer, daily O3 variances were mostly associated with relative humidity, while winter O3 levels were mostly linked to temperature and CO. PCA indicated that CO, temperature and relative humidity were not strongly collinear. GAM also identified CO, temperature and relative humidity as the strongest factors affecting the daily variation of O3. Partial residual plots found that temperature, radiation and nitrogen oxides most likely have a non-linear relationship with O3,while the relationship with relative humidity and CO is probably linear. An inter-comparison between O3 levels modelled with the three statistical models compared to measured O3 concentrations showed that the GAM model offered a slight improvement over the MLR model. These findings emphasise the critical role of regional-scale O3 precursors coupled with meteorological conditions in daily variances of O3 levels in continental South Africa.

multiple linear regression (MLR)↗

LinkWinds: An Approach to Visual Data Analysis

The Linked Windows Interactive Data System (LinkWinds) is a prototype visual data exploration and analysis system resulting from a NASA/JPL program of research into graphical methods for rapidly accessing, displaying and analyzing large multivariate multidisciplinary datasets. It is an integrated multi-application execution environment allowing the dynamic interconnection of multiple windows containing visual displays and/or controls through a data-linking paradigm. This paradigm, which results in a system much like a graphical spreadsheet, is not only a powerful method for organizing large amounts of data for analysis, but provides a highly intuitive, easy to learn user interface on top of the traditional graphical user interface.

Jacobson, Allan S.↗

Lessons Learned from Assimilating Altimeter Data into a Coupled General Circulation Model with the GMAO Augmented Ensemble Kalman Filter

Satellite altimetry measurements have provided global, evenly distributed observations of the ocean surface since 1993. However, the difficulties introduced by the presence of model biases and the requirement that data assimilation systems extrapolate the sea surface height (SSH) information to the subsurface in order to estimate the temperature, salinity and currents make it difficult to optimally exploit these measurements. This talk investigates the potential of the altimetry data assimilation once the biases are accounted for with an ad hoc bias estimation scheme. Either steady-state or state-dependent multivariate background-error covariances from an ensemble of model integrations are used to address the problem of extrapolating the information to the sub-surface. The GMAO ocean data assimilation system applied to an ensemble of coupled model instances using the GEOS-5 AGCM coupled to MOM4 is used in the investigation. To model the background error covariances, the system relies on a hybrid ensemble approach in which a small number of dynamically evolved model trajectories is augmented on the one hand with past instances of the state vector along each trajectory and, on the other, with a steady state ensemble of error estimates from a time series of short-term model forecasts. A state-dependent adaptive error-covariance localization and inflation algorithm controls how the SSH information is extrapolated to the sub-surface. A two-step predictor corrector approach is used to assimilate future information. Independent (not-assimilated) temperature and salinity observations from Argo floats are used to validate the assimilation. A two-step projection method in which the system first calculates a SSH increment and then projects this increment vertically onto the temperature, salt and current fields is found to be most effective in reconstructing the sub-surface information. The performance of the system in reconstructing the sub-surface fields is particularly impressive for temperature, but not as satisfactory for salt.

Keppenne, Christian↗

Probabilistic Machine Learning Estimation of Ocean Mixed Layer Depth from Dense Satellite and Sparse In-Situ Observations

The ocean mixed layer plays an important role in the coupling between the upper ocean and atmosphere across a wide range of time scales. Estimation of the variability of the ocean mixed layer is therefore important for atmosphere-ocean prediction and analysis. The increasing coverage of in situ Argo profile data allows for an increasingly accurate analysis of the mixed layer depth (MLD) variability associated with deviations from the seasonal climatology. However, sampling rates are not sufficient to fully resolve subseasonal (<90 day) MLD variability. Yet, many multivariate observations-based analyses include implicit modeled subseasonal MLD variability. One analysis method is optimal interpolation of in situ data, but the interior analysis can be improved by leveraging surface data with regression or variational approaches. Here, we demonstrate how machine learning methods and satellite sea surface temperature, salinity, and height facilitate MLD estimation in a pilot study of two regions: the mid-latitude southern Indian and the eastern equatorial Pacific Oceans. We construct multiple machine learning architectures to produce weekly 1/2° gridded MLD anomaly fields (relative to a monthly climatology) with uncertainty estimates. We test multiple traditional and probabilistic machine learning techniques to compare both accuracy and probabilistic calibration. We validate our methodology by applying it to ocean model simulations. We find that incorporating sea surface data through a machine learning model improves the performance of spatiotemporal MLD variability estimation compared to optimal interpolation of Argo observations alone. These preliminary results are a promising first step for the application of machine learning to MLD prediction.

Machine Learning↗

Remote sensing of earth terrain

Two monographs and 85 journal and conference papers on remote sensing of earth terrain have been published, sponsored by NASA Contract NAG5-270. A multivariate K-distribution is proposed to model the statistics of fully polarimetric data from earth terrain with polarizations HH, HV, VH, and VV. In this approach, correlated polarizations of radar signals, as characterized by a covariance matrix, are treated as the sum of N n-dimensional random vectors; N obeys the negative binomial distribution with a parameter alpha and mean bar N. Subsequently, and n-dimensional K-distribution, with either zero or non-zero mean, is developed in the limit of infinite bar N or illuminated area. The probability density function (PDF) of the K-distributed vector normalized by its Euclidean norm is independent of the parameter alpha and is the same as that derived from a zero-mean Gaussian-distributed random vector. The above model is well supported by experimental data provided by MIT Lincoln Laboratory and the Jet Propulsion Laboratory in the form of polarimetric measurements.

Kong, J. A.↗

Improved Regolith-Landform and Geological Mapping using AIRSAR Data as an Aid to Mineral Exploration in the North-Eastern Goldfields Region, Western Australia

The distinctive contribution of AIRSAR data in characterizing the regolith-landforms in a relatively vegetation free environment are discussed. AIRSAR frame processed data were initially MAF-cleaned to enhance the signal content of the data before geocoding to AMG coordinates. Colors in a three frequency single polarization combination image C-, L- P- bands, and in an enhancement of pedestal height, relate directly to the scale of surface roughness of the various regolith units Examination of the AIRSAR enhancements reveals that in mafic terrain, and to a lesser extent, in felsic terrain, AIRSAR data provides discrimination between the principle geomorphic regimes, relict, erosional and depositional. A multivariate statistical technique called an all-possible subsets calculation was used to examine the degree of polarimetric separation between selected regolith-landforms for all combinations of the nine band AIRSAR radar. An unanticipated aspect of the research was the identification on the AIRSAR imagery of previously unmapped structural features.

Tapley, Ian J.↗

Predictability of Malaria Transmission Intensity in the Mpumalanga Province, South Africa, Using Land Surface Climatology and Autoregressive Analysis

There has been increasing effort in recent years to employ satellite remotely sensed data to identify and map vector habitat and malaria transmission risk in data sparse environments. In the current investigation, available satellite and other land surface climatology data products are employed in short-term forecasting of infection rates in the Mpumalanga Province of South Africa, using a multivariate autoregressive approach. The climatology variables include precipitation, air temperature and other land surface states computed by the Off-line Land-Surface Global Assimilation System (OLGA) including soil moisture and surface evaporation. Satellite data products include the Normalized Difference Vegetation Index (NDVI) and other forcing data used in the Goddard Earth Observing System (GEOS-1) model. Predictions are compared to long- term monthly records of clinical and microscopic diagnoses. The approach addresses the high degree of short-term autocorrelation in the disease and weather time series. The resulting model is able to predict 11 of the 13 months that were classified as high risk during the validation period, indicating the utility of applying antecedent climatic variables to the prediction of malaria incidence for the Mpumalanga Province.

Grass, David↗

Identification of aerodynamic indicial functions using flight data

It is pointed out that the use of indicial function representation provides a model superior to the aerodynamic derivative model. Specific derivatives can be approximated from the indicial models. The model can also be used to compute equivalent stability and control parameters not usually available from flight data. It is shown that derivatives regarding the angle-of-attack and the side slip angle can be derived directly from the indicial functions without any identifiability problem. Attention is given to the pitch moment coefficient, linear indicial function representation, the identification problem for the pitch moment equation, the identifiability of linear systems, parametric representations of the indicial functions, an identification technique, angle-of-attack and pitch rate dynamics in the pitch plane, multivariate linear models, nonlinear aerodynamic indicial functions, measurement system accuracy, and poststall and spin-entry data from a scaled research vehicle.

Gupta, N. K.↗

Mapping Arid Vegetation Species Distributions in the White Mountains, Eastern California, Using AVIRIS, Topography, and Geology

Our challenge is to model plant species distributions in complex montane environments using disparate sources of data, including topography, geology, and hyperspectral data. From an ecologist's point of view, species distributions are determined by local environment and disturbance history, while spectral data are 'ancillary.' However, a remote sensor's perspective says that spectral data provide picture of what vegetation is there, topographic and geologic data are ancillary. In order to bridge the gap, all available data should be used to get the best possible prediction of species distributions using complex multivariate techniques implemented on a GIS. Vegetation reflects local climatic and nutrient conditions, both of which can be modeled, allowing predictive mapping of vegetation distributions. Geologic substrate strongly affects chemical, thermal, and physical properties of soils, while climatic conditions are determined by local topography. As elevation increases, precipitation increases and temperature decreases. Aspect, slope, and surrounding topography determine potential insolation, so that south-facing slopes are warmer and north-facing slopes cooler at a given elevation. Topographic position (ridge, slope, canyon, or meadow) and slope angle affect sediment accumulation and soil depth. These factors combine as complex environmental gradients, and underlie many features of plant distributions. Airborne Visible/Infrared Imaging Spectrometer (AVIRIS) data, digital elevation models, digitized geologic maps, and 378 ground control points were used to predictively map species distributions in the central and southern White Mountains, along the western boundary of the Basin and Range province. Minimum Noise Fraction (MNF) bands were calculated from the visible and near-infrared AVIRIS bands, and combined with digitized geologic maps and topographic variables using Canonical Correspondence Analysis (CCA). CCA allows for modeling species 'envelopes' in multidimensional environmental space, which can then be projected across entire landscapes.

VandeVen, C.↗

K-distribution and polarimetric terrain radar clutter

A multivariate K-distribution is proposed to model the statistics of fully polarimetric radar data from earth terrain with polarizations HH, HV, VH, and VV. In this approach, correlated polarizations of radar signals, as characterized by a covariance matrix, are treated as the sum of N n-dimensional random vectors; N obeys the negative binomial distribution with a parameter alpha and mean N-bar. Subsequently, an n-dimensional K-distribution, with either zero or nonzero mean, is developed in the limit of infinite N-bar or illuminated area. The probability density function (PDF) of the K-distributed vector normalized by its Euclidean norm is independent of the parameter alpha and is the same as that derived from a zero-mean Gaussian-distributed random vector.

Yueh, S. H.↗

Parametric Analysis of a Hover Test Vehicle using Advanced Test Generation and Data Analysis

Large complex aerospace systems are generally validated in regions local to anticipated operating points rather than through characterization of the entire feasible operational envelope of the system. This is due to the large parameter space, and complex, highly coupled nonlinear nature of the different systems that contribute to the performance of the aerospace system. We have addressed the factors deterring such an analysis by applying a combination of technologies to the area of flight envelop assessment. We utilize n-factor (2,3) combinatorial parameter variations to limit the number of cases, but still explore important interactions in the parameter space in a systematic fashion. The data generated is automatically analyzed through a combination of unsupervised learning using a Bayesian multivariate clustering technique (AutoBayes) and supervised learning of critical parameter ranges using the machine-learning tool TAR3, a treatment learner. Covariance analysis with scatter plots and likelihood contours are used to visualize correlations between simulation parameters and simulation results, a task that requires tool support, especially for large and complex models. We present results of simulation experiments for a cold-gas-powered hover test vehicle.

Gundy-Burlet, Karen↗

Fast Multivariate Search on Large Aviation Datasets

Multivariate Time-Series (MTS) are ubiquitous, and are generated in areas as disparate as sensor recordings in aerospace systems, music and video streams, medical monitoring, and financial systems. Domain experts are often interested in searching for interesting multivariate patterns from these MTS databases which can contain up to several gigabytes of data. Surprisingly, research on MTS search is very limited. Most existing work only supports queries with the same length of data, or queries on a fixed set of variables. In this paper, we propose an efficient and flexible subsequence search framework for massive MTS databases, that, for the first time, enables querying on any subset of variables with arbitrary time delays between them. We propose two provably correct algorithms to solve this problem (1) an R-tree Based Search (RBS) which uses Minimum Bounding Rectangles (MBR) to organize the subsequences, and (2) a List Based Search (LBS) algorithm which uses sorted lists for indexing. We demonstrate the performance of these algorithms using two large MTS databases from the aviation domain, each containing several millions of observations Both these tests show that our algorithms have very high prune rates (>95%) thus needing actual

Bhaduri, Kanishka↗

New Constraints on Titan's Stratospheric n-Butane Abundance

Curiously, n-butane has yet to be detected at Titan, though it is predicted to be present in a wide range of abundances that span over 2.5 orders of magnitude. We have searched infrared spectroscopic observations of Titan for signals from n-butane (n-C4H10) in Titan's stratosphere. Three sets of Cassini Composite Infrared Spectrometer Focal Plane 4 (1050–1500 cm−1) observations were selected for modeling, having been collected from different flybys and pointing latitudes. We modeled the observations with the Nonlinear Optimal Estimator for MultivariatE Spectral AnalySIS radiative transfer tool. Temperature profiles were retrieved for each of the data sets by modeling the ν4 emission from methane near 1305 cm−1. Then, incorporating the temperature profiles, we retrieved abundances of all of Titan's known trace gases that are active in this spectral region, reliably reproducing the observations. We then systematically tested a set of models with varying abundances of n-butane, investigating how the addition of this gas affected the fits. We did this for several different photochemically predicted abundance profiles from the literature, as well as for a constant-with-altitude profile. Ultimately, though we did not produce any firm detection of n-butane, we derived new upper limits on its abundance specific to the use of each profile and to multiple different ranges of stratospheric altitudes. These results will tightly constrain the C4 chemistry of future photochemical modeling of Titan's atmosphere and also motivate the continued search for n-butane and its isomer, isobutane.

Brendan L. Steffens↗