Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “multivariate data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Nearest-Neighbor Machine Learning Feature Selection for Interpretation of Microbial Molecular Signatures from Isotope Ratio Mass Spectrometry Data

Mass spectrometry (MS) promises to be a powerful tool for potential biosignature detection during astrobiological missions on ocean worlds in our solar system. Accurate and generalizable machine learning methods could enhance science return on investment by predicting seawater chemistry and classifying isotopic biosignatures, either as a signature consistent with microbial life (biotic) or as a novelty (unclassified/unique). However, machine learning models are likely to be complex and involve interactions between MS features, making biosignatures difficult to interpret. Feature selection methods provide biological and chemical context that help interpret the mechanisms of machine learning models, but these methods also need the ability to detect complex interactions. Previously, we developed a machine learning feature selection algorithm called nearest-neighbor projected distance regression (NPDR) that has the ability to identify important model features that involve complex interactions and automatically reduce correlation and the dimensionality in a high-dimensional variable space. The standard distance metrics used in NPDR – Manhattan and Euclidean – assume the multivariate data are isotropic, which is often violated in real data due to differences in the covariance between variables. Thus, we extend NPDR to include a random forest distance, and other anisotropic distance metrics, for computing nearest neighbors. We also augment the isotope-ratio MS data with time-series features from the raw MS signal to improve biotic classification. We test NPDR on our novel experimental ocean world seawater analog MS data. We measure isotope fractionations of volatile CO 2 that could be measured in exospheres or plumes. Samples include baseline abiotic conditions using a range of possible seawater chemistry consistent with Europa and Enceladus, and biotic samples that include microbes in these seawaters. We use penalized NPDR with random forest proximity to identify interpretable microbial molecular signatures. We compare features with random forest importance, and we train a classifier that discriminates between biotic and abiotic samples with high accuracy. These ML-trained ocean-world analog MS data could be used to assist in identifying biosignatures during future missions.

geochemistry↗

Validation of MODIS FLH and In Situ Chlorophyll a from Tampa Bay, Florida (USA)

Satellite observation of phytoplankton concentration or chlorophyll-a (chla) is an important characteristic, critically integral to monitoring coastal water quality. However, the optical properties of estuarine and coastal waters are highly variable and complex and pose a great challenge for accurate analysis. Constituents such as suspended solids and dissolved organic matter and the overlapping and uncorrelated absorptions in the blue region of the spectrum renders the blue-green ratio algorithms for estimating chl-a inaccurate. Measurement of suninduced chlorophyll fluorescence, on the other hand, which utilizes the near infrared portion of the electromagnetic spectrum may, provide a better estimate of phytoplankton concentrations. While modelling and laboratory studies have illustrated both the utility and limitations of satellite algorithms based on the sun induced chlorophyll fluorescence signal, few have examined the empirical validity of these algorithms or compared their accuracy against bluegreen ratio algorithms . In an unprecedented analysis using a long term (2003-2011) in situ monitoring data set from Tampa Bay, Florida (USA), we assess the validity of the FLH product from the Moderate Resolution Imaging Spectrometer against a suite of water quality parameters taken in a variety of conditions throughout this large optically complex estuarine system. . Overall, the results show a 106% increase in the validity of chla concentration estimation using FLH over the standard chla estimate from the blue-green OC3M algorithm. Additionally, a systematic analysis of sampling sites throughout the bay is undertaken to understand how the FLH product responds to varying conditions in the estuary and correlations are conducted to see how the relationships between satellite FLH and in situ chlorophyll-a change with depth, distance from shore, from structures like bridges, and nutrient concentrations and turbidity. Such analysis illustrates that the correlations between FLH and in situ chla measurements increases with increasing distance between monitoring sites and structures like bridges and shore. Due probably to confounding factors, expected improvement in the FLH- chla relationship was not clearly noted when increasing depth and distance from shore alone (not including bridges). Correlations between turbidity and nutrient concentrations are discussed further and principle component analyses are employed to address the relationships between the multivariate data sets. A thorough understanding of how satellite FLH algorithms relate to in situ water quality parameters will enhance our understanding of how MODIS s global FLH algorithm can be used empirically to monitor coastal waters worldwide.

Fischer, Andrew↗

A method of determining spectral dye densities in color films

A mathematical analysis technique called characteristic vector analysis, reported by Simonds (1963), is used to determine spectral dye densities in multiemulsion film such as color or color-IR imagery. The technique involves examining a number of sets of multivariate data and determining linear transformations of these data to a smaller number of parameters which contain essentially all of the information contained in the original set of data. The steps involved in the actual procedure are outlined. It is shown that integral spectral density measurements of a large number of different color samples can be accurately reconstructed from the calculated spectral dye densities.

Friederichs, G. A.↗

Computer program documentation: ISOCLS iterative self-organizing clustering program, program C094

The author has identified the following significant results. This program implements an algorithm which, ideally, sorts a given set of multivariate data points into similar groups or clusters. The program is intended for use in the evaluation of multispectral scanner data; however, the algorithm could be used for other data types as well. The user may specify a set of initial estimated cluster means to begin the procedure, or he may begin with the assumption that all the data belongs to one cluster. The procedure is initiatized by assigning each data point to the nearest (in absolute distance) cluster mean. If no initial cluster means were input, all of the data is assigned to cluster 1. The means and standard deviations are calculated for each cluster.

Minter, R. T.↗

A fast routine for computing

A routine for calculating multidimensional histograms of multivariate data using a combination table look up and search procedure is described. The software was originally developed to computer four-dimensional histograms from LANDSAT multispectral imagery, but the concept can be used on other types of data and the program can be modified for the desired type of output information.

Jayroe, R. R., Jr.↗

Computer program documentation for the patch subsampling processor

The programs presented are intended to provide a way to extract a sample from a full-frame scene and summarize it in a useful way. The sample in each case was chosen to fill a 512-by-512 pixel (sample-by-line) image since this is the largest image that can be displayed on the Integrated Multivariant Data Analysis and Classification System. This sample size provides one megabyte of data for manipulation and storage and contains about 3% of the full-frame data. A patch image processor computes means for 256 32-by-32 pixel squares which constitute the 512-by-512 pixel image. Thus, 256 measurements are available for 8 vegetation indexes over a 100-mile square.

Nieves, M. J.↗

Configuration space representation in parallel coordinates

By means of a system of parallel coordinates, a nonprojective mapping from R exp N to R squared is obtained for any positive integer N. In this way multivariate data and relations can be represented in the Euclidean plane (embedded in the projective plane). Basically, R squared with Cartesian coordinates is augmented by N parallel axes, one for each variable. The N joint variables of a robotic device can be represented graphically by using parallel coordinates. It is pointed out that some properties of the relation are better perceived visually from the parallel coordinate representation, and that new algorithms and data structures can be obtained from this representation. The main features of parallel coordinates are described, and an example is presented of their use for configuration space representation of a mechanical arm (where Cartesian coordinates cannot be used).

Fiorini, Paolo↗

User's manual for the Gaussian windows program

'Gaussian Windows' is a method for exploring a set of multivariate data, in order to estimate the shape of the underlying density function. The method can be used to find and describe structural features in the data. The method is described in two earlier papers. I assume that the reader has access to both of these papers, so I will not repeat material from them. The program described herein is written in BASIC and it runs on an IBM PC or PS/2 with the DOS 3.3 operating system. Although the program is slow and has limited memory space, it is adequate for experimenting with the method. Since it is written in BASIC, it is relatively easy to modify. The program and some related files are available on a 3-inch diskette. A listing of the program is also available. This user's manual explains the use of the program. First, it gives a brief tutorial, illustrating some of the program's features with a set of artificial data. Then, it describes the results displayed after the program does a Gaussian window, and it explains each of the items on the various menus.

Jaeckel, Louis A.↗

The Grand Tour via Geodesic Interpolation of 2-frames

Grand tours are a class of methods for visualizing multivariate data, or any finite set of points in n-space. The idea is to create an animation of data projections by moving a 2-dimensional projection plane through n-space. The path of planes used in the animation is chosen so that it becomes dense, that is, it comes arbitrarily close to any plane. One of the original inspirations for the grand tour was the experience of trying to comprehend an abstract sculpture in a museum. One tends to walk around the sculpture, viewing it from many different angles. A useful class of grand tours is based on the idea of continuously interpolating an infinite sequence of randomly chosen planes. Visiting randomly (more precisely: uniformly) distributed planes guarantees denseness of the interpolating path. In computer implementations, 2-dimensional orthogonal projections are specified by two 1-dimensional projections which map to the horizontal and vertical screen dimensions, respectively. Hence, a grand tour is specified by a path of pairs of orthonormal projection vectors. This paper describes an interpolation scheme for smoothly connecting two pairs of orthonormal vectors, and thus for constructing interpolating grand tours. The scheme is optimal in the sense that connecting paths are geodesics in a natural Riemannian geometry.

Asimov, Daniel↗

Logistic Risk Model for the Unique Effects of Inherent Aerobic Capacity on (+)G(sub z) Tolerance Before and After Simulated Weightlessness

Small sample size (n less than 1O) and inappropriate analysis of multivariate data have hindered previous attempts to describe which physiologic and demographic variables are most important in determining how long humans can tolerate acceleration. Data from previous centrifuge studies conducted at NASA/Ames Research Center, utilizing a 7-14 d bed rest protocol to simulate weightlessness, were included in the current investigation. After review, data on 25 women and 22 men were available for analysis. Study variables included gender, age, weight, height, percent body fat, resting heart rate, mean arterial pressure, Vo(sub 2)max and plasma volume. Since the dependent variable was time to greyout (failure), two contemporary biostatistical modeling procedures (proportional hazard and logistic discriminant function) were used to estimate risk, given a particular subject's profile. After adjusting for pro-bed-rest tolerance time, none of the profile variables remained in the risk equation for post-bed-rest tolerance greyout. However, prior to bed rest, risk of greyout could be predicted with 91% accuracy. All of the profile variables except weight, MAP, and those related to inherent aerobic capacity (Vo(sub 2)max, percent body fat, resting heart rate) entered the risk equation for pro-bed-rest greyout. A cross-validation using 24 new subjects indicated a very stable model for risk prediction, accurate within 5% of the original equation. The result for the inherent fitness variables is significant in that a consensus as to whether an increased aerobic capacity is beneficial or detrimental has not been satisfactorily established. We conclude that tolerance to +Gz acceleration before and after simulated weightlessness is independent of inherent aerobic fitness.

Ludwig, David A.↗

Low power, lightweight vapor sensing using arrays of conducting polymer composite chemically-sensitive resistors

Arrays of broadly responsive vapor detectors can be used to detect, identify, and quantify vapors and vapor mixtures. One implementation of this strategy involves the use of arrays of chemically-sensitive resistors made from conducting polymer composites. Sorption of an analyte into the polymer composite detector leads to swelling of the film material. The swelling is in turn transduced into a change in electrical resistance because the detector films consist of polymers filled with conducting particles such as carbon black. The differential sorption, and thus differential swelling, of an analyte into each polymer composite in the array produces a unique pattern for each different analyte of interest, Pattern recognition algorithms are then used to analyze the multivariate data arising from the responses of such a detector array. Chiral detector films can provide differential detection of the presence of certain chiral organic vapor analytes. Aspects of the spaceflight qualification and deployment of such a detector array, along with its performance for certain analytes of interest in manned life support applications, are reviewed and summarized in this article.

NASA Discipline Life Sciences Technologies↗

Interval Predictor Models for Robust System Identification

This paper proposes a framework for the identification and uncertainty quantification of plant models according to multivariable data. The only restriction imposed upon such models is for their outputs to depend continuously on their parameters. An Interval Predictor Model (IPM) prescribes the parameters of a computational model as a path-connected set thereby making each predicted output an interval-valued function of its inputs. The formulation proposed seeks the parameter set for which the predicted outputs tightly enclose the data. This set, which is modeled as a semi-algebraic set of low-degree polynomials, enables the characterization of possibly strong parameter dependencies commonly found in practice. This uncertainty characterization makes the resulting plant model amenable to robust control approaches using polynomial optimization. Furthermore, we use non-convex scenario theory to assess the reliability of the resulting IPM. This assessment yields a distribution-free upper bound on the probability that future data will fall outside the predicted intervals.

interval↗

Development of A Multidecadal Land Reanalysis Over High Mountain Asia

Anthropogenic and climatic changes affect the water and energy cycles in High Mountain Asia (HMA), home to over two billion people and the largest reservoirs of freshwater outside the polar zone. Despite their significant importance for water management, consistent and reliable estimates of water storage and fluxes over the region are lacking because of the high uncertainties associated with the estimates of atmospheric conditions and human management. Here, we relied on multivariate data assimilation (MVDA) to provide estimates of energy and water storage and fluxes that reflect the processes occurring in the region such as greening and irrigation-driven groundwater depletion. We developed and employed an ensemble precipitation estimate by blending different precipitation products thereby reducing the uncertainties and inconsistencies associated with precipitation in HMA. Then, we assimilated five variables that capture the changes in hydrology in response to climate change and anthropogenic activities. Overall, our results have shown that MVDA has allowed a better representation of the land surface processes including greening and irrigation-driven groundwater depletion in HMA.

Fadji Z. Maina↗

Identification of multivariable high performance turbofan engine dynamics from closed loop data

The multivariable instrumental variable/approximate maximum likelihood (IV/AML) method or recursive time-series analysis is used to identify the multivariable (four inputs-three outputs) dynamics of the Pratt and Whitney F100 engine. A detailed nonlinear engine simulation is used to determine linear engine model structures and parameters at an operating point using open loop data. Also, the IV/AML method is used in a direct identification mode to identify models from actual closed loop engine test data. Models identified from simulated and test data are compared to determine a final model structure and parameterization that can predict engine response for a wide class of inputs. The ability of the IV/AML algorithm to identify useful dynamic models from engine test data is assessed.

Merrill, W.↗

Identification of multivariable high performance turbofan engine dynamics from closed loop data

The multivariable instrumental variable/approximate maximum likelihood (IV/AML) method of recursive time-series analysis is used to identify the multivariable (four inputs-three outputs) dynamics of the Pratt and Whitney F100 engine. A detailed nonlinear engine simulation is used to determine linear engine model structures and parameters at an operating point using open loop data. Also, the IV/AML method is used in a direct identification made to identify models from actual closed loop engine test data. Models identified from simulated and test data are compared to determine a final model structure and parameterization that can predict engine response for a wide class of inputs. The ability of the IV/AML algorithm to identify useful dynamic models from engine test data is assessed. Previously announced in STAR as N82-20339

Merrill, W.↗

Linear Multivariable Regression Models for Prediction of Eddy Dissipation Rate from Available Meteorological Data

Linear multivariable regression models for predicting day and night Eddy Dissipation Rate (EDR) from available meteorological data sources are defined and validated. Model definition is based on a combination of 1997-2000 Dallas/Fort Worth (DFW) data sources, EDR from Aircraft Vortex Spacing System (AVOSS) deployment data, and regression variables primarily from corresponding Automated Surface Observation System (ASOS) data. Model validation is accomplished through EDR predictions on a similar combination of 1994-1995 Memphis (MEM) AVOSS and ASOS data. Model forms include an intercept plus a single term of fixed optimal power for each of these regression variables; 30-minute forward averaged mean and variance of near-surface wind speed and temperature, variance of wind direction, and a discrete cloud cover metric. Distinct day and night models, regressing on EDR and the natural log of EDR respectively, yield best performance and avoid model discontinuity over day/night data boundaries.

MCKissick, Burnell T.↗

A Network-Based Algorithm for Clustering Multivariate Repeated Measures Data

The National Aeronautics and Space Administration (NASA) Astronaut Corps is a unique occupational cohort for which vast amounts of measures data have been collected repeatedly in research or operational studies pre-, in-, and post-flight, as well as during multiple clinical care visits. In exploratory analyses aimed at generating hypotheses regarding physiological changes associated with spaceflight exposure, such as impaired vision, it is of interest to identify anomalies and trends across these expansive datasets. Multivariate clustering algorithms for repeated measures data may help parse the data to identify homogeneous groups of astronauts that have higher risks for a particular physiological change. However, available clustering methods may not be able to accommodate the complex data structures found in NASA data, since the methods often rely on strict model assumptions, require equally-spaced and balanced assessment times, cannot accommodate missing data or differing time scales across variables, and cannot process continuous and discrete data simultaneously. To fill this gap, we propose a network-based, multivariate clustering algorithm for repeated measures data that can be tailored to fit various research settings. Using simulated data, we demonstrate how our method can be used to identify patterns in complex data structures found in practice.

Koslovsky, Matthew↗

A distributed system for visualizing and analyzing multivariate and multidisciplinary data

The Linked Windows Interactive Data System (Link Winds) is being developed with NASA support. The objective of this proposal is to adapt and apply that system in a complex network environment containing elements to be found by scientists working multidisciplinary teams on very large scale and distributed data sets. The proposed three year program will develop specific visualization and analysis tools, to be exercised locally and remotely in the Link Winds environment, to demonstrate visual data analysis, interdisciplinary data analysis and cooperative and interactive televisualization and analysis of data by geographically separated science teams. These demonstrations will involve at least two science disciplines with the aim of producing publishable results.

Jacobson, Allan S.↗