Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “multivariate data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Statistical methods and neural network approaches for classification of data from multiple sources

Statistical methods for classification of data from multiple data sources are investigated and compared to neural network models. A problem with using conventional multivariate statistical approaches for classification of data of multiple types is in general that a multivariate distribution cannot be assumed for the classes in the data sources. Another common problem with statistical classification methods is that the data sources are not equally reliable. This means that the data sources need to be weighted according to their reliability but most statistical classification methods do not have a mechanism for this. This research focuses on statistical methods which can overcome these problems: a method of statistical multisource analysis and consensus theory. Reliability measures for weighting the data sources in these methods are suggested and investigated. Secondly, this research focuses on neural network models. The neural networks are distribution free since no prior knowledge of the statistical distribution of the data is needed. This is an obvious advantage over most statistical classification methods. The neural networks also automatically take care of the problem involving how much weight each data source should have. On the other hand, their training process is iterative and can take a very long time. Methods to speed up the training procedure are introduced and investigated. Experimental results of classification using both neural network models and statistical methods are given, and the approaches are compared based on these results.

Benediktsson, Jon Atli↗

Model-Agnostic Algorithm for Real-Time Attack Identification in Power Grid using Koopman Modes

Malicious activities on measurements from sensors like Phasor Measurement Units (PMUs) can mislead the control center operator into taking wrong control actions resulting in disruption of operation, financial losses, and equipment damage. In particular, false data attacks initiated during power systems transients caused due to abrupt changes in load and generation can fool the conventional model-based detection methods relying on thresholds comparison to trigger an anomaly. In this paper, we propose a Koopman mode decomposition (KMD) based algorithm to detect and identify false data attacks in real-time. The Koopman modes (KMs) are capable of capturing the nonlinear modes of oscillation in the transient dynamics of the power networks and reveal the spatial embedding of both natural and anomalous modes of oscillations in the sensor measurements. The Koopman-based spatio-temporal nonlinear modal analysis is used to filter out the false data injected by an attacker. The performance of the algorithm is illustrated on the IEEE 68-bus test system using synthetic attack scenarios generated on GridSTAGE, a recently developed multivariate spatio-temporal data generation framework for simulation of adversarial scenarios in cyber-physical power systems.

Nandanoori, Sai Pushpak↗

Identification of multivariable high-performance turbofan engine dynamics from closed loop data

A typical engine control design cycle consists of developing a dynamic engine simulation from steady-state component performance data, designing a control based upon this simulation, and then testing and modifying the control in an engine test cell to meet performance requirements. This design cycle was successful for state-of-the-art engines. However, for more advanced multivariable engines that exhibit strong variable interactions, this procedure will result in substantial trial and error modification of the control during the testing phase. One method to automate the design process and reduce control modification testing and development cost would be to identify accurate dynamic models directly from the closed-loop test data. These identified models would then be used in conjunction with a synthesis procedure to systematically refine the control. Recent advances in closed-loop identifiability present a methodology for this direct identification of engine model dynamics from closed-loop test data. The application of an identification method to simulated and actual closed-loop F100 engine data is described. This study was undertaken to determine if useful dynamic engine models could be identified directly from closed-loop engine test data.

Merrill, W. C.↗

A Graphical Model for Fusing Diverse Microbiome Data

This paper develops a Bayesian graphical model for fusing disparate types of count data. The motivating application is the study of bacterial communities from diverse high-dimensional features, in this case, transcripts, collected from different treatments. In such datasets, there are no explicit correspondences between the communities and each corresponds to different factors, making data fusion challenging. We introduce a flexible multinomial-Gaussian generative model for jointly modeling such count data. This latent variable model jointly characterizes the observed data through a common multivariate Gaussian latent space that parameterizes the set of multinomial probabilities of the transcriptome counts. The covariance matrix of the latent variables induces a covariance matrix of co-dependencies between all the transcripts, effectively fusing multiple data sources. We present a computationally scalable variational Expectation-Maximization (EM) algorithm for inferring the latent variables and the parameters of the model. Here, the inferred latent variables provide a common dimensionality reduction for visualizing the data and the inferred parameters provide a predictive posterior distribution. In addition to simulation studies that demonstrate the variational EM procedure, we apply our model to a bacterial microbiome dataset.

59 BASIC BIOLOGICAL SCIENCES↗

Rapid measurement of soluble xylo-oligomers using near-infrared spectroscopy (NIRS) and multivariate statistics: calibration model development and practical approaches to model optimization

Rapid monitoring of biomass conversion processes using techniques such as near-infrared (NIR) spectroscopy can be substantially quicker and less labor-, resource-, and energy-intensive than conventional measurement techniques such as gas or liquid chromatography (GC or LC) due to the lack of solvents and preparation methods, as well as removing the need to transfer samples to an external lab for analytical evaluation. The purpose of this study was to determine the feasibility of rapid monitoring of a biomass conversion process using NIR spectroscopy combined with multivariate statistical modeling, and to examine the impact of (1) subsetting the samples in the original dataset by process location and (2) reducing the spectral range used in the calibration model on model performance. We develop multivariate calibration models for the concentrations of soluble xylo-oligosaccharides (XOS), monomeric xylose, and total solids at multiple points in a biomass conversion process which produces and then purifies XOS compounds from sugar cane bagasse. A single model using samples from multiple locations in the process stream showed acceptable performance as measured by standard statistical measures. However, compared to the single model, we show that separate models built by segregating the calibration samples according to process location show improved performance. We also show that combining an understanding of the sample spectra with simple multivariate analysis tools can result in a calibration model with a substantially smaller spectral range that provides essentially equal performance to the full-range model. We demonstrate that real-time monitoring of soluble xylo-oligosaccharides (XOS), monomeric xylose, and total solids concentration at multiple points in a process stream using NIR spectroscopy coupled with multivariate statistics is feasible. Segregation of sample populations by process location improves model performance. Models using a reduced spectral range containing the most relevant spectral signatures show very similar performance to the full-range model, reinforcing the importance of performing robust exploratory data analysis before beginning multivariate modeling.

09 BIOMASS FUELS↗

Regression Model Optimization for the Analysis of Experimental Data

A candidate math model search algorithm was developed at Ames Research Center that determines a recommended math model for the multivariate regression analysis of experimental data. The search algorithm is applicable to classical regression analysis problems as well as wind tunnel strain gage balance calibration analysis applications. The algorithm compares the predictive capability of different regression models using the standard deviation of the PRESS residuals of the responses as a search metric. This search metric is minimized during the search. Singular value decomposition is used during the search to reject math models that lead to a singular solution of the regression analysis problem. Two threshold dependent constraints are also applied. The first constraint rejects math models with insignificant terms. The second constraint rejects math models with near-linear dependencies between terms. The math term hierarchy rule may also be applied as an optional constraint during or after the candidate math model search. The final term selection of the recommended math model depends on the regressor and response values of the data set, the user s function class combination choice, the user s constraint selections, and the result of the search metric minimization. A frequently used regression analysis example from the literature is used to illustrate the application of the search algorithm to experimental data.

Ulbrich, N.↗

Aeroservoelastic Modeling and Validation of a Thrust-Vectoring F/A-18 Aircraft

An F/A-18 aircraft was modified to perform flight research at high angles of attack (AOA) using thrust vectoring and advanced control law concepts for agility and performance enhancement and to provide a testbed for the computational fluid dynamics community. Aeroservoelastic (ASE) characteristics had changed considerably from the baseline F/A-18 aircraft because of structural and flight control system amendments, so analyses and flight tests were performed to verify structural stability at high AOA. Detailed actuator models that consider the physical, electrical, and mechanical elements of actuation and its installation on the airframe were employed in the analysis to accurately model the coupled dynamics of the airframe, actuators, and control surfaces. This report describes the ASE modeling procedure, ground test validation, flight test clearance, and test data analysis for the reconfigured F/A-18 aircraft. Multivariable ASE stability margins are calculated from flight data and compared to analytical margins. Because this thrust-vectoring configuration uses exhaust vanes to vector the thrust, the modeling issues are nearly identical for modem multi-axis nozzle configurations. This report correlates analysis results with flight test data and makes observations concerning the application of the linear predictions to thrust-vectoring and high-AOA flight.

Brenner, Martin J.↗

Online evolutionary neural architecture search for multivariate non-stationary time series forecasting

Time series forecasting (TSF) is one of the most important tasks in data science. TSF models are usually pre-trained with historical data and then applied on future unseen datapoints. However, real-world time series data is usually non-stationary and models trained offline usually face problems from data drift. Models trained and designed in an offline fashion can not quickly adapt to changes quickly or be deployed in real-time. To address these issues, this work presents the Online NeuroEvolution-based Neural Architecture Search (ONE-NAS) algorithm, which is a novel neural architecture search method capable of automatically designing and dynamically training recurrent neural networks (RNNs) for online forecasting tasks. Without any pre-training, ONE-NAS utilizes populations of RNNs that are continuously updated with new network structures and weights in response to new multivariate input data. ONE-NAS is tested on real-world, large-scale multivariate wind turbine data as well as the univariate Dow Jones Industrial Average (DJIA) dataset. These results demonstrate that ONE-NAS outperforms traditional statistical time series forecasting methods, including online linear regression, fixed long short-term memory (LSTM) and gated recurrent unit (GRU) models trained online, as well as state-of-the-art, online ARIMA strategies. Additionally, results show that utilizing multiple populations of RNNs which are periodically repopulated provide significant performance improvements, allowing this online neural network architecture design and training to be successful.

97 MATHEMATICS AND COMPUTING↗

Measuring watershed runoff capability with ERTS data

Parameters of most equations used to predict runoff from an ungaged area are based on characteristics of the watershed and subject to the biases of a hydrologist. Digital multispectral scanner, MSS, data from ERTS was reduced with the aid of computer programs and a Dicomed display. Multivariate analyses of the MSS data indicate that discrimination between watersheds with different runoff capabilities is possible using ERTS data. Differences between two visible bands of MSS data can be used to more accurately evaluate the parameters than present subjective methods, thus reducing construction cost due to overdesign of flood detention structures.

Blanchard, B. J.↗

Areal Distribution of the Oxygen-Isotope Ratio in Greenland

Mean values of the oxygen-isotope ratio relative to standard mean ocean water reported for 46 sites on the Greenland ice sheet are compiled together with data on mean annual surface temperature, latitude, 6180 elevation, and mean annual shortest distance to the open ocean denoted by the 10% sea-ice concentration boundary. Stepwise regression analyses, with 6180 as the dependent variable, define two robust models. In the forward mode at the 99.9% confidence level, only temperature enters the model. In the backward mode at the 95% confidence level, only temperature, latitude, and distance to the open ocean remain in the model. Inversions of the models on the basis of 160 gridpoint locations 100 km apart in the area delimited by the surface equilibrium line produce four contoured distributions of 6"0. Two distributions are based on the bivariate model and two on the multivariate model. The second distribution for each model is obtained substituting mean annual surface-temperature values obtained from the Nimbus-7 Temperature Humidity Infrared Radiometer (THIR) database. All four distributions are considered valid, and differences between them are evaluated using contoured anomaly maps. It is suggested that the inversion of the multivariate model using THIR data provides the more reliable pattern for studies of atmospheric advection or for the derivation of ice-flow adjustments for 6180 series obtained from deep-core or ablation-zone sites.

Zwally, H. Jay↗

MAD: Self-Supervised Masked Anomaly Detection Task for Multivariate Time Series

In this paper, we introduce Masked Anomaly Detection (MAD), a general self-supervised learning task for multivariate time series anomaly detection. With the increasing availability of sensor data from industrial systems, being able to detecting anomalies from streams of multivariate time series data is of significant importance. Given the scarcity of anomalies in real-world applications, the majority of literature has been focusing on modeling normality. The learned normal representations can empower anomaly detection as the model has learned to capture certain key underlying data regularities. A typical formulation is to learn a predictive model, i.e., use a window of time series data to predict future data values. In this paper, we propose an alternative self-supervised learning task. By randomly masking a portion of the inputs and training a model to estimate them using the remaining ones, MAD is an improvement over the traditional left-to-right next step prediction (NSP) task. Our experimental results demonstrate that MAD can achieve better anomaly detection rates over traditional NSP approaches when using exactly the same neural network (NN) base models, and can be modified to run as fast as NSP models during test time on the same hardware, thus making it an ideal upgrade for many existing NSP-based NN anomaly detection models.

97 MATHEMATICS AND COMPUTING↗

Accelerated Probabilistic Marching Cubes by Deep Learning for Time-Varying Scalar Ensembles

Visualizing the uncertainty of ensemble simulations is challenging due to the large size and multivariate and temporal features of en-semble data sets. One popular approach to studying the uncertainty of ensembles is analyzing the positional uncertainty of the level sets. Probabilistic marching cubes is a technique that performs Monte Carlo sampling of multivariate Gaussian noise distributions for positional uncertainty visualization of level sets. However, the technique suffers from high computational time, making interactive visualization and analysis impossible to achieve. This paper introduces a deep-learning-based approach to learning the level-set uncertainty for two-dimensional ensemble data with a multivariate Gaussian noise assumption. We train the model using the first few time steps from time-varying ensemble data in our workflow. We demonstrate that our trained model accurately infers uncertainty in level sets for new time steps and is up to 170X faster than that of the original probabilistic model with serial computation and 10X faster than that of the original parallel computation.

Han, Mengjiao↗

Latent-Space Dynamics for Prediction and Fault Detection in Geothermal Power Plant Operations

This paper presents a latent-space dynamic neural network (LSDNN) model for the multi-step-ahead prediction and fault detection of a geothermal power plant’s operation. The model was trained to learn the dynamics of the power generation process from multivariate time-series data and the effects of exogenous variables, such as control adjustment and ambient temperature. In the LSDNN model, an encoder–decoder architecture was designed to capture cross-correlation among different measured variables. In addition, a latent space dynamic structure was proposed to propagate the dynamics in the latent space to enable prediction. The prediction power of the LSDNN was utilized for monitoring a geothermal power plant and detecting abnormal events. The model was integrated with principal component analysis (PCA)-based process monitoring techniques to develop a fault-detection procedure. The performance of the proposed LSDNN model and fault detection approach was demonstrated using field data collected from a geothermal power plant.

15 GEOTHERMAL ENERGY↗

Vegetation monitoring and classification using NOAA/AVHRR satellite data

A vegetation gradient model, based on a new surface hydrologic index and NOAA/AVHRR meteorological satellite data, has been analyzed along a 1300 km east-west transect across the state of Texas. The model was developed to test the potential usefulness of such low-resolution data for vegetation stratification and monitoring. Normalized Difference values (ratio of AVHRR bands 1 and 2, considered to be an index of greenness) were determined and evaluated against climatological and vegetation characteristics at 50 sample locations (regular intervals of 0.25 deg longitude) along the transect on five days in 1980. Statistical treatment of the data indicate that a multivariate model incorporating satellite-measured spectral greenness values and a surface hydrologic factor offer promise as a new technique for regional-scale vegetation stratification and monitoring.

Greegor, D. H., Jr.↗

Modeling and managing risk early in software development

In order to improve the quality of the software development process, we need to be able to build empirical multivariate models based on data collectable early in the software process. These models need to be both useful for prediction and easy to interpret, so that remedial actions may be taken in order to control and optimize the development process. We present an automated modeling technique which can be used as an alternative to regression techniques. We show how it can be used to facilitate the identification and aid the interpretation of the significant trends which characterize 'high risk' components in several Ada systems. Finally, we evaluate the effectiveness of our technique based on a comparison with logistic regression based models.

Briand, Lionel C.↗

A multiparametric analysis of the Einstein sample of early-type galaxies. 1: Luminosity and ISM parameters

We have conducted bivariate and multivariate statistical analysis of data measuring the luminosity and interstellar medium of the Einstein sample of early-type galaxies (presented by Fabbiano, Kim, & Trinchieri 1992). We find a strong nonlinear correlation between L(sub B) and L(sub X), with a power-law slope of 1.8 +/- 0.1, steepening to 2.0 +/- if we do not consider the Local Group dwarf galaxies M32 and NGC 205. Considering only galaxies with log L(sub X) less than or equal to 40.5, we instead find a slope of 1.0 +/- 0.2 (with or without the Local Group dwarfs). Although E and S0 galaxies have consistent slopes for their L(sub B)-L(sub X) relationships, the mean values of the distribution functions of both L(sub X) and L(sub X)/L(sub B) for the S0 galaxies are lower than those for the E galaxies at the 2.8 sigma and 3.5 sigma levels, respectively. We find clear evidence for a correlation between L(sub X) and the X-ray color C(sub 21), defined by Kim, Fabbiano, & Trinchieri (1992b), which indicates that X-ray luminosity is correlated with the spectral shape below 1 keV in the sense that low-L(sub X) systems have relatively large contributions from a soft component compared with high-L(sub X) systems. We find evidence from our analysis of the 12 micron IRAS data for our sample that our S0 sample has excess 12 micron emission compared with the E sample, scaled by their optical luminosities. This may be due to emission from dust heated in star-forming regions in S0 disks. This interpretation is reinforced by the existence of a strong L(sub 12)-L(sub 100) correlation for our S0 sample that is not found for the E galaxies, and by an analysis of optical-IR colors. We find steep slopes for power-law relationships between radio luminosity and optical, X-ray, and far-IR (FIR) properties. This last point argues that the presence of an FIR-emitting interstellar medium (ISM) in early-type galaxies is coupled to their ability to generate nonthermal radio continuum, as previously argued by, e.g., Walsh et al. (1989). We also find that, for a given L(sub 100), galaxies with larger L(sub X)/L(sub B) tend to be stronger nonthermal radio sources, as originally suggested by Kim & Fabbiano (1990). We note that, while L(sub B) is most strongly correlated with L(sub 6), the total radio luminosity, both L(sub X) and L(sub X)/L(sub B) are more strongly correlated with L(sub 6 CO), the core radio luminosity. These points support the argument (proposed by Fabbiano, Gioia, & Trinchieri 1989) that radio cores in early-type galaxies are fueled by the hot ISM.

Eskridge, Paul B.↗

A multiparametric analysis of the Einstein sample of early-type galaxies. 2: Galaxy formation history and properties of the interstellar medium

We have conducted bivariate and multivariate statistical analysis of data measuring the integrated luminosity, shape, and potential depth of the Einstein sample of early-type galaxies (presented by Fabbiano et al. 1992). We find significant correlations between the X-ray properties and the axial ratios (a/b) of our sample, such that the roundest systems tend to have the highest L(sub x) and L(sub x)/L(sub B). The most radio-loud objects are also the roundest. We confirm the assertion of Bender et al. (1989) that galaxies with high L(sub x) are boxy (have negative a(sub 4)). Both a/b and a(sub 4) are correlated with L(sub B), but not with IRAS 12 um and 100 um luminosities. There are strong correlations between L(sub x), Mg(sub 2), and sigma(sub nu) in the sense that those systems with the deepest potential wells have the highest L(sub x) and Mg(sub 2). Thus the depth of the potential well appears to govern both the ability to reatin an ISM at the present epoch and to retain the enriched ejecta of early star formation bursts. Both L(sub x)/L(sub B) and L(sub 6) (the 6 cm radio luminosity) show threshold effects with sigma(sub nu) exhibiting sharp increases at log sigma(sub nu) approximately = 2.2. Finally, there is clearly an interrelationship between the various stellar and structural parameters: The scatter in the bivariate relationships between the shape parameters (a/b and a(sub 4)) and the depth parameter sigma(sub nu) is a function of abundance in the sense that, for a given a(sub 4) or a/b, the systems with the highest sigma(sub nu) also have the highest Mg(sub 2). Furthermore, for a constant sigma(sun nu), disky galaxies tend to have higher Mg(sub 2) than boxy ones. Alternatively, for a given abundance, boxy ellipticals tend to be more massive than disky ellipticals. One possibility is that early-type galaxies of a given mass, originating from mergers (boxy ellipticals), have lower abundances than 'primordial' (disky) early-type galaxies. Another is that disky inner isophotes are due not to primordial dissipation collapse, but to either the self-gravitating inner disks of captured spirals or the dissipational collapse of new disk structures from the premerger ISM. The high measured nuclear Mg(sub 2) values would thus be due to enrichment from secondary bursts of star formation triggered by the merging event.

Eskridge, Paul B.↗

Visual data mining for quantized spatial data

In previous papers we've shown how a well known data compression algorithm called Entropy-constrained Vector Quantization ( can be modified to reduce the size and complexity of very large, satellite data sets. In this paper, we descuss how to visualize and understand the content of such reduced data sets.

cluster analysis↗