Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “multivariate data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Foundations of machine learning for low-temperature plasmas: methods and case studies

Abstract Machine learning (ML) and artificial intelligence have proven to be an invaluable tool in tackling a vast array of scientific, engineering, and societal problems. The main drivers behind the recent proliferation of ML in practically all aspects of science and technology can be attributed to: (a) improved data acquisition and inexpensive data storage; (b) exponential growth in computing power; and (c) availability of open-source software and resources that have made the use of state-of-the-art ML algorithms widely accessible. The impact of ML on the field of low-temperature plasmas (LTPs) could be particularly significant in the emerging applications that involve plasma treatment of complex interfaces in areas ranging from the manufacture of microelectronics and processing of quantum materials, to the LTP-driven electrification of the chemical industry, and to medicine and biotechnology. This is primarily due to the complex and poorly-understood nature of the plasma-surface interactions in these applications that pose unique challenges to the modeling, diagnostics, and predictive control of LTPs. As the use of ML is becoming more prevalent, it is increasingly paramount for the LTP community to be able to critically analyze and assess the concepts and techniques behind data-driven approaches. To this end, the goal of this paper is to provide a tutorial overview of some of the widely-used ML methods that can be useful, amongst others, for discovering and correlating patterns in the data that may be otherwise impractical to decipher by human intuition alone, for learning multivariable nonlinear data-driven prediction models that are capable of describing the complex behavior of plasma interacting with interfaces, and for guiding the design of experiments to explore the parameter space of plasma-assisted processes in a systematic and resource-efficient manner. We illustrate the utility of various supervised, unsupervised and active learning methods using LTP datasets consisting of commonly-available, information-rich measurements (e.g. optical emission spectra, current–voltage characteristics, scanning electron microscope images, infrared surface temperature measurements, Fourier transform infrared spectra). All the ML demonstrations presented in this paper are carried out using open-source software; the datasets and codes are made publicly available. The FAIR guiding principles for scientific data management and stewardship can accelerate the adoption and development of ML in the LTP community.

Physics↗

TPSAS-NF1676L-36113-DND

Concerns about the effects of extreme heat and poor air quality are increasing in North America’s largest urban centers. In Philadelphia, environmental and public health groups are concerned about how these phenomena disproportionality affect marginalized communities and populations, which often have extensive impervious surfaces and little access to green space. In order to address these concerns, the Philadelphia Department of Public Health and the Office of Sustainability seek to effectively prioritize cooling initiatives to reduce urban heat and decrease air pollutants. We evaluated land surface temperature (LST) and the Normalized Difference Vegetation Index (NDVI), as a measure of overall greenness, obtained from NASA Earth observations Aqua and Terra Moderate Resolution Imaging Spectroradiometer (MODIS), and the Ecosystem Spaceborne Thermal Radiometer Experiment on Space Station (ECOSTRESS). These analyses we recombined with local tree inventory, air quality, and socioeconomic data through a multivariate analysis to identify areas where new trees or cooling adaptations are most needed. The results and data of this project can be used by our partners to inform both short-term heat relief planning and a long-term, multi-agency heat response.

Spencer Nelson↗

Long-term missing value imputation for time series data using deep neural networks

We present an approach that uses a deep learning model, in particular, a MultiLayer Perceptron, for estimating the missing values of a variable in multivariate time series data. We focus on filling a long continuous gap (e.g., multiple months of missing daily observations) rather than on individual randomly missing observations. Our proposed gap filling algorithm uses an automated method for determining the optimal MLP model architecture, thus allowing for optimal prediction performance for the given time series. We tested our approach by filling gaps of various lengths (three months to three years) in three environmental datasets with different time series characteristics, namely daily groundwater levels, daily soil moisture, and hourly Net Ecosystem Exchange. We compared the accuracy of the gap-filled values obtained with our approach to the widely used R-based time series gap filling methods ImputeTS and mtsdi. The results indicate that using an MLP for filling a large gap leads to better results, especially when the data behave nonlinearly. Thus, our approach enables the use of datasets that have a large gap in one variable, which is common in many long-term environmental monitoring observations.

97 MATHEMATICS AND COMPUTING↗

A data based random number generator for a multivariate distribution (using stochastic interpolation)

Let X be a K-dimensional random variable serving as input for a system with output Y (not necessarily of dimension k). given X, an outcome Y or a distribution of outcomes G(Y/X) may be obtained either explicitly or implicity. The situation is considered in which there is a real world data set X sub j sub = 1 (n) and a means of simulating an outcome Y. A method for empirical random number generation based on the sample of observations of the random variable X without estimating the underlying density is discussed.

Thompson, J. R.↗

Astronaut Preflight Cardiovascular Variables Associated with Vascular Compliance are Highly Correlated with Post-Flight Eye Outcome Measures in the Visual Impairment Intracranial Pressure (VIIP) Syndrome Following Long Duration Spaceflight

The detection of the first VIIP case occurred in 2005, and adequate eye outcome measures were available for 31 (67.4%) of the 46 long duration US crewmembers who had flown on the ISS since its first crewed mission in 2000. Therefore, this analysis is limited to a subgroup (22 males and 9 females). A "cardiovascular profile" for each astronaut was compiled by examining twelve individual parameters; eleven of these were preflight variables: systolic blood pressure, pulse pressure, body mass index, percentage body fat, LDL, HDL, triglycerides, use of anti‐lipid medication, fasting serum glucose, and maximal oxygen uptake in ml/kg. Each of these variables was averaged across three preflight annual physical exams. Astronaut age prior to the long duration mission, and inflight salt intake was also included in the analysis. The group of cardiovascular variables for each crew member was compared with seven VIIP eye outcome variables collected during the immediate post‐flight period: anterior-posterior axial length of the globe measured by ultrasound and optical biometry; optic nerve sheath diameter, optic nerve diameter, and optic nerve to sheath ratio‐ each measured by ultrasound and magnetic resonance imaging (MRI), intraocular pressure (IOP), change in manifest refraction, mean retinal nerve fiber layer (RNFL) on optical coherence tomography (OCT), and RNFL of the inferior and superior retinal quadrants. Since most of the VIIP eye outcome measures were added sequentially beginning in 2005, as knowledge of the syndrome improved, data were unavailable for 22.0% of the outcome measurements. To address the missing data, we employed multivariate multiple imputation techniques with predictive mean matching methods to accumulate 200 separate imputed datasets for analysis. We were able to impute data for the 22.0% of missing VIIP eye outcomes. We then applied Rubin's rules for collapsing the statistical results across our 200 multiply imputed data sets to assess the canonical correlation between the eye outcomes and the twelve astronaut cardiovascular variables available for all 31 subjects. Results: A highly significant canonical correlation was observed among the canonical solutions (p<.00001), with an average best canonical correlation of.97. The results suggest a strong association between astronauts' measures of cardiovascular health and the seven eye outcomes of the VIIP syndrome used in this analysis. Furthermore, the "joint test" revealed a significant difference in cardiovascular profile between male and female astronauts (Prob > F = 0.00001). Overall, female astronauts demonstrated a significantly healthier cardiovascular status. Individually, the female astronauts had significantly healthier profiles on seven of twelve cardiovascular variables than the men (p values ranging from <0.0001 to <0.05). Male astronauts did not demonstrate significantly healthier values on any of the twelve cardiovascular variables measured

Otto, Christian↗

Patterns of Element Incorporation in Calcium Carbonate Biominerals Recapitulate Phylogeny for a Diverse Range of Marine Calcifiers

Elemental ratios in biogenic marine calcium carbonates are widely used in geobiology, environmental science, and paleoenvironmental reconstructions. It is generally accepted that the elemental abundance of biogenic marine carbonates reflects a combination of the abundance of that ion in seawater, the physical properties of seawater, the mineralogy of the biomineral, and the pathways and mechanisms of biomineralization. Here we report measurements of a suite of nine elemental ratios (Li/Ca, B/Ca, Na/Ca, Mg/Ca, Zn/Ca, Sr/Ca, Cd/Ca, Ba/Ca, and U/Ca) in 18 species of benthic marine invertebrates spanning a range of biogenic carbonate polymorph mineralogies (low-Mg calcite, high-Mg calcite, aragonite, mixed mineralogy) and of phyla (including Mollusca, Echinodermata, Arthropoda, Annelida, Cnidaria, Chlorophyta, and Rhodophyta) cultured at a single temperature (25°C) and a range of p CO 2 treatments (ca. 409, 606, 903, and 2856 ppm). This dataset was used to explore various controls over elemental partitioning in biogenic marine carbonates, including species-level and biomineralization-pathway-level controls, the influence of internal pH regulation compared to external pH changes, and biocalcification responses to changes in seawater carbonate chemistry. The dataset also enables exploration of broad scale phylogenetic patterns of elemental partitioning across calcifying species, exhibiting high phylogenetic signals estimated from both uni- and multivariate analyses of the elemental ratio data (univariate: λ = 0–0.889; multivariate: λ = 0.895–0.99). Comparing partial R 2 values returned from non-phylogenetic and phylogenetic regression analyses echo the importance of and show that phylogeny explains the elemental ratio data 1.4–59 times better than mineralogy in five out of nine of the elements analyzed. Therefore, the strong associations between biomineral elemental chemistry and species relatedness suggests mechanistic controls over element incorporation rooted in the evolution of biomineralization mechanisms.

58 GEOSCIENCES↗

Search for ${\text {Z}{}{}} {\text {Z}{}{}} $ and ${\text {Z}{}{}} {\text {H}{}{}} $ production in the ${\text {b}{}{}} {\bar{{\text {b}{}{}}}{}{}} {\text {b}{}{}} {\bar{{\text {b}{}{}}}{}{}} $ final state using proton-proton collisions at $\sqrt{s}=13\,\text {Te}\hspace{-.08em}\text {V} $

A search for ${\text {Z}{}{}} {\text {Z}{}{}} $ and ${\text {Z}{}{}} {\text {H}{}{}} $ production in the ${\text {b}{}{}} {\bar{{\text {b}{}{}}}{}{}} {\text {b}{}{}} {\bar{{\text {b}{}{}}}{}{}} $ final state is presented, where H is the standard model (SM) Higgs boson. The search uses an event sample of proton-proton collisions corresponding to an integrated luminosity of 133$\,\text {fb}^{-1}$ collected at a center-of-mass energy of 13$\,\text {Te}\hspace{-.08em}\text {V}$ with the CMS detector at the CERN LHC. The analysis introduces several novel techniques for deriving and validating a multi-dimensional background model based on control samples in data. A multiclass multivariate classifier customized for the ${\text {b}{}{}} {\bar{{\text {b}{}{}}}{}{}} {\text {b}{}{}} {\bar{{\text {b}{}{}}}{}{}} $ final state is developed to derive the background model and extract the signal. The data are found to be consistent, within uncertainties, with the SM predictions. The observed (expected) upper limits at 95% confidence level are found to be 3.8 (3.8) and 5.0 (2.9) times the SM prediction for the ${\text {Z}{}{}} {\text {Z}{}{}} $ and ${\text {Z}{}{}} {\text {H}{}{}} $ production cross sections, respectively.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Multivariate Error Covariance Estimates by Monte-Carlo Simulation for Assimilation Studies in the Pacific Ocean

One of the most difficult aspects of ocean state estimation is the prescription of the model forecast error covariances. The paucity of ocean observations limits our ability to estimate the covariance structures from model-observation differences. In most practical applications, simple covariances are usually prescribed. Rarely are cross-covariances between different model variables used. Here a comparison is made between a univariate Optimal Interpolation (UOI) scheme and a multivariate OI algorithm (MvOI) in the assimilation of ocean temperature. In the UOI case only temperature is updated using a Gaussian covariance function and in the MvOI salinity, zonal and meridional velocities as well as temperature, are updated using an empirically estimated multivariate covariance matrix. Earlier studies have shown that a univariate OI has a detrimental effect on the salinity and velocity fields of the model. Apparently, in a sequential framework it is important to analyze temperature and salinity together. For the MvOI an estimation of the model error statistics is made by Monte-Carlo techniques from an ensemble of model integrations. An important advantage of using an ensemble of ocean states is that it provides a natural way to estimate cross-covariances between the fields of different physical variables constituting the model state vector, at the same time incorporating the model's dynamical and thermodynamical constraints as well as the effects of physical boundaries. Only temperature observations from the Tropical Atmosphere-Ocean array have been assimilated in this study. In order to investigate the efficacy of the multivariate scheme two data assimilation experiments are validated with a large independent set of recently published subsurface observations of salinity, zonal velocity and temperature. For reference, a third control run with no data assimilation is used to check how the data assimilation affects systematic model errors. While the performance of the UOI and MvOI is similar with respect to the temperature field, the salinity and velocity fields are greatly improved when multivariate correction is used, as evident from the analyses of the rms differences of these fields and independent observations. The MvOI assimilation is found to improve upon the control run in generating the water masses with properties close to the observed, while the UOI failed to maintain the temperature and salinity structure.

Borovikov, Anna↗

Anomaly Detection in Flight Operational Data Using Deep Learning

In this session, we demonstrate two recently developed deep learning models for anomaly detection in flight operational data by the Data Sciences Group at NASA Ames Research Center. The first model is Convolutional Variational Auto-Encoder (CVAE) [1], which is an unsupervised deep encoder-decoder model, designed specifically for finding anomalies in heterogeneous multivariate time series data. We will demonstrate its application to finding anomalies in streaming data from NASA’s Digital Information Platform’s Fuser source. CVAE identifies data instances that are not representative of expected nominal behavior as anomalous. Since it is an unsupervised approach, the flagged anomalies will need to be reviewed by the subject matter experts (SMEs) for validation and labeling and is designed to assist with vulnerability discovery within Safety Monitoring System programs. The second model is Robust and Explainable Semi-supervised Anomaly Detection (RESAD) model [2], which builds on CVAE to allow learning from both minimally labeled data (previously reviewed by the SMEs) as well as majority unlabeled data. RESAD takes advantage of graph theoretic techniques to propagate the labels from the labeled data to the unlabeled data based on a pre-defined similarity metric and structures the learned feature space from flight time-series so that data of the same class would cluster tightly together. This model characteristic is enabled by training with an augmented loss function and allows learning of a more informative feature space for down-stream tasks such as search and active learning. We demonstrate RESAD using data from the NASA DASHlink project [3].

anomaly detection↗

Data augmentation for disruption prediction via robust surrogate models

The goal of this work is to generate large statistically representative data sets to train machine learning models for disruption prediction provided by data from few existing discharges. Such a comprehensive training database is important to achieve satisfying and reliable prediction results in artificial neural network classifiers. Here, we aim for a robust augmentation of the training database for multivariate time series data using Student t process regression. We apply Student t process regression in a state space formulation via Bayesian filtering to tackle challenges imposed by outliers and noise in the training data set and to reduce the computational complexity. Thus, the method can also be used if the time resolution is high. We use an uncorrelated model for each dimension and impose correlations afterwards via colouring transformations. We demonstrate the efficacy of our approach on plasma diagnostics data of three different disruption classes from the DIII-D tokamak. To evaluate if the distribution of the generated data is similar to the training data, we additionally perform statistical analyses using methods from time series analysis, descriptive statistics and classic machine learning clustering algorithms.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Data augmentation for disruption prediction via robust surrogate models

The goal of this work is to generate large statistically representative datasets to train machine learning models for disruption prediction provided by data from few existing discharges. Such a comprehensive training database is important to achieve satisfying and reliable prediction results in artificial neural network classifiers. Here, we aim for a robust augmentation of the training database for multivariate time series data using Student-t process regression. We apply Student-t process regression in a state space formulation via Bayesian filtering to tackle challenges imposed by outliers and noise in the training data set and to reduce the computational complexity. Thus, the method can also be used if the time resolution is high. We use an uncorrelated model for each dimension and impose correlations afterwards via coloring transformations. We demonstrate the efficacy of our approach on plasma diagnostics data of three different disruption classes from the DIII-D tokamak. To evaluate if the distribution of the generated data is similar to the training data, we additionally perform statistical analyses using methods from time series analysis, descriptive statistics, and classic machine learning clustering algorithms.

97 MATHEMATICS AND COMPUTING↗

Evaluation of a multivariate variational assimilation of conventional and satellite data for the diagnosis of cyclone systems

A variational data assimilation method for the study of cyclone-scale weather systems is described. The variational data assimilation method is to incorporate primitive equations for a moist, convectively unstable atmosphere and the radiative transfer equation. The variables to be adjusted include the three-dimensional vector wind, height, temperature, and moisture from rawinsonde data, and cloud-wind vectors, moisture, and radiance from satellite data. The development of variational model 1 which contains two nonlinear horizontal momentum equations, an integrated continuity equation, and a hydrostatic equation is examined. Examples applying the assimilation model to rawinsonde and satellite data are presented.

Achtemeier, Gary L.↗

Modeling the High Speed Research Cycle 2B Longitudinal Aerodynamic Database Using Multivariate Orthogonal Functions

The data for longitudinal non-dimensional, aerodynamic coefficients in the High Speed Research Cycle 2B aerodynamic database were modeled using polynomial expressions identified with an orthogonal function modeling technique. The discrepancy between the tabular aerodynamic data and the polynomial models was tested and shown to be less than 15 percent for drag, lift, and pitching moment coefficients over the entire flight envelope. Most of this discrepancy was traced to smoothing local measurement noise and to the omission of mass case 5 data in the modeling process. A simulation check case showed that the polynomial models provided a compact and accurate representation of the nonlinear aerodynamic dependencies contained in the HSR Cycle 2B tabular aerodynamic database.

Morelli, E. A.↗

Transport in the Subtropical Lowermost Stratosphere during CRYSTAL-FACE

We use in situ measurements of water vapor (H2O), ozone (O3), carbon dioxide (CO2), carbon monoxide (CO), nitric oxide (NO), and total reactive nitrogen (NO(y)) obtained during the CRYSTAL-FACE campaign in July 2002 to study summertime transport in the subtropical lowermost stratosphere. We use an objective methodology to distinguish the latitudinal origin of the sampled air masses despite the influence of convection, and we calculate backward trajectories to elucidate their recent geographical history. The methodology consists of exploring the statistical behavior of the data by performing multivariate clustering and agglomerative hierarchical clustering calculations, and projecting cluster groups onto principal component space to identify air masses of like composition and hence presumed origin. The statistically derived cluster groups are then examined in physical space using tracer-tracer correlation plots. Interpretation of the principal component analysis suggests that the variability in the data is accounted for primarily by the mean age of air in the stratosphere, followed by the age of the convective influence, and lastly by the extent of convective influence, potentially related to the latitude of convective injection [Dessler and Sherwuud, 2004]. We find that high-latitude stratospheric air is the dominant source region during the beginning of the campaign while tropical air is the dominant source region during the rest of the campaign. Influence of convection from both local and non-local events is frequently observed. The identification of air mass origin is confirmed with backward trajectories, and the behavior of the trajectories is associated with the North American monsoon circulation.

Pittman, Jasna V.↗

Transport in the Subtropical Lowermost Stratosphere during the Cirrus Regional Study of Tropical Anvils and Cirrus Layers-Florida Area Cirrus Experiment

We use in situ measurements of water vapor (H2O), ozone (O3), carbon dioxide (CO2), carbon monoxide (CO), nitric oxide (NO), and total reactive nitrogen (NOy) obtained during the CRYSTAL-FACE campaign in July 2002 to study summertime transport in the subtropical lowermost stratosphere. We use an objective methodology to distinguish the latitudinal origin of the sampled air masses despite the influence of convection, and we calculate backward trajectories to elucidate their recent geographical history. The methodology consists of exploring the statistical behavior of the data by performing multivariate clustering and agglomerative hierarchical clustering calculations and projecting cluster groups onto principal component space to identify air masses of like composition and hence presumed origin. The statistically derived cluster groups are then examined in physical space using tracer-tracer correlation plots. Interpretation of the principal component analysis suggests that the variability in the data is accounted for primarily by the mean age of air in the stratosphere, followed by the age of the convective influence, and last by the extent of convective influence, potentially related to the latitude of convective injection (Dessler and Sherwood, 2004). We find that high-latitude stratospheric air is the dominant source region during the beginning of the campaign while tropical air is the dominant source region during the rest of the campaign. Influence of convection from both local and nonlocal events is frequently observed. The identification of air mass origin is confirmed with backward trajectories, and the behavior of the trajectories is associated with the North American monsoon circulation.

Pittman, Jasna V.↗

Supplemental Data for the Manuscript: Quantification of Manganese for ChemCam Mars and Laboratory Spectra Using a Multivariate Model

This dataset includes all of the data needed to validate and/or reproduce the manganese calibration model described in the manuscript. The reference database contains metadata for the new Mn-bearing standards, minerals, and mixtures that are > 2.9 wt.% MnO. In addition, the files include the MnO composition data for all standards used, non-normalized spectral data, mean peak area spectrum, results of outlier determination, RMSECV data, regression vectors, and Test Set predictions.

58 GEOSCIENCES↗

Long–short-term memory encoder–decoder with regularized hidden dynamics for fault detection in industrial processes

The ability of recurrent neural networks (RNN) to model nonlinear dynamics of high dimensional process data has enabled data-driven RNN-based fault detection algorithms. Previous studies have focused on detecting faults by identifying the discrepancies in data distribution between the faulty and normal data, as reflected in prediction errors generated by RNN models. However, in industrial processes, variations in data distribution can also result from changes in normal control setpoints and compensatory control adjustments in response to disturbances, making it hard to differentiate between normal and faulty conditions. This paper proposes a fault detection method utilizing a long short-term memory (LSTM) encoder–decoder structure with regularized hidden dynamics and reversible instance normalization (RevIN) to compactly represent high-dimensional measurements for effective monitoring. During training, the hidden states of the model are regularized to form a low-dimensional latent space representation of the original multivariate time series data. As a result, the prediction errors of the latent states can be used to monitor the abnormal dynamic variations, while the reconstruction errors of the measured variables are used to monitor the abnormal static variations. Furthermore, the proposed indices can reflect operating conditions, even when the distribution of test data changes, which helps distinguish faults from normal adjustments and disturbances that controllers can settle. Here, data from numerical simulation and the Tennessee Eastman process are used to illustrate the effectiveness of the proposed fault detection method.

42 ENGINEERING↗

Estimation of terrain iso-gradients from a stochastic range data measurement matrix

The problem of estimating terrain iso-gradients for an autonomous roving vehicle is complicated by the measurement error inherent in the range data matrix. In this paper, the in-path and cross-path slopes at each data point in the multivariable range matrix are expressed in terms of parameters at that data point and those at its surrounding points. The sensitivity in the change of these slopes - due to the errors in measurements in range, elevation angle and azimuth angle are formulated. These sensitivities are used to determine the appropriate statistics of the gradient, which are then used to describe the terrain with a known probability of accuracy.

Shen, C. N.↗