Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “multivariate data analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Nonparametric analysis of Minnesota spruce and aspen tree data and LANDSAT data

The application of nonparametric methods in data-intensive problems faced by NASA is described. The theoretical development of efficient multivariate density estimators and the novel use of color graphics workstations are reviewed. The use of nonparametric density estimates for data representation and for Bayesian classification are described and illustrated. Progress in building a data analysis system in a workstation environment is reviewed and preliminary runs presented.

Scott, D. W.↗

PlasmoData.jl — A Julia framework for modeling and analyzing complex data as graphs

Datasets encountered in scientific and engineering applications appear in complex formats (e.g., images, multivariate time series, molecules, video, text strings, networks). Graph theory provides a unifying framework to model such datasets and enables the use of powerful tools that can help analyze, visualize, and extract value from data. In this work, we present PlasmoData.jl, an open-source, Julia framework that uses concepts of graph theory to facilitate the modeling and analysis of complex datasets. The core of our framework is a general data modeling abstraction, which we call a DataGraph. We show how the abstraction and software implementation can be used to represent diverse data objects as graphs and to enable the use of tools from topology, graph theory, and machine learning (e.g., graph neural networks) to conduct a variety of tasks. We illustrate the versatility of the framework by using real datasets: (i) an image classification problem using topological data analysis to extract features from the graph model to train machine learning models; (ii) a disease outbreak problem where we model multivariate time series as graphs to detect abnormal events; and (iii) a technology pathway analysis problem where we highlight how we can use graphs to navigate connectivity. Further, our discussion also highlights how PlasmoData.jl leverages native Julia capabilities to enable compact syntax, scalable computations, and interfaces with diverse packages. Overall, we show that the DataGraph abstraction and PlasmoData.jl Julia package are able to model data within graphs and enable useful analysis.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

A computer analysis of ERTS data of the Lake Gregory area of South Australia with particular emphasis on its role in terrain classification for engineering

A digital computer and multivariate statistical techniques were used to analyze 4-band multispectral data. A representation of the original data for each of the four bands allows a certain degree of terrain interpretation; however, variations in appearance of sites within and between bands, without additional criteria for deciding which representation should be preferred, create difficulties for classification. Investigation of the video data groups produced by principal components analysis and cluster analysis techniques shows that effective correlations with classifications of terrain produced by conventional methods could be carried out. The analyses also highlighted underlying relationships between the various elements. The approach used allows large areas (185 cm by 185 cm) to be classified into fundamental units within a matter of hours and can be applied to those parts of the Earth where facilities for conventional studies are poor or lacking.

Lodwick, G. D.↗

Probabilistic Machine Learning Estimation of Ocean Mixed Layer Depth from Dense Satellite and Sparse In-Situ Observations

The ocean mixed layer plays an important role in the coupling between the upper ocean and atmosphere across a wide range of time scales. Estimation of the variability of the ocean mixed layer is therefore important for atmosphere-ocean prediction and analysis. The increasing coverage of in situ Argo profile data allows for an increasingly accurate analysis of the mixed layer depth (MLD) variability associated with deviations from the seasonal climatology. However, sampling rates are not sufficient to fully resolve subseasonal (<90 day) MLD variability. Yet, many multivariate observations-based analyses include implicit modeled subseasonal MLD variability. One analysis method is optimal interpolation of in situ data, but the interior analysis can be improved by leveraging surface data with regression or variational approaches. Here, we demonstrate how machine learning methods and satellite sea surface temperature, salinity, and height facilitate MLD estimation in a pilot study of two regions: the mid-latitude southern Indian and the eastern equatorial Pacific Oceans. We construct multiple machine learning architectures to produce weekly 1/2° gridded MLD anomaly fields (relative to a monthly climatology) with uncertainty estimates. We test multiple traditional and probabilistic machine learning techniques to compare both accuracy and probabilistic calibration. We validate our methodology by applying it to ocean model simulations. We find that incorporating sea surface data through a machine learning model improves the performance of spatiotemporal MLD variability estimation compared to optimal interpolation of Argo observations alone. These preliminary results are a promising first step for the application of machine learning to MLD prediction.

Machine Learning↗

New Constraints on Titan's Stratospheric n-Butane Abundance

Curiously, n-butane has yet to be detected at Titan, though it is predicted to be present in a wide range of abundances that span over 2.5 orders of magnitude. We have searched infrared spectroscopic observations of Titan for signals from n-butane (n-C4H10) in Titan's stratosphere. Three sets of Cassini Composite Infrared Spectrometer Focal Plane 4 (1050–1500 cm−1) observations were selected for modeling, having been collected from different flybys and pointing latitudes. We modeled the observations with the Nonlinear Optimal Estimator for MultivariatE Spectral AnalySIS radiative transfer tool. Temperature profiles were retrieved for each of the data sets by modeling the ν4 emission from methane near 1305 cm−1. Then, incorporating the temperature profiles, we retrieved abundances of all of Titan's known trace gases that are active in this spectral region, reliably reproducing the observations. We then systematically tested a set of models with varying abundances of n-butane, investigating how the addition of this gas affected the fits. We did this for several different photochemically predicted abundance profiles from the literature, as well as for a constant-with-altitude profile. Ultimately, though we did not produce any firm detection of n-butane, we derived new upper limits on its abundance specific to the use of each profile and to multiple different ranges of stratospheric altitudes. These results will tightly constrain the C4 chemistry of future photochemical modeling of Titan's atmosphere and also motivate the continued search for n-butane and its isomer, isobutane.

Brendan L. Steffens↗

Resolving the Evolution of Atomic Layer-Deposited Thin-Film Growth by Continuous In Situ X-Ray Absorption Spectroscopy

In situ synchrotron X-ray absorption near-edge structure characterization of thin-film titania growth by atomic layer deposition (ALD) over ZnO nanowires reveals persistent low-coordinated Ti motifs leading to a new picture of ALD growth. Through the design of growth and measurement cycles, Ti K-edge spectral data are continuously recorded so as to characterize the film evolution as a function of ALD cycle number and the surface changes within the time scale of the ALD cycle. A unified set of analysis tools is developed to interpret the time-series of spectral data. A prenucleation stage of growth, a transition region, and then a steady-state growth stage are observed with distinguishable features. Multivariate curve resolution analysis, that is physically constrained, demonstrates two specific spectral components with associated, time-dependent concentrations. The bulk-film component tracks the stages of growth. The surface and interface components, present throughout the stages of growth, reveal a significant coverage of relatively isolated or loosely networked tetrahedrally coordinated Ti atomic motifs. Lastly, spectral signatures for the intra-cycle growth kinetics are reconstructed at a time resolution of ~1 s and demonstrate that the transient Ti motifs on the growing surface stabilize within a few seconds of the Ti precursor pulse.

36 MATERIALS SCIENCE↗

New methods for trace analysis of gamma-irradiated pentaerythritol tetranitrate

High explosives (HEs) are used in a diverse range of applications in which they could be exposed to various radiation levels that may cause potential chemical changes. This study further evaluated pentaerythritol tetranitrate (PETN) that was previously aged with a low-level 2 kGy dose of gamma irradiation in order to understand chemical changes caused by irradiation. Both unirradiated PETN and gamma-irradiated PETN were analyzed using ultra high-pressure liquid chromatography coupled to quadrupole time of flight mass spectrometry (UHPLC-QTOF). The resulting data were processed in a non-targeted manner using Fisher's ratio analysis and multivariate curve resolution-alternating least squares (MCR-ALS) to aid in discovery and identification of the chemical changes brought about by irradiation without a priori knowledge. The application of using UHPLC-QTOF in combination with chemometric techniques for the analysis of irradiated samples has not previously been performed. Here, in this work, we show how to use this method to provide chemical information that would otherwise not be discernible, such as the discovery of degradation of the various homologues of PETN. Major differences identified with radiolytic aging of the PETN sample included decomposition products that resulted from the degradation of the trigger linkage – the O–NO 2 bonds – resulting in the formation of alcohol and aldehyde groups. Similar degradation was also observed in the PETN homologues as well as interconversion from one homologue to another.

38 RADIATION CHEMISTRY, RADIOCHEMISTRY, AND NUCLEA↗

Reflectance of vegetation, soil, and water

There are no author-identified significant results in this report. This report deals with the selection of the best channels from the 24-channel aircraft data to represent crop and soil conditions. A three-step procedure has been developed that involves using univariate statistics and an F-ratio test to indicate the best 14 channels. From the 14, the 10 best channels are selected by a multivariate stochastic process. The third step involves the pattern recognition procedures developed in the data analysis plan. Indications are that the procedures in use are satsifactory and will extract the desired information from the data.

Wiegand, C. L.↗

The second-moment climatology of the GATE rain rate data

The first part of this paper presents the description of the GARP (Global Atmospheric Research Program) Atlantic Tropical Experiment (GATE) 1 rain-rate data and its two-dimensional spectral and correlation characteristics, which has made it possible to accomplish the following: to show the concentration of a significant power along the frequency axis in the spatiotemporal spectra; to detect a diurnal cycle (which has a range of variation of about 3.4-5.4 mm/n) as one of the sources of bias in the rain statistics of satellite data; to study the distinction between the north-south and east-west transport of spatial rain-rate field and character of its anisotropy; to evaluate the scales of the distinction between second-moment estimates associated with ground and satellite samples; and to determine the appropriate spatial and temporal scales of simple linear stochastic models fitted to averaged rain-rate fields. The second part of this paper is devoted to an analysis of the diffusion of the rain rate by establishing a relationship between the parameters of the multivariate autoregressive model and the coefficients of a diffusion equation. This analysis led to the use of rain data to estimate the rain advection velocity as well as other coefficients of the diffusion equation of the corresponding field. The results obtained can be used for comparison with corresponding estimates of other sources of data (satellite, Tropical Oceans Global Atmosphere Coupled Ocean - Atmosphere Response Experiment (TOGA, COARE) or simulated by physical models), for generating multiple samples of any size, for solving the inverse problems of some of the hydrodynamic equations, and in some other areas of rain data analysis and modeling.

Polyak, Ilya↗

Statistical Analysis of Factors Riving Surface Ozone Variability over Continental South Africa

Statistical relationships between surface ozone (O3) concentration, precursor species and meteorological conditions in continental South Africa were examined from data obtained from measurement stations in north-eastern South Africa. Three multivariate statistical methods were applied in the investigation, i.e. multiple linear regression (MLR), principal component analysis (PCA) and –regression (PCR), and generalised additive model (GAM) analysis. The daily maximum 8-h moving average O3 concentrations were considered in these statistical models (dependent variable). MLR models indicated that meteorology and precursor species concentrations are able to explain ~50% of the variability in daily maximum O3 levels. MLR analysis revealed that atmospheric carbon monoxide (CO), temperature and relative humidity were the strongest factors affecting the daily O3 variability. In summer, daily O3 variances were mostly associated with relative humidity, while winter O3 levels were mostly linked to temperature and CO. PCA indicated that CO, temperature and relative humidity were not strongly collinear. GAM also identified CO, temperature and relative humidity as the strongest factors affecting the daily variation of O3. Partial residual plots found that temperature, radiation and nitrogen oxides most likely have a non-linear relationship with O3,while the relationship with relative humidity and CO is probably linear. An inter-comparison between O3 levels modelled with the three statistical models compared to measured O3 concentrations showed that the GAM model offered a slight improvement over the MLR model. These findings emphasise the critical role of regional-scale O3 precursors coupled with meteorological conditions in daily variances of O3 levels in continental South Africa.

multiple linear regression (MLR)↗

Robust framework and software implementation for fast speciation mapping

One of the greatest benefits of synchrotron radiation is the ability to perform chemical speciation analysis through X-ray absorption spectroscopies (XAS). XAS imaging of large sample areas can be performed with either full-field or raster-scanning modalities. A common practice to reduce acquisition time while decreasing dose and/or increasing spatial resolution is to compare X-ray fluorescence images collected at a few diagnostic energies. In this work, several authors have used different multivariate data processing strategies to establish speciation maps. Furthermore, the theoretical aspects and assumptions that are often made in the analysis of these datasets are focused on. A robust framework is developed to perform speciation mapping in large bulk samples at high spatial resolution by comparison with known references. Two fully operational software implementations are provided: a user-friendly implementation within the MicroAnalysis Toolkit software, and a dedicated script developed under the R environment. The procedure is exemplified through the study of a cross section of a typical fossil specimen. Additionally, the algorithm provides accurate speciation and concentration mapping while decreasing the data collection time by typically two or three orders of magnitude compared with the collection of whole spectra at each pixel. Whereas acquisition of spectral datacubes on large areas leads to very high irradiation times and doses, which can considerably lengthen experiments and generate significant alteration of radiation-sensitive materials, this sparse excitation energy procedure brings the total irradiation dose greatly below radiation damage thresholds identified in previous studies. This approach is particularly adapted to the chemical study of heterogeneous radiation-sensitive samples encountered in environmental, material, and life sciences.

47 OTHER INSTRUMENTATION↗

Chemistry imaging and distribution analysis of rare earth elements in coal using LIBS and LA-ICP-MS instruments

Currently, demand for rare earth elements (REEs) increased significantly. Coal is actively evaluated as potential economic sources for extraction of REEs. Here, in this work, laser-induced breakdown spectroscopy (LIBS) was evaluated for rapid estimation of REEs content and their distribution in the natural coal samples. The results were compared with similar laser ablation–inductively coupled plasma–mass spectrometry (LA-ICP-MS) measurements. Thirteen coal samples (nine standard samples and five natural samples) were used in this study. Powder samples were pressed into pellets while coal chunks were directly ablated for data recording. Pellets of the powder standard samples were used to optimize the data acquisition system and then data recorded with this optimized system was used to identify the proper data acquisition and analysis models. After establishing the proper data acquisition system and analysis model using the standard samples, natural coal samples in powder form and their chunks were utilized to record LIBS and LA-ICP-MS spectra. Multivariate calibration models were developed using four of the natural samples, which were evaluated by predicting the REE content in the fifth sample. Principal component analysis was performed on the LIBS data obtained from the natural samples and it classified all the samples with high accuracy. Two-dimensional (2D) elemental mapping on coal chunk samples was also performed using both LIBS and LA-ICP-MS to study the distribution of REEs in the samples. The resulting elemental images and their correlations can be used to infer mineral distributions.

01 COAL, LIGNITE, AND PEAT↗

Stochastic Simulation of Daily Suspended Sediment Concentration Using Multivariate Copulas

Estimation of daily suspended sediment concentration (SSC) is required for water resources and environment management. In this paper, a copula-based stochastic method was proposed for daily SSC simulation. Here, the multivariate copula function, constructed based on a bivariate copula and two bivariate conditional probability distributions, was used to model the temporal and cross dependence structures in daily SSCs. Then, the daily SSCs were generated by sampling from the multivariate conditional distribution. As a result, synthetic long-term SSCs data beyond the limited observation period can be provided for water resources managers, which plays a critical role in accurately estimating frequency and magnitude of extreme SSCs events. The proposed method was under rigorous examination by applying to a case study at Pingshan station in the Jinsha River Basin, China. Results showed that the generated daily SSC sequences not only had a high degree of accuracy in preserving the statistical characteristics of the daily SSC observations, but also captured both the temporal correlation and the cross-correlation between the daily streamflow and daily SSC. Specifically, the average daily relative error values corresponding to mean, standard deviation, skewness, lag-1 temporal correlation, and cross correlation were 0.87%, 4.24%, 7.52%, 0.51% and 2.02%, respectively. The multivariate copula framework proposed here can accurately and efficiently generate long-term daily SSC data for water resources management such as frequency analysis and risk assessment of extreme SSC events.

54 ENVIRONMENTAL SCIENCES↗

An Integrated Analysis of the Physiological Effects of Space Flight: Executive Summary

A large array of models were applied in a unified manner to solve problems in space flight physiology. Mathematical simulation was used as an alternative way of looking at physiological systems and maximizing the yield from previous space flight experiments. A medical data analysis system was created which consist of an automated data base, a computerized biostatistical and data analysis system, and a set of simulation models of physiological systems. Five basic models were employed: (1) a pulsatile cardiovascular model; (2) a respiratory model; (3) a thermoregulatory model; (4) a circulatory, fluid, and electrolyte balance model; and (5) an erythropoiesis regulatory model. Algorithms were provided to perform routine statistical tests, multivariate analysis, nonlinear regression analysis, and autocorrelation analysis. Special purpose programs were prepared for rank correlation, factor analysis, and the integration of the metabolic balance data.

Leonard, J. I.↗

Multivariate space - time analysis of PRE-STORM precipitation

This paper presents the methodologies and results of the multivariate modeling and two-dimensional spectral and correlation analysis of PRE-STORM rainfall gauge data. Estimated parameters of the models for the specific spatial averages clearly indicate the eastward and southeastward wave propagation of rainfall fluctuations. A relationship between the coefficients of the diffusion equation and the parameters of the stochastic model of rainfall fluctuations is derived that leads directly to the exclusive use of rainfall data to estimate advection speed (about 12 m/s) as well as other coefficients of the diffusion equation of the corresponding fields. The statistical methodology developed here can be used for confirmation of physical models by comparison of the corresponding second-moment statistics of the observed and simulated data, for generating multiple samples of any size, for solving the inverse problem of the hydrodynamic equations, and for application in some other areas of meteorological and climatological data analysis and modeling.

Polyak, Ilya↗

M-DAS: System for multispectral data analysis

M-DAS is a ground data processing system designed for analysis of multispectral data. M-DAS operates on multispectral data from LANDSAT, S-192, M2S and other sources in CCT form. Interactive training by operator-investigators using a variable cursor on a color display was used to derive optimum processing coefficients and data on cluster separability. An advanced multivariate normal-maximum likelihood processing algorithm was used to produce output in various formats: color-coded film images, geometrically corrected map overlays, moving displays of scene sections, coverage tabulations and categorized CCTs. The analysis procedure for M-DAS involves three phases: (1) screening and training, (2) analysis of training data to compute performance predictions and processing coefficients, and (3) processing of multichannel input data into categorized results. Typical M-DAS applications involve iteration between each of these phases. A series of photographs of the M-DAS display are used to illustrate M-DAS operation.

Johnson, R. H.↗

Noise and drift analysis of non-equally spaced timing data

Generally, it is possible to obtain equally spaced timing data from oscillators. The measurement of the drifts and noises affecting oscillators is then performed by using a variance (Allan variance, modified Allan variance, or time variance) or a system of several variances (multivariance method). However, in some cases, several samples, or even several sets of samples, are missing. In the case of millisecond pulsar timing data, for instance, observations are quite irregularly spaced in time. Nevertheless, since some observations are very close together (one minute) and since the timing data sequence is very long (more than ten years), information on both short-term and long-term stability is available. Unfortunately, a direct variance analysis is not possible without interpolating missing data. Different interpolation algorithms (linear interpolation, cubic spline) are used to calculate variances in order to verify that they neither lose information nor add erroneous information. A comparison of the results of the different algorithms is given. Finally, the multivariance method was adapted to the measurement sequence of the millisecond pulsar timing data: the responses of each variance of the system are calculated for each type of noise and drift, with the same missing samples as in the pulsar timing sequence. An estimation of precision, dynamics, and separability of this method is given.

Vernotte, F.↗