Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Mahalanobis distance”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

An AERONET-Based Aerosol Classification Using the Mahalanobis Distance

We present an aerosol classification based on AERONET aerosol data from 1993 to 2012. We used the AERONET Level 2.0 almucantar aerosol retrieval products to define several reference aerosol clusters which are characteristic of the following general aerosol types: Urban-Industrial, Biomass Burning, Mixed Aerosol, Dust, and Maritime. The classification of a particular aerosol observation as one of these aerosol types is determined by its five-dimensional Mahalanobis distance to each reference cluster. We have calculated the fractional aerosol type distribution at 190 AERONET sites, as well as the monthly variation in aerosol type at those locations. The results are presented on a global map and individually in the supplementary material. Our aerosol typing is based on recognizing that different geographic regions exhibit characteristic aerosol types. To generate reference clusters we only keep data points that lie within a Mahalanobis distance of 2 from the centroid. Our aerosol characterization is based on the AERONET retrieved quantities, therefore it does not include low optical depth values. The analysis is based on point sources (the AERONET sites) rather than globally distributed values. The classifications obtained will be useful in interpreting aerosol retrievals from satellite borne instruments.

Hamill, Patrick↗

On the Use of Mahalanobis Distance in Particle Image Velocimetry Post-Processing

Particle Image Velocimetry (PIV) is a method of flow measurement that has become increasingly popular as an experimental tool. New technology has made high-speed and higher-dimension (stereoscopic, tomographic, etc) methods available to an ever-growing population of researchers. These advanced methods can provide significantly more data than traditional low-speed planar PIV, but these larger data sets also require more resources to process and store. A major time sink in the post-processing of large PIV data sets is the identification and rejection of “bad” or ”spurious” vectors that survive an initial processing step in commercial software, which attempts to identify spurious vectors on an image-by-image basis and not with respect to repeated trials. The Mahalanobis distance, an almost 100-year-old statistical function, was determined to be well suited for this task for its computational efficiency and higher-dimensional nature. This work includes a summary of the Mahalanobis distance and validation of its usefulness as a tool for outlier rejection in the post-processing of PIV data.

Experimental Methods↗

Structure-informed clustering for population stratification in association studies

Background: Identifying variants associated with complex traits is a challenging task in genetic association studies due to linkage disequilibrium (LD) between genetic variants and population stratification, unrelated to the disease risk. Existing methods of population structure correction use principal component analysis or linear mixed models with a random effect when modeling associations between a trait of interest and genetic markers. However, due to stringent significance thresholds and latent interactions between the markers, these methods often fail to detect genuinely associated variants. Results: To overcome this, we propose CluStrat, which corrects for complex arbitrarily structured populations while leveraging the linkage disequilibrium induced distances between genetic markers. It performs an agglomerative hierarchical clustering using the Mahalanobis distance covariance matrix of the markers. In simulation studies, we show that our method outperforms existing methods in detecting true causal variants. Applying CluStrat on WTCCC2 and UK Biobank cohorts, we found biologically relevant associations in Schizophrenia and Myocardial Infarction. CluStrat was also able to correct for population structure in polygenic adaptation of height in Europeans. Conclusions: CluStrat highlights the advantages of biologically relevant distance metrics, such as the Mahalanobis distance, which captures the cryptic interactions within populations in the presence of LD better than the Euclidean distance.

59 BASIC BIOLOGICAL SCIENCES↗

Hybrid Cyber-attack Detection in Photovoltaic Farms

Here, to address the cyber-physical security in PV farms, a hybrid cyber-attack detection is proposed in this manuscript. To secure PV farms, the proposed method integrates model-based and data-driven methods by fusing the detection score at the device and system levels. First, a model-based cyber-attack detection method is developed for each PV inverter. A residual between the estimation of the Kalman filter and measurement is calculated. By leveraging the calculated residual from all inverters, a squared Mahalanobis distance is developed for device detection score generation. At the system level, a convolutional neural network (CNN) is proposed to detect cyber-attack using the waveform data at the point of common coupling (PCC) in PV farms. To improve the CNN detection accuracy, a set of well-designed features are extracted from the raw waveform data. Finally, a weighted detection score fusion method is proposed to combine device and system detection scores by using their complementary strength. The feasibility and robustness of the proposed method are validated by testing cases and a comparative experiment.

14 SOLAR ENERGY↗

Multidimensional Risk Analysis: MRISK

Multidimensional Risk (MRISK) calculates the combined multidimensional score using Mahalanobis distance. MRISK accounts for covariance between consequence dimensions, which de-conflicts the interdependencies of consequence dimensions, providing a clearer depiction of risks. Additionally, in the event the dimensions are not correlated, Mahalanobis distance reduces to Euclidean distance normalized by the variance and, therefore, represents the most flexible and optimal method to combine dimensions. MRISK is currently being used in NASA's Environmentally Responsible Aviation (ERA) project o assess risk and prioritize scarce resources.

McCollum, Raymond↗

Gradient-Based Novelty Detection Boosted by Self-Supervised Binary Classification

Novelty detection aims to automatically identify out-of-distribution (OOD) data, without any prior knowledge of them. It is a critical step in data monitoring, behavior analysis and other applications, helping enable continual learning in the field. Conventional methods of OOD detection perform multi-variate analysis on an ensemble of data or features, and usually resort to the supervision with OOD data to improve the accuracy. In reality, such supervision is impractical as one cannot anticipate the anomalous data. In this paper, we propose a novel, self-supervised approach that does not rely on any pre-defined OOD data: (1) The new method evaluates the Mahalanobis distance of the gradients between the in-distribution and OOD data. (2) It is assisted by a self-supervised binary classifier to guide the label selection to generate the gradients, and maximize the Mahalanobis distance. In the evaluation with multiple datasets, such as CIFAR-10, CIFAR-100, SVHN and TinyImageNet, the proposed approach consistently outperforms state-of-the-art supervised and unsupervised methods in the area under the receiver operating characteristic (AUROC) and area under the precision-recall curve (AUPR) metrics. We further demonstrate that this detector is able to accurately learn one OOD class in continual learning.

Sun, Jingbo↗

Substructure in the stellar halo near the Sun: I. Data-driven clustering in integrals-of-motion space

Context. Merger debris is expected to populate the stellar haloes of galaxies. In the case of the Milky Way, this debris should be apparent as clumps in a space defined by the orbital integrals of motion of the stars. Aims. Our aim is to develop a data-driven and statistics-based method for finding these clumps in integrals-of-motion space for nearby halo stars and to evaluate their significance robustly. Methods. We used data from Gaia EDR3, extended with radial velocities from ground-based spectroscopic surveys, to construct a sample of halo stars within 2.5 kpc from the Sun. We applied a hierarchical clustering method that makes exhaustive use of the single linkage algorithm in three-dimensional space defined by the commonly used integrals of motion energy E, together with two components of the angular momentum, L z and L ⊥ . To evaluate the statistical significance of the clusters, we compared the density within an ellipsoidal region centred on the cluster to that of random sets with similar global dynamical properties. By selecting the signal at the location of their maximum statistical significance in the hierarchical tree, we extracted a set of significant unique clusters. By describing these clusters with ellipsoids, we estimated the proximity of a star to the cluster centre using the Mahalanobis distance. Additionally, we applied the HDBSCAN clustering algorithm in velocity space to each cluster to extract subgroups representing debris with different orbital phases. Results. Our procedure identifies 67 highly significant clusters (> 3σ), containing 12% of the sources in our halo set, and 232 subgroups or individual streams in velocity space. In total, 13.8% of the stars in our data set can be confidently associated with a significant cluster based on their Mahalanobis distance. Inspection of the hierarchical tree describing our data set reveals a complex web of relations between the significant clusters, suggesting that they can be tentatively grouped into at least six main large structures, many of which can be associated with previously identified halo substructures, and a number of independent substructures. This preliminary conclusion is further explored in a companion paper, in which we also characterise the substructures in terms of their stellar populations. Conclusions. Our method allows us to systematically detect kinematic substructures in the Galactic stellar halo with a data-driven and interpretable algorithm. The list of the clusters and the associated star catalogue are provided in two tables available at the CDS.

79 ASTRONOMY AND ASTROPHYSICS↗

NMR Spectroscopy for Protein Higher Order Structure Similarity Assessment in Formulated Drug Products

Peptide and protein drug molecules fold into higher order structures (HOS) in formulation and these folded structures are often critical for drug efficacy and safety. Generic or biosimilar drug products (DPs) need to show similar HOS to the reference product. The solution NMR spectroscopy is a non-invasive, chemically and structurally specific analytical method that is ideal for characterizing protein therapeutics in formulation. However, only limited NMR studies have been performed directly on marketed DPs and questions remain on how to quantitively define similarity. Here, NMR spectra were collected on marketed peptide and protein DPs, including calcitonin-salmon, liraglutide, teriparatide, exenatide, insulin glargine and rituximab. The 1D 1 H spectral pattern readily revealed protein HOS heterogeneity, exchange and oligomerization in the different formulations. Principal component analysis (PCA) applied to two rituximab DPs showed consistent results with the previously demonstrated similarity metrics of Mahalanobis distance (D M ) of 3.3. The 2D 1 H- 13 C HSQC spectral comparison of insulin glargine DPs provided similarity metrics for chemical shift difference (Δδ) and methyl peak profile, i.e., 4 ppb for 1 H, 15 ppb for 13 C and 98% peaks with equivalent peak height. Finally, 2D 1 H- 15 N sofast HMQC was demonstrated as a sensitive method for comparison of small protein HOS. The application of NMR procedures and chemometric analysis on therapeutic proteins offer quantitative similarity assessments of DPs with practically achievable similarity metrics.

59 BASIC BIOLOGICAL SCIENCES↗

Time Series Classification for Locating Forced Oscillation Sources

Here, this article presents a machine learning based time-series classification method for using synchrophasor measurements to locate the source of forced oscillation (FO) for fast disturbance removal. First, multivariate time series (MTS) matrices are constructed by the most informative measurements selected by sequential feature selection from each power plant. Then, the Mahalanobis matrix is trained such that the Mahalanobis distance between the MTSs from the same class (i.e., with the same FO source location) are minimized and from different classes (i.e., with different FO source locations) are maximized. This allows MTSs to be classified by classifiers with class membership corresponding to the location of each FO source. To meet the runtime requirements of online matching, class templates are constructed to reduce data size and improve matching efficiency. To account for uncertainty in identifying the exact beginning of an FO event, dynamic time warping is used to align the out-of-sync MTSs. IEEE 39bus and WECC 179bus systems are used for algorithm development and validation. Simulation results demonstrate that the algorithm meets online operation runtime requirement with high accuracy using misaligned data sets.

42 ENGINEERING↗

Driving mode analysis—How uncertain functional inputs propagate to an output

Abstract Driving mode analysis elucidates how correlated features of uncertain functional inputs jointly propagate to produce uncertainty in the output of a computation. Uncertain input functions are decomposed into three terms: the mean functions, a zero‐mean driving mode, and zero‐mean residual. The random driving mode varies along a single direction, having fixed functional shape and random scale. It is uncorrelated with the residual, and under linear error propagation, it produces an output variance equal to that of the full input uncertainty. Finally, the driving mode best represents how input uncertainties propagate to the output because it minimizes expected squared Mahalanobis distance amongst competitors. These characteristics recommend interpretation of the driving mode as the single‐degree‐of‐freedom component of input uncertainty that drives output uncertainty. We derive the functional driving mode, show its superiority to other seemingly sensible definitions, and demonstrate the utility of driving mode analysis in an application. The application is the simulation of neutron transport in criticality experiments. The uncertain input functions are nuclear data that describe how Pu reacts to bombardment by neutrons. Visualization of the driving mode helps scientists understand what aspects of correlated functional uncertainty have effects that either reinforce or cancel one another in propagating to the output of the simulation.

97 MATHEMATICS AND COMPUTING↗

Alteration mapping at Goldfield, Nevada, by cluster and discriminant analysis of LANDSAT digital data

The ability of Landsat multispectral digital data to differentiate among 62 combinations of rock and alteration types at the Goldfield mining district of Western Nevada was investigated by using statistical techniques of cluster and discriminant analysis. Multivariate discriminant analysis was not effective in classifying each of the 62 groups, with classification results essentially the same whether data of four channels alone or combined with six ratios of channels were used. Bivariate plots of group means revealed a cluster of three groups including mill tailings, basalt and all other rock and alteration types. Automatic hierarchical clustering based on the fourth dimensional Mahalanobis distance between group means of 30 groups having five or more samples was performed. The results of the cluster analysis revealed hierarchies of mill tailings vs. natural materials, basalt vs. non-basalt, highly reflectant rocks vs. other rocks and exclusively unaltered rocks vs. predominantly altered rocks. The hierarchies were used to determine the order in which sets of multiple discriminant analyses were to be performed and the resulting discriminant functions were used to produce a map of geology and alteration which has an overall accuracy of 70 percent for discriminating exclusively altered rocks from predominantly altered rocks.

Ballew, G.↗

Alteration mapping at Goldfield, Nevada, by cluster and discriminant analysis of Landsat digital data

The ability of Landsat multispectral digital data to differentiate among 62 combinations of rock and alteration types at the Goldfield mining district of Western Nevada was investigated by using statistical techniques of cluster and discriminant analysis. Multivariate discriminant analysis was not effective in classifying each of the 62 groups, with classification results essentially the same whether data of four channels alone or combined with six ratios of channels were used. Bivariate plots of group means revealed a cluster of three groups including mill tailings, basalt and all other rock and alteration types. Automatic hierarchical clustering based on the fourth dimensional Mahalanobis distance between group means of 30 groups having five or more samples was performed using Johnson's HICLUS program. The results of the cluster analysis revealed hierarchies of mill tailings vs. natural materials, basalt vs. non-basalt, highly reflectant rocks vs. other rocks and exclusively unaltered rocks vs. predominantly altered rocks. The hierarchies were used to determine the order in which sets of multiple discriminant analyses were to be performed and the resulting discriminant functions were used to produce a map of geology and alteration which has an overall accuracy of 70 percent for discriminating exclusively altered rocks from predominantly altered rocks.

Ballew, G.↗

Unsupervised Classification of Global Radar Units on Venus

Characterization of the Venusian surface in terms of its radar properties was accomplished by application of an unsupervised, linear discriminant algorithm to two Pioneer-Venus (PV) Orbiter radar data sets: the RMS-slope (surface roughness) and reflectivity. Both databases were spatially filtered to the same effective resolution of 100 km prior to classification. A recent supervised classification study using these data was based on presupposed morphologic significance of selected data ranges. The knowledge of both Venusian geology and the geologic significance of the radar data is so limited that the data warrant a more unsupervised approach; for this study a linear discriminant classifier was chosen. This approach is purely statistical, thereby removing any observer bias. Statistical significance of the resulting clusters was evaluated by an ancillary program in which an F test utilizing the Mahalanobis' distance.

Kozak, R. C.↗

Earth Observing System Covariance Realism

The purpose of covariance realism is to properly size a primary object's covariance in order to add validity to the calculation of the probability of collision. The covariance realism technique in this paper consists of three parts: collection/calculation of definitive state estimates through orbit determination, calculation of covariance realism test statistics at each covariance propagation point, and proper assessment of those test statistics. An empirical cumulative distribution function (ECDF) Goodness-of-Fit (GOF) method is employed to determine if a covariance is properly sized by comparing the empirical distribution of Mahalanobis distance calculations to the hypothesized parent 3-DoF chi-squared distribution. To realistically size a covariance for collision probability calculations, this study uses a state noise compensation algorithm that adds process noise to the definitive epoch covariance to account for uncertainty in the force model. Process noise is added until the GOF tests pass a group significance level threshold. The results of this study indicate that when outliers attributed to persistently high or extreme levels of solar activity are removed, the aforementioned covariance realism compensation method produces a tuned covariance with up to 80 to 90% of the covariance propagation timespan passing (against a 60% minimum passing threshold) the GOF tests-a quite satisfactory and useful result.

Realism↗