Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Mahalanobis distance”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

23 records · Page 2

Linking Remotely Sensed Aerosol Types to Their Chemical Composition

Aerosol types measured during the Ship-Aircraft Bio-Optical Research (SABOR) experiment are related to GEOS-Chem model chemical composition. The application for this procedure to link model chemical components to aerosol type is desirable for understanding aerosol evolution over time. The Mahalanobis distance (DM) statistic is used to cluster model groupings of five chemical components (organic carbon, black carbon, sea salt, dust and sulfate) in a way analogous to the methods used by Burton et al. [2012] and Russell et al. [2014]. First, model-to-measurement evaluation is performed by collocating vertically resolved aerosol extinction from SABOR High Spectral Resolution LiDAR (HSRL) to the GEOS-Chem nested high-resolution data. Comparisons of modeled-to-measured aerosol extinction are shown to be within 35% +/- 14%. Second, the model chemical components are calculation into five variables to calculate the DM and cluster means and covariances for each HSRL-retrieved aerosol type. The layer variables from the model are aerosol optical depth (AOD) ratios of (i) sea salt and (ii) dust to total AOD, mass ratios of (iii) total carbon (i.e. sum of organic and black carbon) to the sum of total carbon and sulfate (iv) organic carbon to black carbon, and (v) the natural log of the aerosol-to-molecular extinction ratio. Third, the layer variables and at most five out of twenty SABOR flights are used to form the pre-specified clusters for calculating DM and to assign an aerosol type. After determining the pre-specified clusters, model aerosol types are produced for the entire vertically resolved GEOS-Chem nested domain over the United States and the model chemical component distributions relating to each type are recorded. Resulting aerosol types are Dust/Dusty Mix, Maritime, Smoke, Urban and Fresh Smoke (separated into 'dark' and 'light' by a threshold of the organic to black carbon ratio). Model-calculated DM not belonging to a specific type (i.e. not meeting a threshold probability) is termed an outlier and those DM values that can belong to multiple types (i.e. showing weak probability of belonging to a specific cluster) are termed as Overlap. MODIS active fires are overlaid on the model domain to qualitatively evaluate the model-predicted Smoke aerosol types.

Aerosol Types↗

Matrix Theory for Data Association in PVS

Consider a collection of data generated by sensors from a set of aircraft. Data association is the process of connecting each sensor measurement with its corresponding aircraft. Furthermore once the data association has taken place, the state of the aircraft can be approximated using a Kalman filter. This talk aims to explore formal specification and verification of data association in the Prototype Verification System (PVS). Formal specification and verification of data association includes development of Kalman filters, Mahalanobis distance, and other topics of matrix analysis in PVS.

Linear Algebra↗

Self-Supervised Anomaly Detection via Neural Autoregressive Flows with Active Learning

Many self-supervised methods have been proposed with the target of image anomaly detection. These methods often rely on the paradigm of data augmentation with predefined transformations such as flipping, cropping, and rotations. However, it is not straightforward to apply these techniques for non-image data, such as time series or tabular data, while the performance of the existing deep approaches has been under our expectation on tasks beyond images. In this work, we propose a novel active learning (AL) scheme that relied on neural autoregressive flows (NAF) for self-supervised anomaly detection, specifically on small-scale data. Unlike other generative models such as GANs or VAEs, flow-based models allow to explicitly learn the probability density and thus can assign accurate likelihoods to normal data which makes it usable to detect anomalies. The proposed NAF-AL method is achieved by efficiently generating random samples from latent space and transforming them into feature space along with likelihoods via invertible mapping. The samples with lower likelihoods are selected and further checked by outlier detection using Mahalanobis distance. The augmented samples incorporating with normal samples are used for training a better detector so as to approach decision boundaries. Compared with random transformations, NAF-AL can be interpreted as a likelihood-oriented data augmentation that is more efficient and robust. Extensive experiments show that our approach outperforms existing baselines on multiple time series and tabular datasets, and a real-world application in advanced manufacturing, with significant improvement on anomaly detection accuracy and robustness over the state-of-the-art.

Zhang, Jiaxin↗

Prospects for Astrobiology and Technosignature Searches with the Vera C. Rubin Observatory Legacy Survey of Space and Time

The Vera C. Rubin Observatory Legacy Survey of Space and Time (LSST) will map sources in multiband colour--variability space. We present a prototype coherence-based framework for astrobiology and technosignature searches, in which candidates are treated as structured departures from natural astrophysical manifolds rather than isolated photometric outliers. We illustrate the framework with three simulated cases: five Kuiper Belt Object (KBO) surface/activity states, a grid of 649 synthetic exoplanet spectra with vegetation-red-edge-like (VRE) perturbations, and 500 synthetic multiband light curves, each projected into LSST-like observable space and analysed through colour geometry, chromatic variability, and cross-band coherence. Key results include a full-colour Mahalanobis distance $D\approx5.1$ for the weak-coma KBO state (${\sim}5σ$ in the five-dimensional colour vector), an indicative VRE coherence threshold at $f_{\rm crit}\approx0.13$, and an idealised stacking forecast reaching $5σ$ under optimistic assumptions. We show, using a small Gaia~DR3 stellar sample, that stellar colour and photometric stability may inform the prioritisation of Galactic regions for applying such coherence diagnostics.

Kovačević, Andjelka B. [Belgrade U.] (ORCID:000000↗

Comparison of linear regression, k-nearest neighbour and random forest methods in airborne laser-scanning-based prediction of growing stock

Abstract In this study, for five sites around the world, we look at the effects of different model types and variable selection approaches on forest yield modelling performances in an area-based approach (ABA). We compared ordinary least squares regression (OLS), k-nearest neighbours (kNN) and random forest (RF). Our objective was to test if there are systematic differences in accuracy between OLS, kNN and RF in ABA predictions of growing stock volume. The analyses are based on a 5-fold cross-validation at five study sites: an eucalyptus plantation, a temperate forest and three different boreal forests. Two completely independent validation datasets were also available for two of the boreal sites. For the kNN, we evaluated multiple measures of distance including Euclidean, Mahalanobis, most similar neighbour (MSN) and an RF-based distance metric. The variable selection approaches we examined included a heuristic approach (for OLS, kNN and RF), exhaustive search among all combinations (OLS only) and all variables together (RF only). Performances varied by model type and variable selection approaches among sites. OLS and RF had similar accuracies and were more efficient than any of the kNN variants. Variable selection did not affect RF performance. Heuristic and exhaustive variable selection performed similarly for OLS. kNN fared the poorest amongst model types, and kNN with RF distance was prone to overfitting when compared with a validation dataset. Additional caution is therefore required when building kNN models for volume prediction though ABA, being preferable instead to opt for models based on OLS with some variable selection, or RF with all variables together.

Cosenza, Diogo N.↗