Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Outlier detection”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Recent Advances in SMAP RFI Processing

The measurements made by the Soil Moisture Active/Passive (SMAP) mission are affected by the presence of Radio Frequency Interference (RFI) in the protected 1400-1427 MHz band. In SMAP data processing, the main protection against RFI is a sophisticated RFI detection algorithm which flags sub-samples in time and frequency that are contaminated by RFI and removes them before estimating the brightness temperature. This contribution presents two additional approaches that have been developed to address the RFI concern in SMAP. The first consists in locating sources of RFI; once located, it becomes possible to report RFI sources to spectrum management authorities, which can lead to less RFI being experienced by SMAP in the future. The second is a new RFI detection method that is based on detecting outliers in the spatial distribution of measured antenna temperatures.

Radio Frequency Interference↗

PVAnalytics: A Python Package for Automated Processing of Solar Time Series Data

Multiple publicly available software packages exist that analyze solar time series data, including RdTools and Solar Data Tools, among others. Several of these packages contain their own unique quality assurance (QA) and feature recognition algorithms. The python PVAnalytics package was developed to offer an internally consistent source for these analysis tools, making it easier for the end user to deploy these routines on his or her solar data. The PVAnalytics package currently contains routines for outlier detection, inverter clipping detection, irradiance and temperature checks, orientation checks, and data shift detection, among other functions. These functions have been aggregated from various sources including Solar Forecast Arbiter, RdTools, and the QA process developed by NREL's PV Fleets Initiative. We are continuously adding new functionality to the package, including documentation, examples and algorithms. By bundling QA functionality into a single software package, we hope to make PVAnalytics a comprehensive software library to support analysis of solar metadata and time series data.

data cleaning↗

Anomaly Detection in Large Sets of High-Dimensional Symbol Sequences

This paper addresses the problem of detecting and describing anomalies in large sets of high-dimensional symbol sequences. The approach taken uses unsupervised clustering of sequences using the normalized longest common subsequence (LCS) as a similarity measure, followed by detailed analysis of outliers to detect anomalies. As the LCS measure is expensive to compute, the first part of the paper discusses existing algorithms, such as the Hunt-Szymanski algorithm, that have low time-complexity. We then discuss why these algorithms often do not work well in practice and present a new hybrid algorithm for computing the LCS that, in our tests, outperforms the Hunt-Szymanski algorithm by a factor of five. The second part of the paper presents new algorithms for outlier analysis that provide comprehensible indicators as to why a particular sequence was deemed to be an outlier. The algorithms provide a coherent description to an analyst of the anomalies in the sequence, compared to more normal sequences. The algorithms we present are general and domain-independent, so we discuss applications in related areas such as anomaly detection.

Budalakoti, Suratna↗

Genetic monitoring of steelhead in the Klickitat River to estimate productivity, straying, and migration timing

Abstract Objective Salmonids with a complex life history variation present challenges for conservation management, but genetic approaches alongside fisheries monitoring can address questions regarding the viability of the natural populations. Methods We genotyped adult (n = 3108) and juvenile (n = 2624) samples of steelhead Oncorhynchus mykiss that were collected in the Klickitat River, Washington, USA at traps in the lower drainage to examine tributary level productivity, straying from outside sources, and variation in adult migration timing. Result Genetic assignment of steelhead from this system indicated that the majority were produced within or near tributaries of the middle Klickitat River (juvenile mean = 72.8%; adult mean = 87.3%). Analyses with parentage-based tagging identified that most hatchery-origin adults were assigned to the Skamania Hatchery (80.8%), as expected, since this has been the release stock for decades within the Klickitat River drainage. Hatchery-origin adults were also identified from programs operating outside the Klickitat River, which were primarily strays from Snake River hatcheries. Most natural-origin steelhead were assigned to the Klickitat River, but there were also natural-origin fish identified as strays from other regions of the Columbia River (22.3% of natural returns). We also examined genes known to be associated with migration timing in adult steelhead observed at the trap and observed a strong relationship between migration date and alleles for early and late migration, but individual outliers were detected across seasons. Conclusion Our results indicate that genetic variation of steelhead in the Klickitat River has been influenced by hatchery programs as well as natural-origin straying from other subbasins, but genetic diversity remains high throughout the subbasin, and both early and late migration alleles are maintained. The genetic diversity present in Klickitat River steelhead may enable this Endangered Species Act listed (threatened) species to better adapt to stochastic environmental conditions compared to less diverse populations.

Collins, Erin E. (ORCID:0000000260978479)↗

Assessment of Outliers in Alloy Datasets Using Unsupervised Techniques

We report advancements in data analytics techniques have enabled complex, disparate datasets to be leveraged for alloy design. Identifying outliers in a dataset can reduce noise, identify erroneous and/or anomalous records, prevent overfitting, and improve model assessment and optimization. In this work, two alloy datasets (9-12% Cr ferritic martensitic steels, and austenitic stainless steels) have been assessed for outliers using unsupervised techniques and supplemented with domain knowledge. Principal component analysis and k-means clustering were applied to the data, and points were assessed as outliers based on their distance away from other points in the cluster and from other points in the dataset. The outlier characteristics were investigated to determine both cluster-specific and overall trends in the properties of the outlier points. The approach demonstrated here is extensible to other alloy datasets for outlier identification and evaluation to improve the reliability of machine learning and modeling predictions for advanced alloy design.

36 MATERIALS SCIENCE↗

Recent developments with the ORSER system

Additions to the ORSER remote sensing data processing package are described. The ORSER package consists of about 35 individual programs that are grouped into preprocessing, data analysis, and display subsystems. Additional data formats and data management, data transformation, and geometric correlation programs were supplemented to the preprocessing subsystem. Enhancements to the data analysis techniques include a maximum likelihood classifier (MAXCLASS) and a new version of the STATS program which makes delineation of training areas easier and allows for detection of outlier points. Ongoing developments are also described.

Baumer, G. M.↗

An Automated Algorithm to Screen Massive Training Samples for a Global Impervious Surface Classification

An algorithm is developed to automatically screen the outliers from massive training samples for Global Land Survey - Imperviousness Mapping Project (GLS-IMP). GLS-IMP is to produce a global 30 m spatial resolution impervious cover data set for years 2000 and 2010 based on the Landsat Global Land Survey (GLS) data set. This unprecedented high resolution impervious cover data set is not only significant to the urbanization studies but also desired by the global carbon, hydrology, and energy balance researches. A supervised classification method, regression tree, is applied in this project. A set of accurate training samples is the key to the supervised classifications. Here we developed the global scale training samples from 1 m or so resolution fine resolution satellite data (Quickbird and Worldview2), and then aggregate the fine resolution impervious cover map to 30 m resolution. In order to improve the classification accuracy, the training samples should be screened before used to train the regression tree. It is impossible to manually screen 30 m resolution training samples collected globally. For example, in Europe only, there are 174 training sites. The size of the sites ranges from 4.5 km by 4.5 km to 8.1 km by 3.6 km. The amount training samples are over six millions. Therefore, we develop this automated statistic based algorithm to screen the training samples in two levels: site and scene level. At the site level, all the training samples are divided to 10 groups according to the percentage of the impervious surface within a sample pixel. The samples following in each 10% forms one group. For each group, both univariate and multivariate outliers are detected and removed. Then the screen process escalates to the scene level. A similar screen process but with a looser threshold is applied on the scene level considering the possible variance due to the site difference. We do not perform the screen process across the scenes because the scenes might vary due to the phenology, solar-view geometry, and atmospheric condition etc. factors but not actual landcover difference. Finally, we will compare the classification results from screened and unscreened training samples to assess the improvement achieved by cleaning up the training samples. Keywords:

Tan, Bin↗

Optical Navigation Attitude Estimation and Calibration Performance Improvement using Outlier Rejection

Spacecraft optical navigation (OpNav) systems process a sequence of images of celestial bodies against a starfield background to estimate the position and velocity of the vehicle. While attitude is sometimes available from an onboard star tracker, it is often desirable to recognize the background stars in the OpNav images to better align the image. While many image processing algorithms exist for finding stars, efficiency and reliability remain key issues in the presence of extended bodies(e.g. the Moon, Earth), especially when attempting to solve the full lost-in-space problem. Some star outliers(stars identified with high residuals)could appear in the camera field of view, however using them in the attitude estimation or camera calibration would lead to less accurate results. Therefore, we require new and robust approaches to remove these outliers before any further processing. The emphasis of the work is on developing a simple and robust iterative technique to detect and reject the outliers which could be found in any frame during the lost in space attitude determination or during the camera calibration. These outliers are determined based on the residuals of the centroids of the detected stars and the corresponding location using the star catalog. If the residuals exceed a predetermined threshold value, the object will be detected as an outlier and will be removed before another attitude determination and calibration iteration is performed. The performance for both attitude determination and on-orbit camera calibration are improved by an almost two-fold increase in accuracy when applying this outlier rejection technique.

OpNav↗

The Harmonized Landsat and Sentinel-2 Surface Reflectance Data Set

The Harmonized Landsat and Sentinel-2 (HLS) project is a NASA initiative aiming to produce a VirtualConstellation (VC) of surface reflectance (SR) data acquired by the Operational Land Imager (OLI) and MultiSpectral Instrument (MSI) aboard Landsat 8 and Sentinel-2 remote sensing satellites, respectively. The HLS products are based on a set of algorithms to obtain seamless products from both sensors (OLI and MSI): atmospheric correction, cloud and cloud-shadow masking, spatial co-registration and common gridding, bidirectional reflectance distribution function normalization and spectral bandpass adjustment. Three products are derivedfrom the HLS processing chain: (i) S10: full resolution MSI SR at 10 m, 20 m and 60 m spatial resolutions; (ii)S30: a 30 m MSI Nadir BRDF (Bidirectional Reflectance Distribution Function)-Adjusted Reflectance (NBAR);(iii) L30: a 30 m OLI NBAR. All three products are processed for every Level-1 input products from Landsat 8/OLI (L1T) and Sentinel-2/MSI (L1C). As of version 1.3, the HLS data set covers 10.35 million km2 and spans from first Landsat 8 data (2013); Sentinel-2 data spans from October 2015. The L30 and S30 show a good consistency with coarse spatial resolution products, in particular MODIS Collection 6 MCD09CMG products (overall deviations do not exceed 11%) that are used as a reference for quality assurance. The spatial co-registration of the HLS is improved compared to original Landsat 8 L1T and Sentinel 2A L1C products, for which misregistration issues between multi-temporal data are known. In particular, the resulting computed circular errors at 90% for the HLS product are 6.2 m and 18.8 m, for S10 and L30 products, respectively. The main known issue of the current data set remains the Sentinel-2 cloud mask with many cloud detection omissions. The cross-comparison with MODIS was used to flag products with most evident non-detected clouds. A time series outlier filtering approach is suggested to detect remaining clouds. Finally, several time series are presented to highlight the high potential of the HLS data set for crop monitoring.

Landsat Sentinel-2↗

Automatic Feature Tracking on Small Bodies for Autonomous Approach

Abstract—The autonomous approach of a spacecraft to an asteroid or comet (a small body) relies heavily on visual feature tracking to aid in estimating relative trajectories and the properties of the small body. Feature tracking for small bodies brings several challenges, including changing lighting, poor visual texture, and a concentration of features in a small part of an image. Six existing, open-source algorithms for feature tracking were tested on a simulated dataset and compared to the ground truth in the path of features. The main finding is that none of the algorithms provide all of the desired characteristics of long feature tracks with low errors and few outliers. Instead, there is a trade-off between long feature tracks and low error. The feature-matching algorithms SIFT, and BRISK provide good error characteristics, but short feature tracks, whereas the optical flow algorithm KLT provides long feature tracks, but with many features of large error. Given the challenges in feature tracking, it is recommended to focus development on each component of a feature tracking system: detection, description, and outlier rejection.

Morrell, Benjamin J↗

A Probabilistic Autoencoder for Type Ia Supernova Spectral Time Series

We construct a physically parameterized probabilistic autoencoder (PAE) to learn the intrinsic diversity of Type Ia supernovae (SNe Ia) from a sparse set of spectral time series. The PAE is a two-stage generative model, composed of an autoencoder that is interpreted probabilistically after training using a normalizing flow. We demonstrate that the PAE learns a low-dimensional latent space that captures the nonlinear range of features that exists within the population and can accurately model the spectral evolution of SNe Ia across the full range of wavelength and observation times directly from the data. By introducing a correlation penalty term and multistage training setup alongside our physically parameterized network, we show that intrinsic and extrinsic modes of variability can be separated during training, removing the need for the additional models to perform magnitude standardization. We then use our PAE in a number of downstream tasks on SNe Ia for increasingly precise cosmological analyses, including the automatic detection of SN outliers, the generation of samples consistent with the data distribution, and solving the inverse problem in the presence of noisy and incomplete data to constrain cosmological distance measurements. We find that the optimal number of intrinsic model parameters appears to be three, in line with previous studies, and show that we can standardize our test sample of SNe Ia with an rms of 0.091 ± 0.010 mag, which corresponds to 0.074 ± 0.010 mag if peculiar velocity contributions are removed.

79 ASTRONOMY AND ASTROPHYSICS↗

Searching for Novel Chemistry in Exoplanetary Atmospheres Using Machine Learning for Anomaly Detection

Abstract The next generation of telescopes will yield a substantial increase in the availability of high-quality spectroscopic data for thousands of exoplanets. The sheer volume of data and number of planets to be analyzed greatly motivate the development of new, fast, and efficient methods for flagging interesting planets for reobservation and detailed analysis. We advocate the application of machine learning (ML) techniques for anomaly (novelty) detection to exoplanet transit spectra, with the goal of identifying planets with unusual chemical composition and even searching for unknown biosignatures. We successfully demonstrate the feasibility of two popular anomaly detection methods (local outlier factor and one-class support vector machine) on a large public database of synthetic spectra. We consider several test cases, each with different levels of instrumental noise. In each case, we use receiver operating characteristic curves to quantify and compare the performance of the two ML techniques.

Astronomy & Astrophysics↗

Mixture-Tuned, Clutter Matched Filter for Remote Detection of Subpixel Spectral Signals

Mapping localized spectral features in large images demands sensitive and robust detection algorithms. Two aspects of large images that can harm matched-filter detection performance are addressed simultaneously. First, multimodal backgrounds may thwart the typical Gaussian model. Second, outlier features can trigger false detections from large projections onto the target vector. Two state-of-the-art approaches are combined that independently address outlier false positives and multimodal backgrounds. The background clustering models multimodal backgrounds, and the mixture tuned matched filter (MT-MF) addresses outliers. Combining the two methods captures significant additional performance benefits. The resulting mixture tuned clutter matched filter (MT-CMF) shows effective performance on simulated and airborne datasets. The classical MNF transform was applied, followed by k-means clustering. Then, each cluster s mean, covariance, and the corresponding eigenvalues were estimated. This yields a cluster-specific matched filter estimate as well as a cluster- specific feasibility score to flag outlier false positives. The technology described is a proof of concept that may be employed in future target detection and mapping applications for remote imaging spectrometers. It is of most direct relevance to JPL proposals for airborne and orbital hyperspectral instruments. Applications include subpixel target detection in hyperspectral scenes for military surveillance. Earth science applications include mineralogical mapping, species discrimination for ecosystem health monitoring, and land use classification.

Thompson, David R.↗

Groundspeed filtering for CTAS

Ground speed is one of the radar observables which is obtained along with position and heading from NASA Ames Center radar. Within the Center TRACON Automation System (CTAS), groundspeed is converted into airspeed using the wind speeds which CTAS obtains from the NOAA weather grid. This airspeed is then used in the trajectory synthesis logic which computes the trajectory for each individual aircraft. The time history of the typical radar groundspeed data is generally quite noisy, with high frequency variations on the order of five knots, and occasional 'outliers' which can be significantly different from the probable true speed. To try to smooth out these speeds and make the ETA estimate less erratic, filtering of the ground speed is done within CTAS. In its base form, the CTAS filter is a 'moving average' filter which averages the last ten radar values. In addition, there is separate logic to detect and correct for 'outliers', and acceleration logic which limits the groundspeed change in adjacent time samples. As will be shown, these additional modifications do cause significant changes in the actual groundspeed filter output. The conclusion is that the current ground speed filter logic is unable to track accurately the speed variations observed on many aircraft. The Kalman filter logic however, appears to be an improvement to the current algorithm used to smooth ground speed variations, while being simpler and more efficient to implement. Additional logic which can test for true 'outliers' can easily be added by looking at the difference in the a priori and post priori Kalman estimates, and not updating if the difference in these quantities is too large.

Slater, Gary L.↗

Evolutionary Analyses of Gene Expression Divergence in Panicum hallii : Exploring Constitutive and Plastic Responses Using Reciprocal Transplants

Abstract The evolution of gene expression is thought to be an important mechanism of local adaptation and ecological speciation. Gene expression divergence occurs through the evolution of cis- polymorphisms and through more widespread effects driven by trans-regulatory factors. Here, we explore expression and sequence divergence in a large sample of Panicum hallii accessions encompassing the species range using a reciprocal transplantation experiment. We observed widespread genotype and transplant site drivers of expression divergence, with a limited number of genes exhibiting genotype-by-site interactions. We used a modified FST–QST outlier approach (QPC analysis) to detect local adaptation. We identified 514 genes with constitutive expression divergence above and beyond the levels expected under neutral processes. However, no plastic expression responses met our multiple testing correction as QPC outliers. Constitutive QPC outlier genes were involved in a number of developmental processes and responses to abiotic environments. Leveraging earlier expression quantitative trait loci results, we found a strong enrichment of expression divergence, including for QPC outliers, in genes previously identified with cis and cis–environment interactions but found no patterns related to trans-factors. Population genetic analyses detected elevated sequence divergence of promoters and coding sequence of constitutive expression outliers but little evidence for positive selection on these proteins. Our results are consistent with a hypothesis of cis-regulatory divergence as a primary driver of expression divergence in P. hallii.

3′ TagSeq↗

Signatures of Selection for Resistance/Tolerance to Perkinsus olseni in Grooved Carpet Shell Clam ( Ruditapes decussatus ) Using a Population Genomics Approach

ABSTRACT The grooved carpet shell clam ( Ruditapes decussatus ) is a bivalve of high commercial value distributed throughout the European coast. Its production has suffered a decline caused by different factors, especially by the parasite Perkinsus olsenii . Improving production of R . decussatus requires genomic resources to ascertain the genetic factors underlying resistance/tolerance to P. olseni i . In this study, the first reference genome of R . decussatus was assembled through long‐ and short‐read sequencing (1677 contigs; 1.386 Mb) and further scaffolded at chromosome level with Hi‐C (19 superscaffolds; 95.4% of assembly). Repetitive elements were identified (32%) and masked for annotation of 38,276 coding‐ and 13,056 non‐coding genes. This genome was used as a reference to develop a 2bRAD‐Seq 13,438 SNP panel for a genomic screening on six shellfish beds distributed across the Atlantic Ocean and Mediterranean Sea. Beds were selected by perkinsosis prevalence and the infection level was individually evaluated in all the samples. Genetic diversity was significantly higher in the Mediterranean than in the Atlantic region. The main genetic breakage was detected between those regions (F ST = 0.224), being the Mediterranean more heterogeneous than the Atlantic. Several loci under divergent selection (394 outliers; 261 genomic windows) were detected across shellfish beds. Samples were also inspected to detect signals of selection for resistance/tolerance to P. olseni i by using infection‐level and population‐genomics approaches, and 90 common divergent outliers for resistance/tolerance to perkinsosis were identified and used for gene mining. Candidate genes and markers identified provide invaluable information for controlling perkinsosis and for improving production of the grooved carpet shell clam.

Sambade, Inés M. [Department of Zoology, Genetics ↗