Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “outlier detection”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

pyFLANK, a graph neural network based null distribution inference model for F ST outlier detection

Detecting genomic regions under selection is essential for understanding how populations adapt to different environments, yet it remains challenging due to the confounding effects of demographic history and linkage disequilibrium (LD). Fixation index (F ST ) is a widely used statistic to identify genomic regions under adaptation. However, identifying genes under selection by defining F ST outliers often remains challenging, owing to confounding effects of underlying demographic history. Traditional methods assume independence among loci and rely on simple demographic models, while newer models perform much better but are computationally expensive and not easily scalable. Here, we present pyFLANK, an open-source and automated Python implementation which detects F ST outliers using a null distribution inferred from quasi-independent loci. Our tool integrates three approaches to identify loci obeying a null distribution: graph neural network (GNN) inference, linkage disequilibrium (LD)-based inference, and user-defined input. Because pyFLANK uses GNN-based inference of quasi-independent loci, it yields a more accurate null model with less need for user parameter input. In simulation experiments, pyFLANK achieved lower false positive rates than current methods while maintaining comparable detection power, indicating that its refined null model better distinguishes true adaptive loci from background variation. The GNN-based model, in particular, detected additional loci associated with phenotypic variance that were not identified by existing methods. Assessments of simulation and real data from different species demonstrate that pyFLANK achieves lower false positive rates compared with other commonly used F ST outlier detectors, while maintaining comparable detection power and excellent computational performance, providing a robust and user-friendly tool for identifying loci under divergent selection. It extends existing F ST outlier frameworks by incorporating explicit LD-aware strategies for null model calibration. The method is intended as a practical and scalable complement to existing genome scan approaches.

FST↗

Visualisation and outlier detection for probability density function ensembles

Abstract Exploratory data analysis (EDA) for functional data—data objects where observations are entire functions—is a difficult problem that has seen significant attention in recent literature. This surge in interest is motivated by the ubiquitous nature of functional data, which are prevalent in applications across fields such as meteorology, biology, medicine and engineering. Empirical probability density functions (PDFs) can be viewed as constrained functional data objects that must integrate to one and be nonnegative. They show up in contexts such as yearly income distributions, zooplankton size structure in oceanography and in connectivity patterns in the brain, among others. While PDF data are certainly common in modern research, little attention has been given to EDA specifically for PDFs. In this paper, we extend several methods for EDA on functional data for PDFs and compare them on simulated data that exhibit different types of variation, designed to mimic that seen in real‐world applications. We then use our new methods to perform EDA on the breakthrough curves observed in gas transport simulations for underground fracture networks.

97 MATHEMATICS AND COMPUTING↗

A simple data-driven level finding method of many-electron atoms and heavy nuclei based on statistical outlier detection

Here, we report a simple and pure data-driven method to find new energy levels of quantum many-body systems only from observed line wavelengths. In our method, all the possible combinations are computed from known energy levels and wavelengths of unidentified lines. As each excited state exhibits many transition lines to different lower levels, the true levels should be reconstructed coincidentally from many level-line combinations, while the wrong combinations distribute randomly. Such a coincidence can be easily detected statistically. We demonstrate this statistical method by finding new levels for various atomic and nuclear systems from unidentified line lists available online.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

BROOD: Bilevel and Robust Optimization and Outlier Detection for Efficient Tuning of High-Energy Physics Event Generators

The parameters in Monte Carlo (MC) event generators are tuned on experimental measurements by evaluating the goodness of fit between the data and the MC predictions. The relative importance of each measurement is adjusted manually in an often time-consuming, iterative process to meet different experimental needs. In this work, we introduce several optimization formulations and algorithms with new decision criteria for streamlining and automating this process. These algorithms are designed for two formulations: bilevel optimization and robust optimization. Both formulations are applied to the datasets used in the ATLAS A14 tune and to the dedicated hadronization datasets generated by the SHERPA generator, respectively. The corresponding tuned generator parameters are compared using three metrics. We compare the quality of our automatic tunes to the published ATLAS A14 tune. Moreover, we analyze the impact of a pre-processing step that excludes data that cannot be described by the physics models used in the MC event generators.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Advancing Vision-based Feedback and Convolutional Neural Networks for Visual Outlier Detection

Machine learning has matured into a technology that has immediate applicability to the surveillance needs of nuclear material storage containers. These containers at LANL are the barrier preventing release of radioactive material to the workers, public, and environment during the storage period of the material. Annual surveillance activities can only provide coverage on a handful of containers. There is a significant need for surveillance tools to identify potential issues and precursors to containment failure that can be used during opportunistic inspections and, more generally, outside of annual surveillance activities. In this report we provide details on the advancement of our proposed embodiment that combines an automation system for taking pictures and a high-accuracy machine learning-driven object detection software. We further showcase the improvements on the software side with progress on extracting unique identification features and advances in detecting damage. The current state of the system captures subject matter expert training and a space-conscious design whose implementation is envisioned in the near future.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

Detection of Outliers in LiDAR Data Acquired by Multiple Platforms over Sorghum and Maize

High-resolution point cloud data acquired with a laser scanner from any platform contain random noise and outliers. Therefore, outlier detection in LiDAR data is often necessary prior to analysis. Applications in agriculture are particularly challenging, as there is typically no prior knowledge of the statistical distribution of points, plant complexity, and local point densities, which are crop-dependent. The goals of this study were first to investigate approaches to minimize the impact of outliers on LiDAR acquired over agricultural row crops, and specifically for sorghum and maize breeding experiments, by an unmanned aerial vehicle (UAV) and a wheel-based ground platform; second, to evaluate the impact of existing outliers in the datasets on leaf area index (LAI) prediction using LiDAR data. Two methods were investigated to detect and remove the outliers from the plant datasets. The first was based on surface fitting to noisy point cloud data via normal and curvature estimation in a local neighborhood. The second utilized the PointCleanNet deep learning framework. Both methods were applied to individual plants and field-based datasets. To evaluate the method, an F-score was calculated for synthetic data in the controlled conditions, and LAI, the variable being predicted, was computed both before and after outlier removal for both scenarios. Results indicate that the deep learning method for outlier detection is more robust than the geometric approach to changes in point densities, level of noise, and shapes. The prediction of LAI was also improved for the wheel-based vehicle data based on the coefficient of determination (R2) and the root mean squared error (RMSE) of the residuals before and after the removal of outliers.

36 MATERIALS SCIENCE↗

Defining blood hematology reference values in female pig-tailed macaques ( Macaca nemestrina ) using the Isolation Forest algorithm

Background: Pig-tailed macaques (PTMs) are commonly used as preclinical models to assess antiretroviral drugs for HIV prevention research. Drug toxicities and disease pathologies are often preceded by changes in blood hematology. To better assess the safety profile of pharmaceuticals, we defined normal ranges of hematological values in PTMs using an Isolation Forest (iForest) algorithm. Methods: Eighteen female PTMs were evaluated. Blood was collected 1–24 times per animal for a total of 159 samples. Complete blood counts were performed, and iForest was used to analyze the hematology data to detect outliers. Results: Median, IQR, and ranges were calculated for 13 hematology parameters. From all samples, 22 outliers were detected. These outliers were excluded from the reference index. Conclusions: Using iForest, we defined a normal range for hematology parameters in female PTMs. This reference index can be a valuable tool for future studies evaluating drug toxicities in PTMs.

59 BASIC BIOLOGICAL SCIENCES↗

Identifying Outliers in AI-based Image Compression

Image compression using artificial intelligence (AI) is becoming increasingly prevalent across various fields, including scientific research. Scientific instruments can generate hundreds of images per second, and effectively compressing these images with high compression ratios is crucial for facilitating scientific discoveries. However, automatically detecting outlier cases, where compression may not have succeeded or where interesting scientific phenomena are present, poses a significant challenge. To address this, we have developed a methodology based on unsupervised machine learning techniques for detecting outlier compressed images. This methodology utilizes metrics such as peak signal-to-noise ratio (PSNR), structural similarity index measure (SSIM), structural texture similarity index measure (STSIM), and deep image and structural texture similarity index (DISTS). We have evaluated our methodology on several unlabeled datasets, including microscopy and x-ray images, and have successfully identified multiple outlier images using our proposed approach. Furthermore, our approach has enabled us to identify image semantics that are valuable for post-experiment analysis by scientists.

Data Analysis↗

Multi-area parameter error identification for large power systems

Power grid model parameters may contain errors due to various reasons. Detecting and correcting parameter errors typically requires significant computational effort due to the size and complexity of the parameter database. While the normalized Lagrange multiplier (NLM) method can effectively detect, identify and correct parameter errors, its computational burden could rapidly grow with increasing system size. This paper addresses this issue by proposing a multi-area parameter error identification method. Each area has its own outlier detection tool for detecting the incorrect parameters and measurements within the area. On the other hand, due to the reduced redundancy at area boundaries, parameter errors on branches incident to boundary buses may not be detected. Such errors are subsequently detected by a coordination level estimator completing the system-wide parameter detection procedure. In conclusion, performance of the developed method is demonstrated using the IEEE 118-bus and 2000-bus Texas synthetic systems.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Self-Supervised Anomaly Detection via Neural Autoregressive Flows with Active Learning

Many self-supervised methods have been proposed with the target of image anomaly detection. These methods often rely on the paradigm of data augmentation with predefined transformations such as flipping, cropping, and rotations. However, it is not straightforward to apply these techniques for non-image data, such as time series or tabular data, while the performance of the existing deep approaches has been under our expectation on tasks beyond images. In this work, we propose a novel active learning (AL) scheme that relied on neural autoregressive flows (NAF) for self-supervised anomaly detection, specifically on small-scale data. Unlike other generative models such as GANs or VAEs, flow-based models allow to explicitly learn the probability density and thus can assign accurate likelihoods to normal data which makes it usable to detect anomalies. The proposed NAF-AL method is achieved by efficiently generating random samples from latent space and transforming them into feature space along with likelihoods via invertible mapping. The samples with lower likelihoods are selected and further checked by outlier detection using Mahalanobis distance. The augmented samples incorporating with normal samples are used for training a better detector so as to approach decision boundaries. Compared with random transformations, NAF-AL can be interpreted as a likelihood-oriented data augmentation that is more efficient and robust. Extensive experiments show that our approach outperforms existing baselines on multiple time series and tabular datasets, and a real-world application in advanced manufacturing, with significant improvement on anomaly detection accuracy and robustness over the state-of-the-art.

Zhang, Jiaxin↗

Accurate Prediction of Algal Biomass Lipid, Protein, and Carbohydrate Composition with Machine Learning Regression Modelling of Near-IR Spectra

During large scale algal biomass cultivation, it is difficult to reliably control relative composition to target levels. Rapid determination of chemical composition is feasible by using near infrared (NIR) spectral data. We sought to build and improve on reliable high-throughput screening prediction method based on partial least squares regression (PLSR) by the application of artificial neural networks (ANN) and associated optimization strategies. The algal biomass sample set was designed and created in an iterative process of culturing in physiologically diverse conditions at the GAI field site, followed by compositional analyses at NREL. The workflow allowed us to identify gaps in compositional space for informing the subsequent cultivation and sampling efforts and generated a high quality set of 210 unique samples with chemical analysis results, spectral scanning data, and cultivation metadata. We observed a significant improvement in the performance of carbohydrate content predictions using an optimized ANN model compared to PLSR, with > 16% reduction in mean absolute percent error (MAPE) when tested on the same set of reserved data. The optimized ANN models for FAME and protein prediction performed exceptionally well with 5.99% and 5.09% MAPE, respectively. Application of these methods to detection and quantification of minor biomass constituents that are relevant to certain product streams has shown positive preliminary results, opening the possibility for extensions to the outputs of this powerful data type. All models are accompanied by prediction uncertainties and unsupervised spectral outlier detection to alert an operator to unreliable spectral data. These tools can be deployed for rapid determination of algal culture status, and cultivation and biomass quality improvement.

algal biofuels↗

Efficient estimation of the modified Gromov–Hausdorff distance between unweighted graphs

Abstract Gromov–Hausdorff distances measure shape difference between the objects representable as compact metric spaces, e.g. point clouds, manifolds, or graphs. Computing any Gromov–Hausdorff distance is equivalent to solving an NP-hard optimization problem, deeming the notion impractical for applications. In this paper we propose a polynomial algorithm for estimating the so-called modified Gromov–Hausdorff (mGH) distance, a relaxation of the standard Gromov–Hausdorff (GH) distance with similar topological properties. We implement the algorithm for the case of compact metric spaces induced by unweighted graphs as part of Python library , and demonstrate its performance on real-world and synthetic networks. The algorithm finds the mGH distances exactly on most graphs with the scale-free property. We use the computed mGH distances to successfully detect outliers in real-world social and computer networks.

Oles, Vladyslav (ORCID:0000000188727463)↗

Queue wait time prediction in high performance computing (HPC) systems

High Performance Computing (HPC) systems are critical enablers for groundbreaking scientific research across various domains. Efficient resource allocation, facilitated by job scheduling, is paramount for maximizing the utilization of HPC systems. However, the variability in wait times for queued jobs poses challenges for users, necessitating accurate job wait time estimation. This paper explores the influence of job characteristics, including job size (the number of nodes requested and walltime), the queue to which the job is submitted and other resource requirements, on job wait times in leadership-class HPC systems. Focusing on the Theta Cray XC40 and Polaris machines at Argonne National Laboratory, the study evaluates the performance of different supervised learning algorithms in predicting job wait times. It also evaluates the impact of data preprocessing, including outlier detection, Principal Component Analysis (PCA), and feature selection, on the performance of wait time prediction models. The findings reveal insights into the relationship between job characteristics and wait times, offering a foundation for optimizing resource allocation and enhancing user experience. The methodologies and tools developed in this study are adaptable to other leadership-class HPC systems, providing a valuable contribution to the broader HPC community aiming to improve job scheduling efficiency and user satisfaction.

Okafor, Nwamaka↗

Mapping wall-to-wall fractional cover of Arctic tundra plant functional types in Alaska using 20-m spatial resolution satellite imagery and harmonized plot observations

Estimates of fractional cover (fCover) across given land surfaces are used to assess, and often model, vegetation composition and diversity, which are crucial for understanding the health and functioning of terrestrial ecosystems. Remote sensing provides a useful means for scaling local, plot-measured fCover estimates to regional scales. Leveraging a recently synthesized and harmonized plot database, this study generated wall-to-wall maps of fCover for six Alaskan-Arctic plant functional types (PFT), including non-vascular plants, forbs, graminoids, and deciduous and evergreen shrubs, using 20-m satellite data (Sentinel-1, Sentinel-2, ArcticDEM) using a machine learning regression approach, specifically the random forest (RF) algorithm, which is well-suited for handling nonlinear relationships and high-dimensional satellite datasets. This study additionally addressed the spatio-temporal inconsistencies e.g., sampling scale, plot size, and collection year in plot measured fCover by adopting a multivariate outlier detection approach—Cook’s distance—to identify high-quality plots for model training and validation. Our approach achieves high accuracy (R 2 = 0.59–0.93, root mean squared errors = 0.02–0.10 for all PFTs) between plot-observed and satellite-derived fCover when using high-quality plot samples. The mapped fCover characterizes the spatial patterns of different PFTs across the tundra biome at a 20-m resolution, providing key information needed for improved representation of Arctic tundra vegetation in terrestrial biosphere models to better understand climate-vegetation feedback across the Arctic tundra.

Arctic tundra↗

Maximizing efficiency of dataset compression for machine learning potentials with information theory

Machine learning interatomic potentials (MLIPs) balance high accuracy and lower costs compared to density functional theory calculations, but their performance often depends on the size and diversity of training datasets. Large datasets improve model accuracy and generalization but are computationally expensive to produce and train on, while smaller datasets risk discarding rare but important atomic environments and compromising MLIP accuracy/reliability. Here, we develop an information-theoretical framework to quantify the efficiency of dataset compression methods and propose an algorithm that maximizes this efficiency. By framing atomistic dataset compression as an instance of the minimum set cover (MSC) problem over atom-centered environments, our method identifies the smallest subset of structures that contains as much information as possible from the original dataset while pruning redundant information. The approach is extensively demonstrated on the GAP-20 and TM23 datasets and validated on 64 varied datasets from the ColabFit repository. Across all cases, MSC consistently retains outliers, preserves dataset diversity, and reproduces the long-tail distributions of forces even at high compression rates, outperforming other subsampling methods. Furthermore, MLIPs trained on MSC-compressed datasets exhibit reduced error for out-of-distribution data even in low-data regimes. We explain these results using an outlier analysis and show that such quantitative conclusions could not be achieved with conventional dimensionality reduction methods. The algorithm is implemented in the open-source QUESTS package and can be used for several tasks in atomistic modeling, from data subsampling, outlier detection, and training improved MLIPs at a lower cost.

36 MATERIALS SCIENCE↗

Unsupervised atomic data mining via multi-kernel graph autoencoders for machine learning force fields

Constructing a chemically diverse dataset while avoiding sampling bias is critical to training efficient and generalizable force fields. However, in computational chemistry and materials science, many common dataset generation techniques are prone to oversampling regions of the potential energy surface. Furthermore, these regions can be difficult to identify and isolate from each other or may not align well with human intuition, making it challenging to systematically remove bias in the dataset. While traditional clustering and pruning (down-sampling) approaches can be useful for this, they can often lead to information loss or a failure to properly identify distinct regions of the potential energy surface due to difficulties associated with the high dimensionality of atomic descriptors. In this work, we introduce the Multi-kernel Edge Attention-based Graph Autoencoder (MEAGraph) model, an unsupervised approach for analyzing atomic datasets. MEAGraph combines multiple linear kernel transformations with attention-based message passing to capture geometric sensitivity and enable effective dataset pruning without relying on labels or extensive training. Demonstrated applications on niobium, tantalum, and iron datasets show that MEAGraph efficiently groups similar atomic environments, allowing for the use of basic pruning techniques for removing sampling bias. This approach provides an effective method for representation learning and clustering that can be used for data analysis, outlier detection, and dataset optimization.

Materials science↗

Data-centric framework for crystal structure identification in atomistic simulations using machine learning

Atomic-level modeling performed at large scales enables the investigation of mesoscale materials properties with atom-by-atom resolution. The spatial complexity of such cross-scale simulations renders them unsuitable for simple human visual inspection. Instead, specialized structure characterization techniques are required to aid interpretation. These have historically been challenging to construct, requiring significant intuition and effort. Here we propose an alternative framework for a fundamental structural characterization task: classifying atoms according to the crystal structure to which they belong. Our approach is data-centric and favors the employment of Machine Learning over heuristic rules of classification. A group of data-science tools and simple local descriptors of atomic structure are employed together with an efficient synthetic training set. We also introduce the first standard and publicly available benchmark data set for evaluation of algorithms for crystal-structure classification. Further, it is demonstrated that our data-centric framework outperforms all of the most popular heuristic methods—especially at high temperatures when lattices are the most distorted—while introducing a systematic route for generalization to new crystal structures. Moreover, through the use of outlier detection algorithms our approach is capable of discerning between amorphous atomic motifs (i.e., noncrystalline phases) and unknown crystal structures, making it uniquely suited for exploratory materials synthesis simulations.

36 MATERIALS SCIENCE↗