Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “outlier detection”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Cognitive IoT and Edge Computing for Intrusion Detection with Federated TinyML

Internet of Things (IoT) and Edge Computing (EC) are rapidly becoming an integral part of the modern society. By 2030, there is estimated to be over 40 billion active and connected IoT devices [1]. This rapid progress also comes with a significant implication on cybersecurity. Back-end infrastructure and systems have a much broader attack than they did previously due to vulnerable IoT/EC devices being connected to wireless networks. This expanding attack surface is a growing concern because IoT/EC are increasingly being used in critical systems such as power grids, health care, and smart homes. To effectively address a problem of this scale, cognitive cyber methods—which can autonomously detect and react to cyber attacks as they develop—are needed. To address this, we bring Artificial Intelligence (AI) and Machine Learning (ML) to IoT/EC devices, using tinyML to monitor voluminous IoT data against cyber threats, and using Federated Learning (FL) to share local detection knowledge across the system while preserving privacy. We propose a novel three-layer architecture: (1) an IoT layer for tinyML-based inference, (2) an edge layer for ML model training, and (3) a cloud layer for FL operations. Using the publicly available 11-class N-BaIoT dataset [2], we demonstrate that this architecture mitigates resource constraints at the IoT layer while improving detection accuracy over standard two-layer designs. An outlier-resistant scaler, feature reduction, and quantization enable the tinyML model to maintain detection accuracy with a reduced model size. Additionally, federated learning that only utilizes the intersection (across heterogenous devices) of the reduced feature set achieves superior detection accuracy compared to locally trained models.

Li, Mingyan [ORNL] (ORCID:0009000569532640)↗

Out of Distribution Detection with Neural Network Anchoring

This is code to reproduce and build on OOD detection from the paper "Out of Distribution Detection with Neural Network Anchoring". Our goal here is to exploit heteroscedastic temperature scaling as a calibration strategy for out of distribution (OOD) detection. Heteroscedasticity here refers to the fact that the optimal temperature parameter for each sample can be different, as opposed to conventional approaches that use the same value for the entire distribution. To enable this, we propose a new training strategy called anchoring that can estimate appropriate temperature values for each sample, leading to state-of-the-art OOD detection performance across several benchmarks. Using NTK theory, we show that this temperature function estimate is closely linked to the epistemic uncertainty of the classifier, which explains its behavior. In contrast to some of the best-performing OOD detection approaches, our method does not require exposure to additional outlier datasets, custom calibration objectives, or model ensembling. Through empirical studies with different OOD detection settings - far OOD, near OOD, and semantically coherent OOD - we establish a highly effective OOD detection approach.

Thiagarajan, Jayaraman↗

Time series anomaly detection in power electronics signals with recurrent and ConvLSTM autoencoders

The anomalies in the high voltage converter modulator (HVCM) remain a major down time for the spallation neutron source facility, that delivers the most intense neutron beam in the world for scientific materials research. In this work, we propose neural network architectures based on Recurrent AutoEncoders (RAE) to detect anomalies ahead of time in the power signals coming from the HVCM. Bi-directional gated recurrent unit, bi-directional long-short term memory (LSTM), and convolutional LSTM (ConvLSTM) are developed, trained, and tested using real experimental signals from the HVCM module. The results show a good performance of the proposed RAE models, achieving precision up to 91%, recall up to 88%, false omission rate as low as 20% (i.e. 80% of the anomalies were detected), and area under the ROC curve up to 0.9. The three RAE models provide very comparable performance, with LSTM showing slightly better performance than GRU and ConvLSTM. The RAE models are benchmarked against other anomaly detection methods, including isolation forest, support vector machine, local outlier factor, feedforward and convolutional autoencoders, and others; showing a better performance. Here, the results of this study demonstrate the promising potential of RAE in anomaly detection for real-world power systems, and for increasing the reliability of the HVCM modules in the spallation neutron source.

42 ENGINEERING↗

Unique & challenging aspects of plutonium metal standards exchange program for actinide measurements

The Los Alamos National Laboratory exchange program is the only program of its kind for the distribution of plutonium (Pu) standards materials with a range of impurity contents to multiple laboratories for destructive measurements of elemental concentration. This paper discusses statistical methods used to address challenges in Pu metal exchange data by way of two case studies. Challenges include how to evaluate a data set when a large fraction of the values are minimum detection limits (MDLs), and how to determine potential outliers with limited in-formation on the true spread of the data.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Apparatuses and methods for anomalous gas concentration detection

Embodiments of the disclosure are drawn to apparatuses and methods for anomalous gas concentration detection. A spectroscopic system, such as a wavelength modulated spectroscopy (WMS) system may measure gas concentrations in a target area. However, noise, such as speckle noise, may interfere with measuring relatively low concentrations of gas, and may lead to false positives. A noise model, which includes a contribution from a speckle noise model, may be used to process data from the spectroscopic system. An adaptive threshold may be applied based on an expected amount of noise. A speckle filter may remove measurements which are outliers based on a measurement of their noise. Plume detection may be used to determine a presence of gas plumes. Each of these processing steps may be associated with a confidence, which may be used to determine an overall confidence in the processed measurements/gas plumes.

Kreitinger, Aaron Thomas↗

Event-Based Anomaly Detection for Searches for New Physics

This paper discusses model-agnostic searches for new physics at the Large Hadron Collider using anomaly-detection techniques for the identification of event signatures that deviate from the Standard Model (SM). We investigate anomaly detection in the context of a machine-learning approach based on autoencoders. The analysis uses Monte Carlo simulations for the SM background and several selected exotic models. We also investigate the input space for the event-based anomaly detection and illustrate the shapes of invariant masses in the outlier region which will be used to perform searches for resonant phenomena beyond the SM. Challenges and conceptual limitations of this approach are discussed.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Autonomous anomaly detection of proliferation in the AGN-201 nuclear reactor digital twin

The expansion of global nuclear power necessitates advanced methods for analyzing proliferation indicators. This study introduces a novel application of the Isolation Forest Machine Learning (IFML) algorithm within a digital twin (DT) of the AGN-201 nuclear reactor to autonomously detect anomalies. Leveraging real-time operational data from the AGN-201 DT, the IFML algorithm identifies outliers without prior data labeling and operates as a lightweight, complementary approach to traditional physics-based anomaly detection methods for nuclear safeguards. In a simulated Red vs. Blue team exercise, the IFML algorithm successfully detected six significant unseen anomalies related to reactivity changes, achieving an accuracy of 99% for identifying operational deviationxs. These anomalies, caused by deliberate perturbations, were detected alongside known physics-based models, underscoring the potential of IFML to enhance real-time monitoring without displacing traditional methods. Further, this study highlights the applicability of IFML in nuclear environments by providing an additional, redundant layer of anomaly detection to improve safeguards and operational safety in complex systems.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Wasserstein normalized autoencoder for anomaly detection

A novel anomaly detection algorithm is presented. The Wasserstein normalized autoencoder (WNAE) is a normalized probabilistic model that minimizes the Wasserstein distance between the learned probability distribution—a Boltzmann distribution where the energy is the reconstruction error of the autoencoder (AE)—and the distribution of the training data. This algorithm has been developed and applied to the identification of semivisible jets—conical sprays of visible standard model (SM) particles and invisible dark matter states—with the CMS experiment at the CERN LHC. Trained on jets of particles from simulated SM processes, the WNAE is shown to learn the probability distribution of the input data in a fully unsupervised fashion, such that it effectively identifies new physics jets as anomalies. The model exhibits stable, convergent training and recovers strong classification performance for a wide range of signals against the selected background process, for which a standard AE fails because of outlier reconstruction. In addition, the model improves upon standard normalized autoencoders while remaining fully agnostic to the signal. The WNAE directly tackles the problem of outlier reconstruction, a common failure mode of autoencoders in anomaly detection tasks.

Hayrapetyan, Aram [Yerevan Phys. Inst.]↗

Unsupervised anomaly clustering via offset alignment in multivariate grid sensing data

Modern industries increasingly rely on multi-sensor technologies to acquire complex, high-dimensional data streams, enabling advanced monitoring and control systems. One critical application is online anomaly detection in electrical smart grids, where multivariate and multimodal sensing technologies play a vital role. However, detecting anomalies in such time-series data is challenging due to their inherent temporal dependencies and stochastic behavior. Traditional approaches based on supervised and semi-supervised learning methods depend on labeled datasets, which are often unavailable in real-world scenarios. While unsupervised methods have emerged as promising alternatives, these methods are highly susceptible to noise and outliers commonly present in sensing applications. Furthermore, deep learning-based anomaly detection methods, despite their performance, are often criticized for their black-box nature, limiting their applicability in safety-critical and online environments where interpretability and explainability are paramount. In this work, we propose an unsupervised anomaly clustering method leveraging a cyclic alignment-based offset detection algorithm for multivariate time-series signals. The proposed method is applied to multivariate data collected from vibrational, voltage, and magnetic field sensors deployed in a local grid substation. Our results demonstrate the robustness of the algorithm in accurately clustering various anomalies/events across different sensing modalities. Additionally, we compare the effectiveness of the proposed approach against a simple pattern-based anomaly detection method, which performs well for univariate data but fails to generalize to multivariate and multimodal time-series data.

Mukherjee, Subrata [ORNL] (ORCID:0000000309930338)↗

Field intercomparison of ice nucleation measurements: the Fifth International Workshop on Ice Nucleation Phase 3 (FIN-03)

Abstract. The third phase of the Fifth International Ice Nucleation Workshop (FIN-03) was conducted at the Storm Peak Laboratory in Steamboat Springs, Colorado, in September 2015 to facilitate the intercomparison of instruments measuring ice-nucleating particles (INPs) in the field. Instruments included two online and four offline measurement systems for INPs, which are a subset of those utilized in the laboratory study that comprised the second phase of FIN (FIN-02). The composition of the total aerosols was characterized using the Particle Analysis by Laser Mass Spectrometry (PALMS) and Wideband Integrated Bioaerosol Sensor (WIBS) instruments, and aerosol size distributions were measured by a laser aerosol spectrometer (LAS). The dominant total particle compositions present during FIN-03 were composed of sulfates, organic compounds, and nitrates, as well as particles derived from biomass burning. Mineral-dust-containing particles were ubiquitous throughout and represented 67 % of supermicron particles. Total WIBS fluorescing particle concentrations for particles with diameters of > 0.5 µm were 0.04 ± 0.02 cm−3 (0.1 cm−3 highest; 0.02 cm−3 lowest), typical of the warm season in this region and representing ≈ 9 % of all particles in this size range as a campaign average. The primary focus of FIN-03 was the measurement of INP concentrations via immersion freezing at temperatures > −33 °C. Additionally, some measurements were made in the deposition nucleation regime at these same temperatures, representing one of the first efforts to include both mechanisms within a field campaign. INP concentrations via immersion freezing agreed within factors ranging from nearly 1 to 5 times on average between matched (time and temperature) measurements, and disagreements only rarely exceeded 1 order of magnitude for sampling times coordinated to within 3 h. Comparisons were restricted to temperatures lower than −15 °C due to the limits of detection related to sample volumes and very low INP concentrations. Outliers of up to 2 orders of magnitude occurred between −25 and −18 °C; a better agreement was seen at higher and lower temperatures. Although the 5–10 factor agreement of INP measurements found in FIN-03 aligned with the results of the FIN-02 laboratory comparison phase, giving confidence in progress of this measurement field, this level of agreement still equates to temperature uncertainties of 3.5 to 5 °C that may not be sufficient for numerical cloud modeling applications that utilize INP information. INP activity in the immersion-freezing mode was generally found to be an order of magnitude or more, making it more efficient than in the deposition regime at 95 %–99 % water relative humidity, although this limited data set should be augmented in future efforts. To contextualize the study results, an assessment was made of the composition of INPs during the late-summer to early-fall period of this study inferred through comparison to existing ice nucleation parameterizations and through measurement of the influence of thermal and organic carbon digestion treatments on immersion-freezing ice nucleation activity. Consistent with other studies in continental regions, biological INPs dominated at temperatures of > −20 °C and sometimes colder, while arable dust-like or other organic-influenced INPs were inferred to dominate below −20 °C.

54 ENVIRONMENTAL SCIENCES↗

Optical/X-ray/radio view of Abell 1213: A galaxy cluster with anomalous diffuse radio emission

Context. Abell 1213, a low-richness galaxy system, is known to host an anomalous radio halo detected in data of the Very Large Array (VLA). It is an outlier with regard to the relation between the radio halo power and the X-ray luminosity of the parent clusters. Aims. Our aim is to analyze the cluster in the optical, X-ray, and radio bands to characterize the environment of its diffuse radio emission and to shed new light on its nature. Methods. We used optical data from the Sloan Digital Sky Survey to study the internal dynamics of the cluster. We also analyzed archival XMM-Newton X-ray data to unveil the properties of its hot intracluster medium. Finally, we used recent data from the LOw Frequency ARray (LOFAR) at 144 MHz, together with VLA data at 1.4 GHz, to study the spectral behavior of the diffuse radio source. Results. Both our optical and X-ray analysis reveal that this low-mass cluster exhibits disturbed dynamics. In fact, it is composed of several galaxy groups in the peripheral regions and, in particular, in the core, where we find evidence of substructures oriented in the NE–SW direction, with hints of a merger nearly along the line of sight. The analysis of the X-ray emission adds further evidence that the cluster is in an unrelaxed dynamical state. At radio wavelengths, the LOFAR data show that the diffuse emission is ~510 kpc in size. Moreover, there are hints of low-surface-brightness emission permeating the cluster center. Conclusions. The environment of the diffuse radio emission is not what we would expect for a classical halo. The spectral index map of the radio source is compatible with a relic interpretation, possibly due to a merger in the N–S or NE–SW directions, in agreement with the substructures detected through the optical analysis. The fragmented, diffuse radio emissions at the cluster center could be attributed to the surface brightness peaks of a faint central radio halo.

79 ASTRONOMY AND ASTROPHYSICS↗

A Novel Machine Learning Algorithm for Cloud Detection Using AERI Measurement Data

Infrared hyperspectral remote sensing has been widely used in the field of meteorology. Many scientists have carried out research on inversion methods of meteorological elements such as thermodynamic profile, boundary layer height, cloud base height, etc. In this study, a method based on machine learning for cloud detection using ground-based infrared hyperspectral radiation data is proposed. The features of outliers, the cloudy and cloud-free data of Atmospheric Emitted Radiance Interferometer (AERI) radiation are extracted. The “reference values” of cloudy and cloud-free are determined based on the observation data of Vaisala CL31 ceilometer within the time range of 8 min before the corresponding time of AERI. A support vector machine (SVM) algorithm is used for training. The dataset comes from the Atmospheric Radiation Measurement (ARM) Southern Great Plains (SGP) site and North Slope Alaska (NSA) site from 2015 to 2017, and the ARM West Antarctic Radiation Experiment (AWARE) site in 2016 is also analyzed. The instruments used in this paper include AERI, ceilometer, etc. The experimental results reveal that the agreement of cloud detection results between the proposed algorithm and ceilometer is about 93% at each site. However, for high clouds or optically thin clouds, the agreement will decrease.

47 OTHER INSTRUMENTATION↗

Hubble Frontier Field Clusters and Their Parallel Fields: Photometric and Photometric Redshift Catalogs

We present a multiband analysis of the six Hubble Frontier Field clusters and their parallel fields, producing catalogs with measurements of source photometry and photometric redshifts. We release these catalogs to the public along with maps of intracluster light and models for the brightest galaxies in each field. This rich data set covers a wavelength range from 0.2 to 8 μm, utilizing data from the Hubble Space Telescope, Keck Observatories, Very Large Telescope array, and Spitzer Space Telescope. We validate our products by injecting into our fields and recovering a population of synthetic objects with similar characteristics to those in real extragalactic surveys. The photometric catalogs contain a total of over 32,000 entries, with 50% completeness at a threshold of mag AB ~ 29.1 for unblended sources and magAB ~ 29 for blended ones, in the IR-weighted detection band. Photometric redshifts were obtained by means of template fitting and have an average outlier fraction of 10.3% and scatter σ = 0.067 when compared to spectroscopic estimates. The software we devised, after being tested in the present work, will be applied to new data sets from ongoing and future surveys.

79 ASTRONOMY AND ASTROPHYSICS↗

CIRCLEZ : Reliable photometric redshifts for active galactic nuclei computed solely using photometry from Legacy Survey Imaging for DESI

Photometric redshifts for galaxies hosting an accreting supermassive black hole in their center, known as active galactic nuclei (AGNs), are notoriously challenging. At present, they are most optimally computed via spectral energy distribution (SED) fittings, assuming that deep photometry for many wavelengths is available. However, for AGNs detected from all-sky surveys, the photometry is limited and provided by a range of instruments and studies. This makes the task of homogenizing the data challenging, presenting a dramatic drawback for the millions of AGNs that wide surveys such as SRG/eROSITA are poised to detect. This work aims to compute reliable photometric redshifts for X-ray-detected AGNs using only one dataset that covers a large area: the tenth data release of the Imaging Legacy Survey (LS10) for DESI. LS10 provides deep grizW1-W4 forced photometry within various apertures over the footprint of the eROSITA-DE survey, which avoids issues related to the cross-calibration of surveys. We present the results from CIRCLEZ, a machine-learning algorithm based on a fully connected neural network. CIRCLEZ is built on a training sample of 14 000 X-ray-detected AGNs and utilizes multi-aperture photometry, mapping the light distribution of the sources. The accuracy (σNMAD) and the fraction of outliers (η) reached in a test sample of 2913 AGNs are equal to 0.067 and 11.6%, respectively. The results are comparable to (or even better than) what was previously obtained for the same field, but with much less effort in this instance. We further tested the stability of the results by computing the photometric redshifts for the sources detected in CSC2 and Chandra-COSMOS Legacy, reaching a comparable accuracy as in eFEDS when limiting the magnitude of the counterparts to the depth of LS10. The method can be applied to fainter samples of AGNs using deeper optical data from future surveys (for example, LSST, Euclid), granting LS10-like information on the light distribution beyond the morphological type. Along with this paper, we have released an updated version of the photometric redshifts (including errors and probability distribution functions) for eROSITA/eFEDS.

79 ASTRONOMY AND ASTROPHYSICS↗

Photometry on Structured Backgrounds: Local Pixel-wise Infilling by Regression

Photometric pipelines struggle to estimate both the flux and flux uncertainty for stars in the presence of structured backgrounds such as filaments or clouds. However, it is exactly stars in these complex regions that are critical to understanding star formation and the structure of the interstellar medium. We develop a method, similar to Gaussian process regression, which we term local pixel-wise infilling (LPI). Using a local covariance estimate, we predict the background behind each star and the uncertainty of that prediction in order to improve estimates of flux and flux uncertainty. We show the validity of our model on synthetic data and real dust fields. We further demonstrate that the method is stable even in the crowded field limit. While we focus on optical-IR photometry, this method is not restricted to those wavelengths. We apply this technique to the 34 billion detections in the second data release of the Dark Energy Camera Plane Survey. In addition to removing many >3σ outliers and improving uncertainty estimates by a factor of ~2–3 on nebulous fields, we also show that our method is well behaved on uncrowded fields. The entirely post-processing nature of our implementation of LPI photometry allows it to easily improve the flux and flux uncertainty estimates of past as well as future surveys.

79 ASTRONOMY AND ASTROPHYSICS↗

Cy-Phy ADS: Cyber Physical Anomaly Detection Framework for EV Charging Systems

Today’s large-scale Electric Vehicle (EV) infrastructures are heavily dependent on information communication technologies to maintain their operation and to support communication within sub-system components as well as the outside world. These technologies are vulnerable to various cyber and physical threats. Timely identification and mitigation of these threats are critical for improving human safety, avoiding economic losses, and preventing catastrophic system failures. By addressing this, our work presents a ResNet Autoencoder (AE) based Cyber-Physical Anomaly Detection System (Cy-Phy ADS) for detecting anomalies in EV Controller Area Network (CAN) protocol communication. It consists of four main components: Cyber-Physical Feature Extractor, ResNet AE-based Anomaly Detection Framework, Cyber-Physical Health Metric (CPHM), and Visualization Dashboard. The presented framework was trained and tested using CAN data collected from the EV charging system testbed at the Idaho National Laboratory. The presented Cy-Phy ADS compared against six widely used unsupervised anomaly detection algorithms: One Class Support Vector Machine (OCSVM), Variational Autoencoder (VAE), LSTM Autoencoder (LSTM AE), Isolation Forest (IForest), Principle Component Analysis (PCA) and Local Outlier Factor (LOF). Here the presented approach showed the highest accuracy among the compared methods. Further, the proposed approach showed comparable performance in terms of precision, F1, and False positive rate. It also showed the lowest training and inference time compared to the neural network-based baseline algorithms compared against with. Additionally, the Cy-Phy ADS has advantages such as unsupervised training, the ability to provide a holistic metric for system health characterization, and non-linear feature extraction.

99 GENERAL AND MISCELLANEOUS↗

Fast and efficient identification of anomalous galaxy spectra with neural density estimation

ABSTRACT Current large-scale astrophysical experiments produce unprecedented amounts of rich and diverse data. This creates a growing need for fast and flexible automated data inspection methods. Deep learning algorithms can capture and pick up subtle variations in rich data sets and are fast to apply once trained. Here, we study the applicability of an unsupervised and probabilistic deep learning framework, the probabilistic auto-encoder, to the detection of peculiar objects in galaxy spectra from the SDSS survey. Different to supervised algorithms, this algorithm is not trained to detect a specific feature or type of anomaly, instead it learns the complex and diverse distribution of galaxy spectra from training data and identifies outliers with respect to the learned distribution. We find that the algorithm assigns consistently lower probabilities (higher anomaly score) to spectra that exhibit unusual features. For example, the majority of outliers among quiescent galaxies are E+A galaxies, whose spectra combine features from old and young stellar population. Other identified outliers include LINERs, supernovae, and overlapping objects. Conditional modelling further allows us to incorporate additional information. Namely, we evaluate the probability of an object being anomalous given a certain spectral class, but other information such as metrics of data quality or estimated redshift could be incorporated as well. We make our code publicly available.

Böhm, Vanessa↗

Coincident learning for unsupervised anomaly detection of scientific instruments

Abstract Anomaly detection is an important task for complex scientific experiments and other complex systems (e.g. industrial facilities, manufacturing), where failures in a sub-system can lead to lost data, poor performance, or even damage to components. While scientific facilities generate a wealth of data, labeled anomalies may be rare (or even nonexistent), and expensive to acquire. Unsupervised approaches are therefore common and typically search for anomalies either by distance or density of examples in the input feature space (or some associated low-dimensional representation). This paper presents a novel approach called coincident learning for anomaly detection (CoAD), which is specifically designed for multi-modal tasks and identifies anomalies based on coincident behavior across two different slices of the feature space. We define an unsupervised metric, F ^ β , out of analogy to the supervised classification F β statistic. CoAD uses F ^ β to train an anomaly detection algorithm on unlabeled data , based on the expectation that anomalous behavior in one feature slice is coincident with anomalous behavior in the other. The method is illustrated using a synthetic outlier data set and a MNIST-based image data set, and is compared to prior state-of-the-art on two real-world tasks: a metal milling data set and our motivating task of identifying RF station anomalies in a particle accelerator.

43 PARTICLE ACCELERATORS↗