Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “anomalies”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Coincident learning for unsupervised anomaly detection of scientific instruments

Abstract Anomaly detection is an important task for complex scientific experiments and other complex systems (e.g. industrial facilities, manufacturing), where failures in a sub-system can lead to lost data, poor performance, or even damage to components. While scientific facilities generate a wealth of data, labeled anomalies may be rare (or even nonexistent), and expensive to acquire. Unsupervised approaches are therefore common and typically search for anomalies either by distance or density of examples in the input feature space (or some associated low-dimensional representation). This paper presents a novel approach called coincident learning for anomaly detection (CoAD), which is specifically designed for multi-modal tasks and identifies anomalies based on coincident behavior across two different slices of the feature space. We define an unsupervised metric, F ^ β , out of analogy to the supervised classification F β statistic. CoAD uses F ^ β to train an anomaly detection algorithm on unlabeled data , based on the expectation that anomalous behavior in one feature slice is coincident with anomalous behavior in the other. The method is illustrated using a synthetic outlier data set and a MNIST-based image data set, and is compared to prior state-of-the-art on two real-world tasks: a metal milling data set and our motivating task of identifying RF station anomalies in a particle accelerator.

43 PARTICLE ACCELERATORS↗

Robust Spectral Anomaly Detection in EELS Spectral Images via 3D Convolutional Variational Autoencoders

Abstract A 3D Convolutional Variational Autoencoder (3D‐CVAE) is introduced for automated anomaly detection in electron energy‐loss spectroscopy spectrum imaging (EELS‐SI) data. This approach leverages the full 3D structure of EELS‐SI data to detect subtle spectral anomalies while preserving both spatial and spectral correlations across the datacube. By employing cross‐entropy loss and training on bulk spectra, the model learns to reconstruct bulk features characteristic of the defect‐free material. In exploring methods for anomaly detection, both the 3D‐CVAE approach and principal component analysis (PCA) are evaluated, testing their performance using FeL‐edge ΔEpeak shifts designed to simulate material defects. These results show that 3D‐CVAE achieves superior anomaly detection and maintains consistent performance across various shift magnitudes. The method demonstrates clear bimodal separation between bulk and anomalous spectra, enabling reliable classification. Further analysis verifies that lower‐dimensional representations are robust to anomalies in the data. While performance advantages over PCA diminish with decreasing anomaly concentration, our method maintains high reconstruction quality even in challenging, noise‐dominated spectral regions. This approach provides a robust framework for unsupervised automated detection of spectral anomalies in EELS‐SI data, particularly valuable for analyzing complex material systems.

Chemistry↗

Anomalies of 4d SpinG theories

Abstract We consider ’t Hooft anomalies of four-dimensional gauge theories whose fermion matter content admits Spin G (4) generalized spin structure, withGeither gauged or a global symmetry. We discuss methods to directly computew 2 ∪w 3 ’t Hooft anomalies involving Stiefel-Whitney classes of gauge and flavor symmetry bundles that such theories can have on non-spin manifolds, e.g.M 4 = ℂℙ 2 . Such anomalies have been discussed for SU(2) gauge theory with adjoint fermions, where they were shown to give an effect that was originally found in the Donaldson-Witten topological twist of$$ \mathcal{N} $$ N = 2 SYM theory. We directly compute these anomalies for a variety of theories, including generalGgauge theories with adjoint fermions, SU(2) gauge theory with fermions in general representations, and Spin(N) gauge theories with fundamental matter. We discuss aspects of matching these and other ’t Hooft anomalies in the IR phase where global symmetries are spontaneously broken, in particular for generalG gauge theory withN f adjoint Weyl fermions. For example, in the case ofN f = 2 we discuss anomaly matching in the IR phase consisting of$$ {h}_{G_{\textrm{gauge}}}^{\vee } $$ h G gauge ∨ copies of a ℂℙ 1 non-linear sigma model, including for thew 2 w 3 anomalies when formulated with$$ {\textrm{Spin}}_{\textrm{SU}{(2)}_{\textrm{global}}}(4) $$ Spin SU 2 global 4 structure.

Physics↗

Near-global summer circulation response to the spring surface temperature anomaly in Tibetan Plateau –– the GEWEX/LS4P first phase experiment

Subseasonal to seasonal (S2S) prediction of droughts and floods is one of the major challenges of weather and climate prediction. Recent studies suggest that the springtime land surface temperature/subsurface temperature (LST/SUBT) over the Tibetan Plateau (TP) can be a new source of S2S predictability. The project “Impact of Initialized Land Surface Temperature and Snowpack on Subseasonal to Seasonal Prediction (LS4P)” was initiated to study the impact of springtime LST/SUBT anomalies over high mountain areas on summertime precipitation predictions. The present work explores the simulated global scale response of the atmospheric circulation to the springtime TP land surface cooling by 16 current state-of-the-art Earth System Models (ESMs) participating in the LS4P Phase I (LS4P-I) experiment. The LS4P-I results show, for the first time, that springtime TP surface anomalies can modulate a persistent quasi-barotropic Tibetan Plateau-Rocky Mountain Circumglobal (TRC) wave train from the TP via the northeast Asia and Bering Strait to the western part of the North America, along with the springtime westerly jet from TP across the whole North Pacific basin. The TRC wave train modulated by the TP thermal anomaly play a critical role on the early summer surface air temperature and precipitation anomalies in the regions along the wave train, especially over the northwest North America and the southern Great Plains. The participant models that fail in capturing the TRC wave train greatly under-predict climate anomalies in reference to observations and the successful models. These results suggest that the TP LST/SUBT anomaly via the TRC wave train is the first order source of the S2S variability in the regions mentioned. Furthermore, the TP surface temperature anomaly can influence the Southern Hemispheric circulation by generating cross-equator wave trains. However, the simulated propagation pathways from the TP into the Southern Hemisphere show large inter-model differences. More dynamical understanding of the TRC wave train as well as its cross-equator propagation into the Southern Hemisphere will be explored in the newly launched LS4P phase II experiment.

54 ENVIRONMENTAL SCIENCES↗

Multiscale Temporal Variability of the Global Air‐Sea CO 2 Flux Anomaly

Abstract The global air‐sea CO 2 flux (F) impacts and is impacted by a plethora of climate‐related processes operating at multiple time scales. In bulk mass transfer formulations, F is driven by physico‐ and bio‐chemical factors such as the air‐sea partial pressure difference (∆pCO 2 ), gas transfer velocity, sea surface temperature, and salinity–all varying at multiple time scales. To de‐convolve the impact of these factors on variability in F at different time scales, time‐resolved estimates of F were computed using a global data set assembled between 1988 and 2015. The F anomalies were defined as temporal deviations from the 28‐year time‐averaged value. Spectral analysis revealed four dominant timescales of variability in F–subseasonal, seasonal, interannual, and decadal with relative amplitude differences varying across regions. A second‐order Taylor series expansion was then conducted along these four timescales to separate drivers across differing regions. The analysis showed that on subseasonal timescales, wind speed variability explains some 66% of the global F anomaly and is the dominant driver. On seasonal, interannual, and decadal timescales, the ∆pCO 2 effect controlled by the ∆pCO 2 anomaly, explained much of the F anomaly. On decadal timescales, the F anomaly was almost entirely governed by the ∆pCO 2 effect with large contributions from high latitudes. The main drivers across timescales also dominate the regional F anomaly, particularly in the mid‐high latitude regions. Finally, the driver of the ∆pCO 2 effect was closely connected with the relative strength of atmospheric pCO 2 and the nonthermal component of oceanic pCO 2 anomaly associated with dissolved inorganic carbon and alkalinity.

Environmental Sciences & Ecology↗

Extreme Precipitation Over the Southern Slope of the Tibetan Plateau and the Associated Atmospheric Circulation Anomalies

The southern slope of the Tibetan Plateau (SSTP) is one of the rainiest regions in the world where geological hazards caused by extreme precipitation often occur. This study investigates the characteristics and mechanisms of extreme precipitation over SSTP from June to September during 2001–2020 using Global Precipitation Measurement satellite observation. The extreme precipitation days are defined as the days with top 5% of regional-mean daily precipitation over SSTP in this period, which has an average precipitation of 27.2 mm/d. Averaging over the extreme precipitation days, precipitation peaks at an altitude of about 300 m, coinciding with the climatological maximum precipitation, but with a much larger value of 37.2 mm/d than the climatology of 11.9 mm/d. Composite analysis of circulations on extreme days reveals significant circulation anomalies in both the lower and upper troposphere. Specifically, the lower-tropospheric circulations are characterized by significant westerly anomalies over northern India, and the upper-tropospheric circulations are characterized by northerly anomalies over the central Tibetan Plateau, which are statistically independent. The lower-tropospheric westerly anomalies blowing toward SSTP are blocked by the topography, favoring extreme precipitation over SSTP. The upper-tropospheric northerly anomalies, on the other hand, correspond to anomalous northeasterlies north of SSTP in the middle troposphere and southeasterlies to the south in the lower troposphere, and the convergence of these circulation anomalies favors extreme precipitation over SSTP. Lastly, the lower-tropospheric westerly and upper-tropospheric northerly anomalies, respectively, correspond to less precipitation over the South Asian monsoon region and the Tibetan Plateau.

Extreme precipitation, Tibetan Plateau↗

Effect of Rocky Mountains and Tibetan Plateau 1998 Spring Land Temperature on N. American and East Asian Summer Precipitation Anomalies

This work follows up on the GEWEX/LS4P Phase I (LS4P-I) experiments, a community effort highlighting the spring land surface temperature anomalies in the Tibetan Plateau (TP) as a useful source for subseasonal to seasonal (S2S) prediction of summer precipitation in global hot spot regions, particularly in East Asia and North America. This paper extends the investigation to both the US Rocky Mountain (RM) region and the TP, considering the 1998 summer drought/flood event in North America/East Asia, respectively, as a case study. A previously developed initialization method for land surface temperature/subsurface temperature (LST/SUBT) is used in the NCEP Global Forecast System, coupled with a land model, SSiB2 (GFS/SSiB2), to produce observed RM cold May temperature anomaly. Forward simulation yields June precipitation anomalies at five remote locations. Likewise, the TP warm May temperature anomaly also produces June precipitation anomalies at these five locations. The effects of RM (cold) and TP (warm) temperature anomalies are consistent in the US South Coastal regions and the south Yangtze River Basin, yielding 49% (42%) of observed drought and 34% (44%) of observed flood, respectively. These LST/SUBT effects in RM and TP induce a global large-scale wave train linking North America with the TP, affecting the subtropical westerly jet and thereby modulating summer precipitation. Global SST effect is examined for comparison but does not yield statistically significant June precipitation anomalies in GFS/SSiB2. Furthermore, this study adds to evidence that high-mountain LST effects in the RM and TP are first-order sources of S2S precipitation predictability in summer months.

Nayak, Hara Prasad [University of California, Los ↗

Signal Processing Based Method for Real-Time Anomaly Detection in High-Performance Computing

Performance anomalies can manifest as irregular execution times or abnormal execution events for many reasons, including network congestion and resource contention. Detecting such anomalies in real-time by analyzing the details of performance traces at scale is impractical due to the sheer volume of data High-Performance Computing (HPC) applications produce. In this paper, we propose formulating HPC performance anomaly detection as a signal-processing problem where anomalies can be treated as noise. We evaluate our proposed method in comparison with two other commonly used anomaly detection techniques of varying complexity based on their detection accuracy and scalability. Since real-time in-situ anomaly detection at a large scale requires lightweight methods that can handle a large volume of streaming data, we find that our proposed method provides the best trade-off. We then implement the proposed method in Chimbuko, the first online, distributed, and scalable workflow-level performance trace analysis framework. We compare our proposed signal-based anomaly detection algorithm with two other methods using a function of their accuracy, F1 score, and detection overhead. Our experiments demonstrate that our proposed approach achieves a 99% improvement for the benchmark datasets and a 93% improvement with Chimbuko traces.

99 GENERAL AND MISCELLANEOUS↗

Multivariate Time Series Anomaly Detection with Few Positive Samples

Given the scarcity of anomalies in real-world applications, the majority of literature has been focusing on modeling normality. The learned representations enable anomaly detection as the normality model is trained to capture certain key underlying data regularities under normal circumstances. In practical settings, particularly industrial time series anomaly detection, we often encounter situations where a large amount of normal operation data is available along with a small number of anomaly events collected over time. This practical situation calls for methodologies to leverage these small number of anomaly events to create a better anomaly detector. In this paper, we introduce two methodologies to address the needs of this practical situation and compared them with recently developed state of the art techniques. Our proposed methods anchor on representative learning of normal operation with autoregressive (AR) model along with loss components to encourage representations that separate normal versus few positive examples. We applied the proposed methods to two industrial anomaly detection datasets and demonstrated effective performance in comparison with approaches from literature. Our study also points out additional challenges with adopting such methods in practical applications.

Xue, Feng↗

MAD: Self-Supervised Masked Anomaly Detection Task for Multivariate Time Series

In this paper, we introduce Masked Anomaly Detection (MAD), a general self-supervised learning task for multivariate time series anomaly detection. With the increasing availability of sensor data from industrial systems, being able to detecting anomalies from streams of multivariate time series data is of significant importance. Given the scarcity of anomalies in real-world applications, the majority of literature has been focusing on modeling normality. The learned normal representations can empower anomaly detection as the model has learned to capture certain key underlying data regularities. A typical formulation is to learn a predictive model, i.e., use a window of time series data to predict future data values. In this paper, we propose an alternative self-supervised learning task. By randomly masking a portion of the inputs and training a model to estimate them using the remaining ones, MAD is an improvement over the traditional left-to-right next step prediction (NSP) task. Our experimental results demonstrate that MAD can achieve better anomaly detection rates over traditional NSP approaches when using exactly the same neural network (NN) base models, and can be modified to run as fast as NSP models during test time on the same hardware, thus making it an ideal upgrade for many existing NSP-based NN anomaly detection models.

97 MATHEMATICS AND COMPUTING↗

Hyper Spectral Anomaly Detection

The HSA is a statistics based anomaly detection model. The model performs unsupervised anomaly detection, based on a datapoint's density and similarity within a dataset. Density and similarity data are encoded into an affinity matrix. The affinity matrix is evolved to summarize the data's structure on greater topographical scales within the data's function space. The set of evolved affinity matrices and an anomaly score vector are passed to a user defined penalized objective function. The penalized objective function of anomaly scores is then minimized. Data points where the absolute value of the z-scores of anomaly scores greater than a specified threshold are predicted as anomalies. A novel multi-filter feature has also been implemented. To reduce false positive rates, the multi-filter records the indexes of the HSA predictions. A new dataset and data loader are instantiated consisting of all the initial HSA predictions and non-anomalous data points in a 10% and 90% split respectively. The HSA is then run through this data set and a count of number of times a data point is predicted is kept. In this way the initial predictions may be compared with data spanning the entire dataset. After the multi-filter is complete, all datapoints will have an associated anomaly score, as well as a multi-filter prediction count to further filter the anomalous predictions.

Rogers, DempseyD [Idaho National Laboratory (INL),↗

Efficient Anomaly Detection Driven By Different Machine Learning Architectures And Models

The rapid growth and ubiquitous adoption of the internet and cyber-physical systems (CPS) have fundamentally transformed modern communication, work, and human-system interactions. While networks now form the backbone of critical digital ecosystems, enabling seamless data transmission across diverse, interconnected systems, this increased connectivity also expands the attack surface, making real-time detection of network intrusions and anomalies a pressing challenge. Detecting unusual activities within network infrastructure requires advanced data traffic analysis to differentiate between legitimate and malicious interactions. Traditional approaches to network anomaly detectionâ??such as rule-based and signature-based systemsâ??often depend on predefined patterns to identify known anomalies, limiting their effectiveness against emerging, stealthy, or previously unseen threats. These conventional methods suffer from high false alarm rates and fail to adapt to the ever-evolving nature of network traffic, particularly in large-scale, decentralized environments where data volume, velocity, and variety are constantly increasing. This dissertation presents artificial intelligence (AI)-driven approaches to anomaly detection that leverage graphics processing unit (GPU)-enabled high-performance computing (HPC) platforms for processing massive network traffic data and monitoring the components of cyber-physical systems (CPS) for potentially hazardous conditions. The research advances several key contributions: (1) Designing efficient machine learning techniques for CPS condition monitoring and anomaly detection; (2) enabling federated learning (FL) frameworks that enable distributed detection while preserving data privacy and system resilience; (3) exploring graph-based methodologies combining graph neural networks (GNN) and graph machine learning (ML) approaches for the Internet of Things (IoT) and automotive network security, and (4) performing distributed edge computing optimizations that integrate FL with scalable technologies for reduced communication overhead. Through extensive experiments, these methodologies demonstrate that complex anomaly detection and condition monitoring tasks can be achieved while balancing computational efficiency and detection accuracy through fine-grained network information processing. The frameworks developed in this research establish a robust foundation for network anomaly detection, providing scalable, adaptive, and privacy-preserving solutions for safeguarding CPS and IoT networks in an increasingly interconnected digital landscape. The practical implications of these research findings are significant, as they can inform the development of next-generation network security systems and contribute to the protection of critical infrastructure against sophisticated cyber attacks.

Marfo, William↗

CONGO²: Scalable Online Anomaly Detection and Localization in Power Electronics Networks

Rapid and accurate detection and localization of electronic disturbances simultaneously are important for preventing its potential damages and determining potential remedies. Existing anomaly detection methods are severely limited by the low accuracy, the expensive computational cost and the need for highly trained personnel. There is an urgent need for a scalable online algorithm for in-field analysis of large-scale power electronics networks. Here in this paper, we propose a fast and accurate algorithm for anomaly detection and localization of power electronics networks: stratified colored-node graph (CONGO2). This algorithm hierarchically models the change of correlated waveforms and then correlated sensors using the colored-node graph. By aggregating the change of each sensor with its neighbors’ inputs, we can spontaneously identify and localize the anomaly that cannot be detected by data collected from a single sensor. As our proposed method only focuses on the changes within a short time frame, it is highly computational efficient and only needs small data storage. Thus, our method is ideal for online and reliable anomaly detection and localization of large-scale power electronic networks. Compared to existing anomaly detection methods, our method is entirely data-driven without training data, highly accurate and reliable for wide-spectrum anomalies detection, and more importantly, capable of both detection and localization. Thus, it is ideal for infield deployment for large-scale power electronic networks. As illustrated by a distributed energy resources (DERs) power grid with 37-node, our method can effectively detect and localize various cyber and physical attacks.

42 ENGINEERING↗

Capsule network-based semantic segmentation model for thermal anomaly identification on building envelopes

Thermography technology is widely used to inspect thermal anomalies in building façade systems. Computer vision-based techniques provide opportunities to autonomously detect such heat anomalies to significantly improve the efficiency of decision-making for building envelope retrofitting and maintenance. Here, in this work, we propose a novel Capsule Network-based deep learning model – CapsLab – that detects and identifies thermal anomalies by semantic segmentation. CapsLab is built based on our proposed prediction-tuning capsule (PT-Capsule) layer. Different from a traditional capsule layer, which consists of part-whole transformation and capsule-routing process, the proposed layer is composed of a prediction and tuning process, which helps decreasing the number of model parameters significantly. While the applicability of traditional Capsule Networks (CapsNets) has been limited to simpler tasks and smaller datasets due to their scalability issue, we can leverage the lightweight of the proposed PT-Capsule layer, and apply it to the semantic segmentation task. In this work, we also employ our previously presented performance metric, referred to as the Anomaly Identification Metric (AIM) (Kakillioglua et al. 2021), to evaluate the segmentation outputs. Traditional performance metrics do not accurately reflect the true performance of the segmentation models in thermal anomaly identification due to the high subjectivity in the annotation process and higher overlap ratio sensitivity of the standard metrics. AIM, on the other hand, is robust to these drawbacks. Experimental results show, both qualitatively and quantitatively, that our proposed segmentation method can effectively segment the thermal anomalies. Specifically, our model provides 9.38% and 13.53% improvements over the baseline model – DeepLabV3+ – based on traditional mIoU score and the AIM score, respectively, while requiring less model parameters and less computation at the same time. In addition, the scores that the AIM metric generates better align with the scores provided by building performance experts.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Monte Carlo Dropout Uncertainty Quantification of Long Short-Term Memory Autoencoder Anomaly Detection in a Liquid Sodium Cold Trap

Advanced high-temperature fluid reactors, such as sodium-cooled fast reactors (SFRs) and molten salt–cooled reactors (MSCRs), require coolant purification systems to prevent fluid contamination and local freezing that can lead to plugging. Liquid sodium purification can be achieved with a cold trap, where the sodium temperature is reduced to a near-freezing point to precipitate out impurities. Automation of monitoring of the cold trap performance with machine learning algorithms can aid in early detection of incipient anomalies. An efficient approach to loss-of-coolant–type anomaly detection in a cold trap monitored with more than two dozen thermal-hydraulic sensors consists of a long short-term memory (LSTM) autoencoder. This work develops the uncertainty quantification of the LSTM autoencoder performance for cold trap anomaly detection using the Monte Carlo (MC) dropout method. The MC dropout methodology creates a distribution of sister distributions that all slightly differ from each other because of random neurons being turned off for testing. The variances of the sister network distributions are used to make an uncertainty interval. Our analysis shows that the uncertainty in the autoencoder performance is largest near the peak of the anomaly signal. Using the MC dropout method, we investigate the uncertainty in the anomaly detection with missing sensor inputs. This capability allows the reactor operator to evaluate resilience of the anomaly detection system and to make informed decisions about continuity of operation in the event of sensor failure.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Platform for Automated Anomaly Detection in the Mercury Process System at the Target System in the Spallation Neutron Source

The Spallation Neutron Source (SNS) at Oak Ridge National Laboratory accelerates proton beams, which are directed toward a mercury target to generate the world’s most intense neutron beams via spallation. The target system consists of several interconnected subsystems and accounts for a major share of the facility’s overall downtime. Early detection of anomalies in the target system response can thus provide the possibility of taking corrective actions to reduce downtime. Accelerator facilities have largely focused on the beam side for data-driven fault prognostics. On the target side, SNS relies on operational shift technicians (OSTs), who respond to alarms and manually flag anomalies onto the System Tracking and Reliability (STAR) platform. This paper presents one of the first studies of using machine learning (ML) to automate anomaly detection in the target system. The study focused on the mercury process system as the first use case and employed reconstruction-based anomaly detection on minutely sampled time series signals. The pipeline was integrated into the STAR platform to autonomously rank and flag anomalies every week. The STAR platform provides a user interface for the OSTs to evaluate the flagged anomalies, thereby incorporating human feedback.

Anomaly detection↗

RX-ADS: Interpretable Anomaly Detection Using Adversarial ML for Electric Vehicle CAN Data

Recent year has brought considerable advancements in Electric Vehicles (EVs) and associated infrastructures/ communications. Intrusion Detection Systems (IDS) are widely deployed for anomaly detection in such critical infrastructures. This paper presents an Interpretable Anomaly Detection System (RX-ADS) for intrusion detection in CAN protocol communication in EVs. Contributions include: 1) window based feature extraction method; 2) deep Autoencoder based anomaly detection method; and 3) adversarial machine learning based explanation generation methodology. The presented approach was tested on two benchmark CAN datasets: OTIDS and Car Hacking. The anomaly detection performance of RX-ADS was compared against the state-of-the-art approaches on these datasets: HIDS and GIDS. The RX-ADS approach presented performance comparable to the HIDS approach (OTIDS dataset) and has outperformed HIDS and GIDS approaches (Car Hacking dataset). Further, the proposed approach was able to generate explanations for detected abnormal behaviors arising from various intrusions. Furthermore, these explanations were later validated by information used by domain experts to detect anomalies. Other advantages of RX-ADS include: 1) the method can be trained on unlabeled data; 2) explanations help experts in understanding anomalies and root course analysis, and also help with AI model debugging and diagnostics, ultimately improving user trust in AI systems.

42 ENGINEERING↗

A Pattern Dictionary Method for Anomaly Detection

In this paper, we propose a compression-based anomaly detection method for time series and sequence data using a pattern dictionary. The proposed method is capable of learning complex patterns in a training data sequence, using these learned patterns to detect potentially anomalous patterns in a test data sequence. The proposed pattern dictionary method uses a measure of complexity of the test sequence as an anomaly score that can be used to perform stand-alone anomaly detection. We also show that when combined with a universal source coder, the proposed pattern dictionary yields a powerful atypicality detector that is equally applicable to anomaly detection. The pattern dictionary-based atypicality detector uses an anomaly score defined as the difference between the complexity of the test sequence data encoded by the trained pattern dictionary (typical) encoder and the universal (atypical) encoder, respectively. We consider two complexity measures: the number of parsed phrases in the sequence, and the length of the encoded sequence (codelength). Specializing to a particular type of universal encoder, the Tree-Structured Lempel–Ziv (LZ78), we obtain a novel non-asymptotic upper bound, in terms of the Lambert W function, on the number of distinct phrases resulting from the LZ78 parser. This non-asymptotic bound determines the range of anomaly score. As a concrete application, we illustrate the pattern dictionary framework for constructing a baseline of health against which anomalous deviations can be detected.

97 MATHEMATICS AND COMPUTING↗