POWER DATA ANOMALY DETECTION PACKAGE
SF-25-150 Software package for developing anomaly detection models for electric power operational data
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
SF-25-150 Software package for developing anomaly detection models for electric power operational data
SF-25-081 Utility software for creating high-performance data pipelines to extract, load, and transform raw electric power systems measurements. For use with anomaly detection models training workflows. The software supports the project: Adaptive Cybersecurity for DER: A Game-Theoretic and Machine Learning approach for Real-Time Threat Detection and Mitigation
With the rapid evolution of malicious software, cyber threats have become increasingly sophisticated, employing advanced obfuscation techniques to evade traditional detection methods. This study presents a hybrid anomaly detection approach applied to obfuscated malware. Even though there is a large body of research in this field, existing malware detection techniques have some drawbacks, such as requiring large amounts of data, trustworthiness (imprecise results) of algorithms, and advanced obfuscation. To overcome these challenges, there is a need to employ solid and efficient techniques for malware detection. This paper proposes a hybrid approach, combining an autoencoder with traditional machine-learning methods to create an efficient malware detection framework. We used the malware memory dataset (MalMemAnalysis-2022) to evaluate this framework. The results indicate that our proposed approach can detect obfuscated malware when a deep autoencoder used for feature learning is combined with logistic regression, and it is extremely fast with an Accuracy, Detection Rate (DR), Matthew Correlation Coefficient(MCC), and Statistical Parity Difference
The U.S. Department of Energy’s Office of Nuclear Energy Advanced Materials and Manufacturing Technologies (AMMT) program is pursuing qualification of laser powder bed fusion (LPBF) components for nuclear applications. A major focus of this effort is the use of in situ process monitoring and machine learning–based tools to establish real-time quality assurance. The primary objective of this report is to identify and evaluate the most relevant in situ sensor systems for LPBF, and to document the deployment of these systems across platforms critical to the AMMT program. This work demonstrates how in situ monitoring can detect process anomalies, track geometry-dependent flaws, and identify limiting combinations of processing parameters—particularly those related to energy density and complex geometries (e.g., overhanging structures). To support this goal, a diverse suite of sensor modalities was evaluated across LPBF platforms, including visible and near-infrared (NIR) imaging, fringe projection profilometry, long-wavelength infrared (LWIR) thermography, and high-speed photodiode/pyrometry systems. These sensor streams were integrated with Peregrine, a machine-agnostic software platform that, among other capabilities, can generate real-time process anomaly classification. This report documents sensor deployments on multiple AMMT flagship platforms, including the Concept Laser M2 and Renishaw AM400/AM250 systems. Calibration builds with complex, flaw-prone geometries such as unsupported overhangs, stepped features, and thin walls, were used to evaluate how well Peregrine and its associated sensors could detect process anomalies and other instabilities under varied energy densities. It will be shown how Peregrine reliably identifies common process anomalies such as recoater streaking, superelevation, etc., and can be used in post-build analysis for anomaly spatial distributions throughout the build height to better understand the impact of geometry and processing parameter choice on the build. This work demonstrates measurable progress toward the vision that components can be born-qualified by establishing a real-time monitoring framework, identifying limiting process conditions, and laying the foundation for sensor fusion–enabled prediction pipelines that are scalable across platforms and applicable to nuclear-relevant components.
In-situ monitoring and anomaly detection are important components for qualification of directed energy deposition (DED) additive manufacturing (AM) processes and components. The use of in-situ monitoring requires an understanding of anomalies that can be identified during the process and how those anomalies correlate to mechanical properties of the component post-production. There is also a need to qualify the algorithms and software used to interpret the process signals for DED AM. There is no single process signal that can be used with a single algorithm that will identify all anomalies that will translate to a defect in a process. The process signals are affected by changes in material, location, resolution, acquisition rate, component geometry, and the machine itself. It is observed that multiple process signals are required to identify relevant features that can be correlated to mechanical properties.
The US Department of Energy’s Advanced Materials and Manufacturing Technologies (AMMT) program is pursuing rapid qualification of new materials for fabrication of nuclear relevant components using advanced manufacturing techniques. Particular interest is placed on code-qualifying stainless steel (SS) 316H processed by laser powder bed fusion (LPBF) additive manufacturing. A paradigm that incorporates data from in-situ sensing during the printing, ex-situ characterization, and advanced artificial intelligence–based models was established under the Transformation Challenge Reactor (TCR) program to develop a pedigree for each fabricated component that could be tracked from the feedstock to the component’s release for application. Under the TCR program, the Peregrine software was developed as a tool for incorporating the vast amounts of in-situ and ex-situ characterization data collected; all data stored on a rapidly growing digital platform. The digital platform allows for users to link site-specific process anomalies to the macro- and microstructure. The platform will eventually be able to predict component performance, which will be crucial to qualifying materials and components in risk-averse industries such as those supporting and building nuclear reactors. Current in-situ process monitoring techniques that are already integrated with software like Peregrine are advantageous for identifying process anomalies including powder spatter, component edge swelling, recoating-build interactions, and so on. However, additional data are required to fully predict the resulting microstructures needed for identifying relationships to component performance. The rapid cooling rates observed in LPBF are some of the highest of any bulk manufacturing process, resulting in heterogenous microstructures and typically causing anisotropy in mechanical properties. Moreover, evolved residual thermal stresses are high, which can cause severe defects such as delamination or cracking. Therefore, other in-situ monitoring methods are warranted for exploration to measure and map the thermal history, and potentially the stress state, of each build. This report summarizes different in-situ monitoring strategies proposed for LPBF with a focus on the more developed sensor systems. Novel capabilities for measuring melt pool temperatures are also addressed to better inform modeling efforts.
Peregrine, a software tool developed at Oak Ridge National Laboratory (ORNL), was used to collect and analyze in-situ monitoring (ISM) data from a Concept Laser M2 (Colibrium Additive) laser powder bed fusion (L-PBF) printer and an ExOne Innovent (Desktop Metal) binder jet printer. Data for four builds (print jobs) were saved to HDF5 (high performance data) files for release. Additionally, process anomalies were annotated by the authors across 37 image stacks (i.e., print layers) and are also provided as HDF5 files.
Cables are initially qualified for nuclear power plant use for 40 years. As plants extend their operating license to 60 and 80 years, justification for continued cable use must shift to a condition-based approach since it is cost prohibitive to completely replace cables that are likely still capable of performing their design function. The Pacific Northwest National Laboratory (PNNL) Accelerated and Real Time Experimental Nodal Analysis (ARENA) cable motor test bed was used to test the response of a commercial spread spectrum time domain reflectometry (SSTDR) system, a laboratory instrument software-controlled SSTDR, and a vector network analyzer-based frequency domain reflectometry (FDR) system to various cable anomalies. The three instrument systems were able to interrogate cables over a range of frequency bandwidths that can be helpful for human data analysis. Data were subjected to supervised and unsupervised machine learning (ML) analyses to distinguish normal undamaged cable responses from anomalous cable responses. Both supervised and unsupervised ML approaches produced encouraging results with an undamaged/anomalous prediction accuracy from 0.69% to 0.87%. Recommendations for further development and field implementation include increased and more balanced sample sets particularly including more training data.
For a quarter of a century, the LandScan Global (LSG) project has annually released a global, high-resolution gridded population dataset representing the ambient or unwarned population at a 30 arcsecond resolution. LSG supports a range of applications such as emergency management, disaster response, and human health and security for understanding populations at risk. The 2023 release of LSG, the LandScan Silver Edition, represents a major methodological leap forward while also leveraging previous knowledge—the previous year was the baseline for the current annual update carrying forward valuable knowledge of the built environment for the past quarter century—to train the machine learning models. Compared with annual releases over the past 24years, multiple advancements were made to different aspects of the methodology to achieve reproducibility, transparency, and consistent global propagation of solutions to modeling or population distribution issues identified during the review process. These novel changes include incorporation of the latest available geospatial inputs across the globe, machine learning models instead of manual modifications, population feature importance analysis, open-source solutions vs. proprietary software, generation of multiple global versions, analytic validations, and human-in-the-loop revisions to produce the final version. Additionally, algorithms—such as anomaly detection—were introduced to quickly identify areas of focus to develop a new and robust systematic review. Significant changes in modeled population distributions were observed between the 2022 and 2023 releases, largely attributable to improvements in data and methods and discussed thoroughly within this report. In summation, the LandScan Silver Edition leverages the best of the past quarter century of LSG legacy knowledge and continues a tradition of applying cutting-edge enhancements to serve as a new benchmark for accurate, actionable gridded population data
Foundation models use large datasets to build an effective representation of data that can be deployed on diverse downstream tasks. Previous research developed the omnilearn foundation model for jet physics, using unique properties of particle physics, and showed that it could significantly advance discovery potential across collider experiments. This paper introduces a major upgrade, resulting in the omnilearned framework. This framework has three new elements: (1) updates to the model architecture and training, (2) using over 1 × 10 9 jets used for training, and (3) providing well-documented software for accessing all datasets and models. We demonstrate omnilearned with three representative tasks: top-quark jet tagging with the community delphes-based benchmark dataset, b tagging with ATLAS full simulation, and anomaly detection with CMS experimental data. In each case, omnilearned is the state of the art, further expanding the discovery potential of past, current, and future collider experiments.
With increasing exposure to software-based sensing and control, power electronics systems are facing higher risks of cyber-physical attacks. To ensure system stability and minimize potential economic losses, it is critical to monitor the operating states and detect those attacks at the early stage. However, anomaly detection and diagnosis of attacks are still challenging, especially when labeled anomaly data is difficult or even infeasible to obtain. To overcome this problem, we propose a Few-Shot Learning (FSL) based approach for cyber-attack diagnosis leveraging the waveform data. To the best of our knowledge, this work is the first attempt at leveraging FSL for cyber-attack diagnosis in power electronics systems. Extensive experimental results demonstrate that our proposed approach can achieve comparable diagnosis accuracy with the state-of-the-art data-driven methods using less than 0.04% of the training samples.
This dataset provides partitioned evapotranspiration (ET, the combined loss of water from soil and plant surfaces) anomalies during heatwave events—soil evaporation (E) and transpiration (T)—for 268 heatwave events across 32 National Ecological Observatory Network (NEON) flux sites in the contiguous United States from 2019–2021. Using an ensemble of four high-frequency turbulence methods (Flux-variance Similarity, Conditional Eddy Covariance [CEC], CEC with Water-Use Efficiency, and Conditional Eddy Accumulation; see Zahn and Bou-Zeid 2024), half-hourly transpiration-to-evapotranspiration (T/ET) ratios were derived from 20 hertz (Hz, cycles per second) eddy covariance measurements of carbon dioxide (CO₂) and water vapor (H₂O) concentrations. The dataset spans six vegetation types including evergreen and deciduous forests, grasslands, cultivated crops, shrublands, and emergent herbaceous wetlands. Data Package Contents: The dataset includes a single CSV (comma-separated values) file containing daily anomalies (deviations from baseline conditions) for transpiration (Delta_T), evaporation (Delta_E), total evapotranspiration (Delta_ET), and T/ET ratio (Delta_T_ET) during each day of identified heatwave events. The file also includes site codes, dates, heatwave event identifiers, and day-of-heatwave indicators. The CSV file can be opened with spreadsheet software (Microsoft Excel, Google Sheets) or programming environments (Python, R, MATLAB). This resource enables researchers to investigate ecosystem-specific responses to thermal extremes, validate land surface model partitioning of ET fluxes, and examine feedbacks between water cycling and surface energy balance during heatwaves. The dataset is particularly valuable for studies linking vegetation hydraulic strategies to climate resilience, as it captures the divergent responses of shallow-rooted versus deep-rooted ecosystems. Potential applications include improving drought early warning systems, informing irrigation management strategies, and advancing our mechanistic understanding of land-atmosphere interactions under extreme heat conditions.
The ICARUS experiment is part of the Short-Baseline Neutrino program at Fermilab. Its primary objective is to explore the possible existence of sterile neutrinos in the O(1 eV) mass range and to clarify the anomalies observed in the Liquid Scintillator Neutrino Detector and MiniBooNE experiments. The ICARUS-T600 detector is a Liquid Argon Time Projection Chamber, capable of producing high-resolution 3D images and precise calorimetric measurements of ionizing particles. This technology allows for a detailed study of neutrino interactions across a broad energy range, from a few keV to several hundred GeV. The track reconstruction is achieved through a software framework that applies a series of pattern recognition algorithms, transforming raw detector signals into fully reconstructed event topologies. This process involves identifying interaction vertices, particle tracks, and electromagnetic showers within the TPC. However, in certain cases, these algorithms may mistakenly break a single particle track into several shorter segments, interpreting each as a distinct particle. Since track length is used to estimate the particle's energy, such fragmentation can result in an energy underestimation of several hundred MeV. Furthermore, when a track is split into multiple segments, the particle identification (which relies on analyzing the energy loss as a function of the residual range) may fail, potentially leading to the loss of the entire event. To mitigate this problem, we have developed a dedicated algorithm designed to identify and reconnect (“stitch”) the tracks that were erroneously divided into multiple segments.
This study introduces envelope- and machine learning (ML)-based electrical fault type detection algorithms for electrical distribution grids, advancing beyond traditional logic-based methods. The proposed detection model involves three stages: anomaly area detection, ML-based fault presence detection, and ML-based fault type detection. Initially, an envelope-based detector identifying the anomaly region was improved to handle noisier power grid signals from meters. The second stage acts as a switch, detecting the presence of a fault among four classes: normal, motor, switching, and fault. Finally, if a fault is detected, the third stage identifies specific fault types. This study explored various feature extraction methods and evaluated different ML algorithms to maximize prediction accuracy. The performance of the proposed algorithms is tested in an emulated software–hardware electrical grid testbed using different sample rate meters/relays, such as SEL735, SEL421, SEL734, SEL700GT, and SEL351S near and far from an inverter-based photovoltaic array farm. The performance outcomes demonstrate the proposed model’s robustness and accuracy under realistic conditions.
As energy demand rises, nuclear energy, particularly from reactors that use tristructural isotropic (TRISO) fuels, has gained attention due to the fuel’s enhanced resistance to radiation damage and high temperatures. This report investigates the modeling capabilities of the Gamma Detector Response and Analysis Software (GADRAS) for TRISO fuels, focusing on the gamma signatures of TRISO particles, which have not been extensively explored. Using the Monte Carlo N-Particle (MCNP) code as a benchmark, we developed both homogeneous and heterogeneous models of TRISO pebbles to analyze gamma spectra. Our findings reveal that the homogeneous and heterogeneous models produced different gamma signatures. Additionally, the GADRAS heterogeneous model significantly reduces computation times compared to MCNP, enabling effective modeling of gamma signatures for safeguards applications. This advancement is essential for the International Atomic Energy Agency (IAEA) in detecting anomalies and potential smuggling attempts in TRISO reactor fuel elements.
In this paper, we propose a Hardware-in-the-Loop (HIL) simulation testbed suitable for the implementation and testing of realistic cyberattacks on grid-tied smart inverter systems integrated with Distributed Energy Resources (DER) that use the Distributed Network Protocol-3 (DNP3) protocol for communications between grid components. Specifically, our testbed combines a Real-Time Digital Simulator (RTDS) NovaCor device, outfitted with GNETx2 network interface cards, a gridtied DER topology implemented via the RTDS software package RSCAD, and a custom virtual network that emulates a man in the middle attacker. The Man-in-the-Middle (MITM) attacker captures DNP3 traffic and falsifies telemetry data in DNP3 packets to trigger unwarranted commands from a DNP3 controller that exploit smart inverter grid support functions. We choose DNP3 and implement grid support functions according to the IEEE Std. 1547-2018 mandated for the interconnection and interoperability of DER power systems with associated power components. Furthermore, we develop a protocol payload agnostic attack detection framework that leverages the round-trip time (RTT) anomalies between DNP3 requests and responses and can detect the presence of attacks without having to analyze the payload’s contents, while balancing trade-offs between false alarm counts, missed detections, and time to detection. To facilitate further research, we publicly release benign and attack network traffic exchanged between various sensors, controllers, and actuators in our grid-tied inverter testbed.
This project developed a cutting-edge 5G-integrated edge computing framework to enhance operational efficiency and reliability in coal-fired power plants through real-time component monitoring and anomaly detection. The initiative focused on leveraging distributed machine learning, federated learning, and 5G-based dynamic network slicing to support scalable, fault-tolerant monitoring environments to meet the operational requirements in industrial control systems. With a Distributed Edge Computing Service (DECS) orchestration, this project enabled federated learning at edge for condition monitoring and introduced adaptive client selection strategies to minimize communication overhead. Scalable distributed training was achieved using the Horovod framework, thus enhancing performance across edge nodes. In the realm of 5G networking, the project designed and deployed reconfigurable, QoS-aware network slicing tailored for operational technology (OT) environments, integrating software-defined networks to bolster cyber-resilience and enabling dynamic slicing for federated learning workloads. A significant milestone was the development of a virtualized ICS environment with 5G core integration—which allowed elastic and fault tolerant distributed training on real-world datasets such as NASA Bearings, Hydraulic Systems, and TEP. To broaden the impact of the project, a TRL-3 virtualized ICS testbed for research and education was designed. This project engaged several graduate and undergraduate students to conduct research on the cutting-edge technology, and it resulted in one PhD dissertation, one MS thesis, and over 14 peer-reviewed publications. With the support of this project students also participated in national cybersecurity competitions to improve their professional development skills.
Historically, cables are initially qualified for nuclear power plant use for 40 years. As plants extend their operating license to 60 and 80 years, continued use of these cables must shift to a performance-based approach since it is cost prohibitive to completely replace cables that are likely still capable of performing their design function. A variety of cable tests are available and are commonly applied during outages when the cables can be taken out of service. Frequency domain reflectometry (FDR) is one of these test methods that is being more broadly accepted and used because it not only detects anomalies along the cable with a low-voltage signal that does not stress the cable insulation, but the technique also locates the anomalies. This supports follow-up local inspection and local repair or partial replacement of a damaged cable segment. Currently, FDR testing is only applied to cables that are taken out of service since the test instrument would be damaged by operational voltages. A related technology that has found some acceptance in the aircraft and rail industry is spread spectrum time domain reflectometry (SSTDR). This technology has been implemented with a custom commercial instrument by LiveWire Innovation that is designed to operate on live cables up to 1000 volts and with a bandwidth of 48 MHz. Initial evaluation by the Pacific Northwest National Laboratory (PNNL) of the Live Wire system indicated that a broader bandwidth (BW) SSTDR may be better for many kinds of flaws. This led PNNL to develop an SSTDR laboratory instrument suitable for tests up to 500 MHz bandwidth. Testing on energized cables is also desirable for online monitoring systems so an inductive clamshell coupler was developed that allows energized cables to be tested up to at least 5 kV and likely higher voltage levels. Dielectric spectroscopy and tan delta testing plus various laboratory destructive tests were included in this data acquisition campaign directed to feed a machine learning (ML) study. With these kinds of developments, online energized cable tests may be possible with industrial adoption of such hardware advances but it will be completely impractical to have highly skilled data analysts continually examine these complex signals for indications of damage or compromised conditions. If online testing is to be implemented in new test hardware, it must be accompanied by software that can interpret the signals and alert plant operators of changing or degraded conditions. The thermally aged, shielded cable investigated here was separately treated for ML analysis. Visual analysis of electrical data showed generally increasing peaks where the cable entered and exited the oven. These peaks were not exactly aligned with expected locations, but these differences were attributed to velocity of propagation calibration errors. Only supervised ML was applied to the thermally aged data as this data was only available shortly before the committed publication date of this report. The supervised ML was structured to divide the 0 to 70-day responses as ‘normal’ from 0 to 35 days or ‘anomalous’ from 36 to 70 days, based on cable tensile elongation at break (EAB) insulation characterization. Using 80% of the data for training and 20% for testing, the supervised ML predicted normal versus anomalous was 70% accurate. Important conclusions include: • Accuracy to predict the presence of cable damage is improved from the 2023 effort by more training data. Weighted accuracies for comparisons among the instruments ranged from 67 to 89 % for unsupervised ML and 71 to 99% for supervised ML. • Based on the synthetic data tests, the unsupervised models are more generalizable to unseen anomalies. The Multi-Layer Perceptron classifier (MLP) model reported as high as 99.7% accuracy on the test data, but this dropped to 58.3% when tested on the synthetic data. In contrast, the unsupervised Pointwise model only achieved 89.7% accuracy on the experimental data but reported 78.3% accuracy on the synthetic data. • The best anomaly indicators are higher frequency (400 MHz BW) FDR data. Other tests may be interesting but for this study, this was the best predicter.