Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “unsupervised method”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Sensor Anomaly Detection for Nuclear Reactor Systems Utilizing Linear Regression and K-Means Unsupervised Machine Learning

Nuclear reactors and related systems are becoming increasingly complex due to advancing technologies in next-generation power reactors. This increased complexity necessitates enhanced automation and data management capabilities. To successfully realize autonomous systems, methods must be developed to handle vast volumes of data and effectively distinguish anomalous data from noise and expected data. While impressive models utilizing digital twins and similar approaches are under development, here we propose a simplified model for analyzing fundamental methods and techniques. Initially, we created a general dataset by using initial data from PCTRAN in order to represent ideal steady-state conditions. We then inserted anomalies based on prevalent sensor anomaly types (e.g., point anomalies, linear drift, and downward deviations), along with unusual anomalies such as exponential drift and upward deviations. To detect anomalies, we developed a program that employs data partitioning and linear regression to preprocess and filter the anomalous data. A K-Means machine learning (ML) method was then applied to separate and count the data within the anomalous partition. The results from all datasets—apart from exponential growth—demonstrated positive outcomes, with each returning multiple instances of greaterthan-95% accuracy. We conducted further investigations using Idaho National Laboratory’s RAVEN software to perform a sensitivity analysis on the input variables (R 2 Tolerance, Slope Tolerance, and Window Size) and found that the output variables (Accuracy and Time) were most sensitive to the Window Size. Despite the promising results published, further development is required to effectively apply these methods to nuclear systems. Nevertheless, the strengths of this approach are evident and hold promise for future applications in the field.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Distinguishing isotropic and anisotropic signals for X-ray total scattering using machine learning

Understanding structure–property relationships is essential for advancing technologies based on thin films. X-ray pair distribution function (PDF) analysis can access relevant atomic structure details spanning local-, mid- and long-range structure. While X-ray PDF has been adapted for thin films on amorphous substrates, measurements on single-crystal substrates are necessary to accurately determine structure origins for some thin film materials, especially those for which the substrate changes the accessible structure and properties. However, when measuring films on single-crystal substrates, high-intensity anisotropic Bragg spots saturate 2D detector images, overshadowing the thin films' isotropic scattering signal. This renders previous data processing methods for films on amorphous substrates unsuitable for films on single-crystal substrates. To address this measurement need, we developed IsoDAT2D, an innovative data processing approach using unsupervised machine learning algorithms. The program combines dimensionality reduction and clustering algorithms to separate thin film and single-crystal substrate X-ray scattering signals. We use SimDAT2D , a program we developed to generate simulated thin film data, to validate IsoDAT2D . Here we also use IsoDAT2D to isolate X-ray total scattering signal from a thin film on a single-crystal substrate. The resulting PDF data are compared with similar data processed using previous methods, especially substrate subtraction for single-crystal and amorphous substrates. PDF data from IsoDAT2D -identified X-ray total scattering data are significantly better than from single-crystal substrate subtraction, but not as reliable as PDF data from amorphous substrate subtraction. With IsoDAT2D , there are new opportunities to expand PDF to a wider variety of thin films, including those on single-crystal substrates, with which new structure–property relationships can be elucidated to enable fundamental understanding and technological advances.

36 MATERIALS SCIENCE↗

Physics-constrained superresolution diffusion for six-dimensional phase space diagnostics

Adaptive physics-constrained superresolution diffusion is developed for noninvasive virtual diagnostics of the six-dimensional (6D) phase space density of charged particle beams. An adaptive variational autoencoder embeds initial beam condition images and scalar measurements to a low-dimensional latent space from which a 32 6 pixel 6D tensor representation of the beam's 6D phase space density is generated. Projecting from a 6D tensor generates physically consistent two-dimensional projections. Physics-guided superresolution diffusion transforms low-resolution images of the 6D density to high resolution 256 × 256 pixel images. Unsupervised adaptive latent space tuning enables tracking of time-varying beams without knowledge of time-varying initial conditions. The method is demonstrated with experimental data and multiparticle simulations at the HiRES UED. The general approach is applicable to a wide range of complex dynamic systems evolving in high-dimensional phase space. The method is shown to be robust to distribution shift without retraining. Published by the American Physical Society 2025

43 PARTICLE ACCELERATORS↗

Image processing developments and applications for water quality monitoring and trophic state determination

Remote sensing data analysis of water quality monitoring is evaluated. Data anaysis and image processing techniques are applied to LANDSAT remote sensing data to produce an effective operational tool for lake water quality surveying and monitoring. Digital image processing and analysis techniques were designed, developed, tested, and applied to LANDSAT multispectral scanner (MSS) data and conventional surface acquired data. Utilization of these techniques facilitates the surveying and monitoring of large numbers of lakes in an operational manner. Supervised multispectral classification, when used in conjunction with surface acquired water quality indicators, is used to characterize water body trophic status. Unsupervised multispectral classification, when interpreted by lake scientists familiar with a specific water body, yields classifications of equal validity with supervised methods and in a more cost effective manner. Image data base technology is used to great advantage in characterizing other contributing effects to water quality. These effects include drainage basin configuration, terrain slope, soil, precipitation and land cover characteristics.

Blackwell, R. J.↗

Learning and Tuning of Fuzzy Rules

In this chapter, we review some of the current techniques for learning and tuning fuzzy rules. For clarity, we refer to the process of generating rules from data as the learning problem and distinguish it from tuning an already existing set of fuzzy rules. For learning, we touch on unsupervised learning techniques such as fuzzy c-means, fuzzy decision tree systems, fuzzy genetic algorithms, and linear fuzzy rules generation methods. For tuning, we discuss Jang's ANFIS architecture, Berenji-Khedkar's GARIC architecture and its extensions in GARIC-Q. We show that the hybrid techniques capable of learning and tuning fuzzy rules, such as CART-ANFIS, RNN-FLCS, and GARIC-RB, are desirable in development of a number of future intelligent systems.

Berenji, Hamid R.↗

A review on recent machine learning applications for imaging mass spectrometry studies

Imaging mass spectrometry (IMS) is a powerful analytical technique widely used in biology, chemistry, and materials science fields that continue to expand. IMS provides a qualitative compositional analysis and spatial mapping with high chemical specificity. The spatial mapping information can be 2D or 3D depending on the analysis technique employed. Due to the combination of complex mass spectra coupled with spatial information, large high-dimensional datasets (hyperspectral) are often produced. Therefore, the use of automated computational methods for an exploratory analysis is highly beneficial. The fast-paced development of artificial intelligence (AI) and machine learning (ML) tools has received significant attention in recent years. These tools, in principle, can enable the unification of data collection and analysis into a single pipeline to make sampling and analysis decisions on the go. There are various ML approaches that have been applied to IMS data over the last decade. Here, in this review, we discuss recent examples of the common unsupervised (principal component analysis, non-negative matrix factorization, k-means clustering, uniform manifold approximation and projection), supervised (random forest, logistic regression, XGboost, support vector machine), and other methods applied to various IMS datasets in the past five years. The information from this review will be useful for specialists from both IMS and ML fields since it summarizes current and representative studies of computational ML-based exploratory methods for IMS.

47 OTHER INSTRUMENTATION↗

Classifying biophysical subpopulations of insulin secretory granules using quantitative whole-cell structure analysis

Pancreatic beta cells contain insulin secretory granules (ISGs), organelles where proinsulin is converted into insulin. As ISGs mature, they undergo extensive biophysical remodeling, producing a spectrum of subpopulations with heterogeneous molecular and spatial characteristics. However, systematic methods to define ISG subpopulations remain underdeveloped. To address this gap in knowledge, we employed soft X-ray tomography (SXT), which can quantitatively measure the biochemical density of ISGs within whole beta cells. Using unsupervised clustering, we classified subpopulations based on molecular density, size, and spatial positioning. Across different insulin secretory stimuli, we observed shifts toward mature and releasable subtypes, demonstrating that exogenous signals can dynamically remodel ISG subpopulation distributions. We extended this methodology to primary beta cells characterized using volume electron microscopy (vEM). Integrating subpopulations from SXT and vEM uncovered insights inaccessible by a single method in isolation. This strategy establishes a framework for defining therapeutic approaches aimed at enriching physiologically beneficial ISG subpopulations.

dense-core granules↗

Cy-Phy ADS: Cyber Physical Anomaly Detection Framework for EV Charging Systems

Today’s large-scale Electric Vehicle (EV) infrastructures are heavily dependent on information communication technologies to maintain their operation and to support communication within sub-system components as well as the outside world. These technologies are vulnerable to various cyber and physical threats. Timely identification and mitigation of these threats are critical for improving human safety, avoiding economic losses, and preventing catastrophic system failures. By addressing this, our work presents a ResNet Autoencoder (AE) based Cyber-Physical Anomaly Detection System (Cy-Phy ADS) for detecting anomalies in EV Controller Area Network (CAN) protocol communication. It consists of four main components: Cyber-Physical Feature Extractor, ResNet AE-based Anomaly Detection Framework, Cyber-Physical Health Metric (CPHM), and Visualization Dashboard. The presented framework was trained and tested using CAN data collected from the EV charging system testbed at the Idaho National Laboratory. The presented Cy-Phy ADS compared against six widely used unsupervised anomaly detection algorithms: One Class Support Vector Machine (OCSVM), Variational Autoencoder (VAE), LSTM Autoencoder (LSTM AE), Isolation Forest (IForest), Principle Component Analysis (PCA) and Local Outlier Factor (LOF). Here the presented approach showed the highest accuracy among the compared methods. Further, the proposed approach showed comparable performance in terms of precision, F1, and False positive rate. It also showed the lowest training and inference time compared to the neural network-based baseline algorithms compared against with. Additionally, the Cy-Phy ADS has advantages such as unsupervised training, the ability to provide a holistic metric for system health characterization, and non-linear feature extraction.

99 GENERAL AND MISCELLANEOUS↗

VoroClust: Scalable Clustering for Remote Sensing

Although supervised machine learning provides a powerful framework for image classification and segmentation, it requires comprehensive consistent datasets, which are not available for many remote-sensing applications. Remote-sensing datasets are expensive to collect, and each is acquired under different environmental conditions or with significant variations in system operating parameters. Unsupervised clustering algorithms analyze the structure of each dataset independently, rather than drawing on similarities with existing “training” examples, and are thus well suited for practical remote-sensing applications. We introduce VoroClust, a fast density-based unsupervised clustering algorithm applicable to high-resolution and high-dimensional data. VoroClust runs as fast as distance-based clustering methods, while capturing complex regional geometries at least as well as current-density-based methods. It uses a data-centered sphere cover to reduce computational demands, while still capturing data topology. It then propagates clusters outward from local peaks in density. We show that VoroClust provides fast state-of-the-art clustering for both high-resolution polarimetric synthetic aperture radar and high-dimensional hyperspectral imaging datasets.

42 ENGINEERING↗

Earthquake Phase Association Using a Bayesian Gaussian Mixture Model

Earthquake phase association algorithms aggregate picked seismic phases from a network of seismometers into individual seismic events and play an important role in earthquake monitoring and research. Dense seismic networks and improved phase picking methods produce massive seismic phase datasets, particularly for earthquake swarms and aftershocks occurring closely in time and space, making phase association a challenging problem. Here, we present a new association method, the Gaussian Mixture Model Association (GaMMA), that combines the Gaussian mixture model with earthquake location, origin time, and magnitude estimation. We treat earthquake phase association as an unsupervised clustering problem in a probabilistic framework, where each earthquake corresponds to a cluster of P and S phases with a hyperbolic moveout of arrival times and a decay of amplitude with distance. We use the multivariate Gaussian distribution to model the collection of phase picks of an event; and the mean of the multivariate Gaussian distribution is given by the predicted arrival time and amplitude from the causative event. We carry out the pick assignment to each earthquake and determine earthquake source parameters (i.e., earthquake location, origin time, and magnitude) under the maximum likelihood criterion using the Expectation-Maximization algorithm. The GaMMA method does not require typical association steps of other algorithms, such as grid-search or supervised training. The results for both synthetic tests and for the 2019 Ridgecrest earthquake sequence show that GaMMA effectively associates phases from a temporally and spatially dense earthquake sequence while producing useful estimates of earthquake location and magnitude.

58 GEOSCIENCES↗

Coordinate-Based Seismic Interpolation in Irregular Land Survey: A Deep Internal Learning Approach

Physical and budget constraints often result in irregular sampling, which complicates accurate subsurface imaging. Preprocessing approaches, such as missing trace or shot interpolation, are typically employed to enhance seismic data in such cases. Recently, deep learning has been used to address the trace interpolation problem at the expense of large amounts of training data to adequately represent typical seismic events. Nonetheless, most research in this area has focused on trace reconstruction, with little attention having been devoted to shot interpolation. Furthermore, existing methods assume regularly spaced receivers/sources failing in approximating seismic data from real (irregular) surveys. This work presents a novel shot gather interpolation approach which uses a continuous coordinate-based representation of the acquired seismic wavefield parameterized by a neural network. The proposed unsupervised approach, which we call coordinate-based seismic interpolation (CoBSI), enables the prediction of specific seismic characteristics in irregular land surveys without using external data during neural network training. Importantly, experimental results on real and synthetic 3-D data validate the ability of the proposed method to estimate continuous smooth seismic events in the time-space and frequency-wavenumber domains, improving sparsity or low-rank-based interpolation methods.

58 GEOSCIENCES↗

via machinae : Searching for stellar streams using unsupervised machine learning

ABSTRACT We develop a new machine learning algorithm, via machinae, to identify cold stellar streams in data from the Gaia telescope. via machinae is based on ANODE, a general method that uses conditional density estimation and sideband interpolation to detect local overdensities in the data in a model agnostic way. By applying ANODE to the positions, proper motions, and photometry of stars observed by Gaia, via machinae obtains a collection of those stars deemed most likely to belong to a stellar stream. We further apply an automated line-finding method based on the Hough transform to search for line-like features in patches of the sky. In this paper, we describe the via machinae algorithm in detail and demonstrate our approach on the prominent stream GD-1. Though some parts of the algorithm are tuned to increase sensitivity to cold streams, the via machinae technique itself does not rely on astrophysical assumptions, such as the potential of the Milky Way or stellar isochrones. This flexibility suggests that it may have further applications in identifying other anomalous structures within the Gaia data set, for example debris flow and globular clusters.

79 ASTRONOMY AND ASTROPHYSICS↗

STSR-INR: Spatiotemporal super-resolution for multivariate time-varying volumetric data via implicit neural representation

Implicit neural representation (INR) has surfaced as a promising direction for solving different scientific visualization tasks due to its continuous representation and flexible input and output settings. We present STSR-INR, an INR solution for generating simultaneous spatiotemporal super-resolution for multivariate time-varying volumetric data. Inheriting the benefits of the INR-based approach, STSR-INR supports unsupervised learning and permits data upscaling with arbitrary spatial and temporal scale factors. Unlike existing GAN- or INR-based super-resolution methods, STSR-INR focuses on tackling variables or ensembles and enabling joint training across datasets of various spatiotemporal resolutions. Here we achieve this capability via a variable embedding scheme that learns latent vectors for different variables. In conjunction with a modulated structure in the network design, we employ a variational auto-decoder to optimize the learnable latent vectors to enable latent-space interpolation. To combat the slow training of INR, we leverage a multi-head strategy to improve training and inference speed with significant speedup. We demonstrate the effectiveness of STSR-INR with multiple scalar field datasets and compare it with conventional tricubic+linear interpolation and state-of-the-art deep-learning-based solutions (STNet and CoordNet).

97 MATHEMATICS AND COMPUTING↗

Application of remote sensing techniques for identification of irrigated crop lands in Arizona

Satellite imagery was used in a project developed to demonstrate remote sensing methods of determining irrigated acreage in Arizona. The Maricopa water district, west of Phoenix, was chosen as the test area. Band rationing and unsupervised categorization were used to perform the inventory. For both techniques the irrigation district boundaries and section lines were digitized and calculated and displayed by section. Both estimation techniques were quite accurate in estimating irrigated acreage in the 1979 growing season.

Billings, H. A.↗

SRF cavity instability detection with machine learning at CEBAF

During the operation of the Continuous Electron Beam Accelerator Facility (CEBAF), one or more unstable superconducting radio-frequency (SRF) cavities often cause beam loss trips while the unstable cavities themselves do not necessarily trip off. The present RF controls for the legacy cavities report at only 1 Hz, which is too slow to detectfast transient instabilities during these trip events. These challenges make the identification of an unstable cavity out of the hundreds installed at CEBAF a difficult and time-consuming task. To tackle these issues, a fast data acquisition system (DAQ) for the legacy SRF cavities has been developed, which records the sample at 5 kHz. An unsupervised learning framework has been developed to identify anomalous SRF cavity behavior. We will discuss the present status of the DAQ system and our framework, along with recent successes in detecting anomalous cavity behavior. Overall, our method offers a practical solution for identifying unstable SRF cavities, contributing to increased beam availability and machine reliability.

Accelerator Physics↗

Seasonal and Geographical Variations in Fundamental Weather Patterns during Extreme Precipitation as Identified from Omega Equation Forcing

Abstract We investigate the large-scale weather patterns during extreme precipitation (PEx) events over the conterminous United States (CONUS) by applying a version of the quasigeostrophic (QG) omega equation. This work aims to develop a climatology of the weather patterns most related to PEx events during current climate. Extreme events are examined for each of seven regions defined by consistent annual cycles of precipitation and spanning the CONUS. For the CONUS we train several self-organizing maps (SOM) on a pressure–time series of vertical velocity from each of the advective forcing terms in the QG omega equation for each extreme event. The unsupervised learning of the SOM allows us to identify the most descriptive set of nine patterns in vertical velocity associated with precipitation extremes. This method finds multiple frontal- and cyclone-driven patterns while grouping primarily convective events into one pattern. Frontal events include a synoptic pattern consistent with West Coast atmospheric river events as well as pattern groups linked to developing and to mature (“occluded”) frontal cyclones. The primary patterns found during PEx events vary seasonally and geographically. Frontal cyclone patterns are most common during PEx events during summer in the part of the Great Plains and during winter for the Northeast, Southeast, Pacific Northwest, and Southwest. Convection is the most common pattern during summer in all regions. Except in the Southeast, the annual cycles of monthly number of PEx events and average precipitation match well, partially validating our choice of regions to aggregate PEx events.

Meteorology & Atmospheric Sciences↗

The LSST AGN Data Challenge: Selection Methods

Abstract Development of the Rubin Observatory Legacy Survey of Space and Time (LSST) includes a series of Data Challenges (DCs) arranged by various LSST Scientific Collaborations that are taking place during the project's preoperational phase. The AGN Science Collaboration Data Challenge (AGNSC-DC) is a partial prototype of the expected LSST data on active galactic nuclei (AGNs), aimed at validating machine learning approaches for AGN selection and characterization in large surveys like LSST. The AGNSC-DC took place in 2021, focusing on accuracy, robustness, and scalability. The training and the blinded data sets were constructed to mimic the future LSST release catalogs using the data from the Sloan Digital Sky Survey Stripe 82 region and the XMM-Newton Large Scale Structure Survey region. Data features were divided into astrometry, photometry, color, morphology, redshift, and class label with the addition of variability features and images. We present the results of four submitted solutions to DCs using both classical and machine learning methods. We systematically test the performance of supervised models (support vector machine, random forest, extreme gradient boosting, artificial neural network, convolutional neural network) and unsupervised ones (deep embedding clustering) when applied to the problem of classifying/clustering sources as stars, galaxies, or AGNs. We obtained classification accuracy of 97.5% for supervised models and clustering accuracy of 96.0% for unsupervised ones and 95.0% with a classic approach for a blinded data set. We find that variability features significantly improve the accuracy of the trained models, and correlation analysis among different bands enables a fast and inexpensive first-order selection of quasar candidates.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Exploratory analysis of machine learning techniques in the Nevada geothermal play fairway analysis

Play fairway analysis (PFA) is commonly used to generate geothermal potential maps and guide exploration studies, with a particular focus on locating and characterizing blind geothermal systems. This study evaluates the application of machine learning techniques to PFA in the Great Basin region of Nevada. Following the evaluation of various techniques, we identified two approaches to PFA that produced promising results, 1) supervised Bayesian probabilistic neural networks to generate geothermal potential maps with confidence intervals, and 2) unsupervised principal component analysis paired with k-means clustering to generate both cluster maps to help identify spatial patterns, as well as new combined feature inputs. We applied these techniques to perform a comparative analysis between two principal sets of geological and geophysical features related to permeability and heat and a set of positive (known geothermal resources) and negative training sites (known drill sites with unsuitable geothermal conditions). We found that these methods constrain previously unrecognized feature controls on geothermal favorability, many of which are spatially organized within the extent of cluster groups and the major structural-hydrologic domains of the study area. Furthermore, we utilized exploratory unsupervised modeling to highlight spatial relationships between input data and predictive output results of our supervised modeling. As a result, we demonstrate how our models compare to the previous Nevada PFA and how the rapid insights these machine learning techniques offer may support future assessments of both known and undiscovered blind geothermal systems in the Great Basin region of Nevada and beyond.

15 GEOTHERMAL ENERGY↗