Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “unsupervised method”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Unsupervised learning of ferroic variants from atomically resolved STEM images

An approach for the analysis of atomically resolved scanning transmission electron microscopy data with multiple ferroic variants in the presence of imaging non-idealities and chemical variabilities based on a rotationally invariant variational autoencoder (rVAE) is presented. We show that an optimal local descriptor for the analysis is a sub-image centered at specific atomic units, since materials and microscope distortions preclude the use of an ideal lattice as a reference point. The applicability of unsupervised clustering and dimensionality reduction methods is explored and is shown to produce clusters dominated by chemical and microscope effects, with a large number of classes required to establish the presence of rotational variants. Comparatively, the rVAE allows extraction of the angle corresponding to the orientation of ferroic variants explicitly, enabling straightforward identification of the ferroic variants as regions with constant or smoothly changing latent variables and sharp orientational changes. This approach allows further exploration of the chemical variability by separating the rotational degrees of freedom via rVAE and searching for remaining variability in the system. The code used in this article is available at https://github.com/saimani5/ferroelectric_domains_rVAE.

30 DIRECT ENERGY CONVERSION↗

On-line object feature extraction for multispectral scene representation

A new on-line unsupervised object-feature extraction method is presented that reduces the complexity and costs associated with the analysis of the multispectral image data and data transmission, storage, archival and distribution. The ambiguity in the object detection process can be reduced if the spatial dependencies, which exist among the adjacent pixels, are intelligently incorporated into the decision making process. The unity relation was defined that must exist among the pixels of an object. Automatic Multispectral Image Compaction Algorithm (AMICA) uses the within object pixel-feature gradient vector as a valuable contextual information to construct the object's features, which preserve the class separability information within the data. For on-line object extraction the path-hypothesis and the basic mathematical tools for its realization are introduced in terms of a specific similarity measure and adjacency relation. AMICA is applied to several sets of real image data, and the performance and reliability of features is evaluated.

Ghassemian, Hassan↗

Clonal Selection Based Artificial Immune System for Generalized Pattern Recognition

The last two decades has seen a rapid increase in the application of AIS (Artificial Immune Systems) modeled after the human immune system to a wide range of areas including network intrusion detection, job shop scheduling, classification, pattern recognition, and robot control. JPL (Jet Propulsion Laboratory) has developed an integrated pattern recognition/classification system called AISLE (Artificial Immune System for Learning and Exploration) based on biologically inspired models of B-cell dynamics in the immune system. When used for unsupervised or supervised classification, the method scales linearly with the number of dimensions, has performance that is relatively independent of the total size of the dataset, and has been shown to perform as well as traditional clustering methods. When used for pattern recognition, the method efficiently isolates the appropriate matches in the data set. The paper presents the underlying structure of AISLE and the results from a number of experimental studies.

pattern recognition↗

GeoThermalCloud: Machine Learning for Geothermal Resource Exploration

Geothermal is a renewable energy source that can provide reliable and flexible electricity generation for the world. In the past decade, the U.S. Geological Survey's resource assessments, Play Fairway Analyses (PFA), and GeoVision report by the U.S. Department of Energy's Geothermal Technologies Office provided insights on enormous untapped potential for geothermal energy to contribute to the U.S. domestic energy needs. The past studies identified that geothermal resources without surface expression (e.g., blind/hidden hydrothermal systems) comprise a huge potential. These blind systems can significantly increase power generation. But a primary challenge is locating and quantifying these hidden resources, which do not have any thermal manifestations on the surface. PFA has successfully identified some blind systems in the western USA (e.g., specific locations in the Great Basin region within Nevada). However, a comprehensive search for these blind systems can be time-consuming, expensive, and resource-intensive with a low probability of success. Accelerated discovery of these blind resources is needed with growing energy needs and higher chances of exploration success. Recent advances in machine learning (ML) have shown promise in shortening the timeline for this discovery. This paper presents a novel ML-based methodology for geothermal exploration towards PFA applications. Our methodology is provided through our open-source ML framework called GeoThermalCloud \url{https://github.com/SmartTensors/GeoThermalCloud.jl}. GeoThermalCloud uses a series of unsupervised, supervised, and physics-informed ML methods available in SmartTensors AI platform \url{https://github.com/SmartTensors}. Here, the presented analyses are performed using our unsupervised ML algorithm called NMF$k$, which is available in the SmartTensors AI platform. Our ML algorithm facilitates the discovery of new phenomena, hidden patterns, and mechanisms that helps us to make informed decisions. Moreover, the GeoThermalCloud enhances the collected PFA data and discovers signatures representative of geothermal resources. Through GeoThermalCloud, we were able to identify hidden patterns in the geothermal field data needed for the efficient discovery of blind systems. Crucial geothermal signatures often overlooked in traditional PFA are extracted using GeoThermalCloud and analyzed by the subject matter experts to provide ML-enhanced PFA, which is informative for efficient exploration. We applied our ML methodology on various open-source geothermal datasets within the U.S. (some of these are collected by past PFA work), and the results provide valuable insights on resource types within those explored regions. This ML-enhanced workflow makes GeoThermalCloud attractive for the geothermal community to improve existing datasets and extract valuable information often unnoticed during geothermal exploration.

machine learning (ML), geothermal energy↗

Optimizing training trajectories in variational autoencoders via latent Bayesian optimization approach *

Unsupervised and semi-supervised ML methods such as variational autoencoders (VAE) have become widely adopted across multiple areas of physics, chemistry, and materials sciences due to their capability in disentangling representations and ability to find latent manifolds for classification and/or regression of complex experimental data. Like other ML problems, VAEs require hyperparameter tuning, e.g. balancing the Kullback–Leibler and reconstruction terms. However, the training process and resulting manifold topology and connectivity depend not only on hyperparameters, but also their evolution during training. Because of the inefficiency of exhaustive search in a high-dimensional hyperparameter space for the expensive-to-train models, here we have explored a latent Bayesian optimization (zBO) approach for the hyperparameter trajectory optimization for the unsupervised and semi-supervised ML and demonstrated for joint-VAE with rotational invariances. We have demonstrated an application of this method for finding joint discrete and continuous rotationally invariant representations for modified national institute of standards and technology database (MNIST) and experimental data of a plasmonic nanoparticles material system. The performance of the proposed approach has been discussed extensively, where it allows for any high dimensional hyperparameter trajectory optimization of other ML models.

42 ENGINEERING↗

Unsupervised Learning Based Interaction Force Model for Nonspherical Particles in Incompressible Flows

This project provides a neural network-based interaction force model for gas-solid flows from low to intermediate Reynolds numbers and concentration, which can be linked to MFiX-DEM. We have constructed a database of the interaction force between the irregular-shaped particles using a spherical harmonic method and the fluid phase based on the particle-resolved direct numerical simulation (PR-DNS) with immersed boundary-based gas kinetic scheme. Unsupervised learning method, i.e., variational auto-encoder (VAE) has been applied to extract the primitive shape factors determining the drag force, lifting forces, and torque. The interaction force model has been trained and validated with a simple but effective multi-layer feed-forward neural network: multi-layer perceptron (MLP), which will be concatenated after the encoder of the previously trained VAE for geometry feature extraction for single, irregular particles. We have trained transpose convolutional neural networks with the PR-DNS data to predict the velocity and pressure gradient of the single particle systems and utilized them to calculate drag force of multi-particle systems. This model can provide high computational efficiency because it does not require collecting multiparticle system data from PR-DNS.

99 GENERAL AND MISCELLANEOUS↗

Spread Spectrum Time Domain Reflectometry (SSTDR) and Frequency Domain Reflectometry (FDR) for Detection of Cable Anomalies Using Machine Learning

Cables are initially qualified for nuclear power plant use for 40 years. As plants extend their operating license to 60 and 80 years, continued use of these cables must shift to a performance-based approach since it is cost prohibitive to completely replace cables that are likely still capable of performing their design function. A variety of cable tests are available and are commonly applied during outages when the cables can be taken out of service. Frequency domain reflectometry (FDR) is one of these test methods that is being more broadly accepted and used because it not only detects anomalies along the cable with a low-voltage signal that does not stress the cable insulation, but the technique also locates the anomalies. This supports follow-up local inspection and local repair or partial replacement of a damaged cable segment. Currently, FDR testing is only applied to cables that are taken out of service since the test instrument would be damaged by operational voltages. A related technology that has found some acceptance in the aircraft and rail industry is spread spectrum time domain reflectometry (SSTDR). This technology has been implemented with a custom commercial instrument by LiveWire Innovation Inc. that is designed to operate on live cables up to 1000 volts. One of the main conclusions of a previous effort was that cable reflectometry plots can be difficult for humans to analyze due to baseline noise, low or noisy anomaly response peaks, or large responses from cable ends. Detection of cable anomalies for many of these frequencies and test conditions was challenging for manual analysis. This presented an ideal opportunity for ML analysis to distinguish undamaged cable indications from anomalous cable indications. This research discusses application of machine learning (ML) to reflectometry cable test methods. The goal was to assess feasibility to distinguish undamaged cable reflectometry responses from damaged or anomalous cable reflectometry responses. The assessment considered the 3 instruments, multiple frequency bandwidths from each instrument, multiple cable anomalies and test conditions, and both supervised and unsupervised ML approaches. Although approaches and analysis methods were not identical or directly comparable, both outputs were encouraging. The unsupervised prediction weighted accuracy was assessed by instrument and by frequency. It performed better at high frequencies with the highest prediction accuracy of 0.84 for the higher frequency FDR, 0.79 for the 48-MHz LiveWire SSTDR, and 0.77 for 300-MHz PNNL SSTDR. The initial weighted accuracy average across all frequencies for using supervised ML was 0.56 to 0.68. The supervised analysis was repeated with noisier training data removed resulting in weighted accuracies of 0.69 to 0.87. These weighted accuracies are not directly comparable due to differences in the supervised and unsupervised analysis details but do indicate an encouraging trend. Even with limited and unbalanced data, strong prediction accuracies seem encouraging for further work including more data under a wider range of conditions.

42 ENGINEERING↗

Evaluating lightweight unsupervised online IDS for masquerade attacks in CAN

Vehicular controller area networks (CANs) are susceptible to masquerade attacks by malicious adversaries. In masquerade attacks, adversaries silence a targeted ID and then send malicious frames with forged content at the expected timing of benign frames. As masquerade attacks could seriously harm vehicle functionality and are the stealthiest attacks to detect in CAN, recent work has devoted attention to compare frameworks for detecting masquerade attacks in CAN. However, most existing works report offline evaluations using CAN logs already collected using simulations that do not comply with the domain’s real-time constraints. Here we contribute to advance the state of the art by presenting a comparative evaluation of four different non-deep learning (DL)-based unsupervised online intrusion detection systems (IDS) for masquerade attacks in CAN. Our approach differs from existing comparative evaluations in that we analyze the effect of controlling streaming data conditions in a sliding window setting. In doing so, we use realistic masquerade attacks being replayed from the ROAD dataset. We show that although evaluated IDS are not effective at detecting every attack type, the method that relies on detecting changes in the hierarchical structure of clusters of time series produces the best results at the expense of higher computational overhead. We discuss limitations, open challenges, and how the evaluated methods can be used for practical unsupervised online CAN IDS for masquerade attacks.

Anomaly detection↗

Combustion machine learning: Principles, progress and prospects

Progress in combustion science and engineering has led to the generation of large amounts of data from large-scale simulations, high-resolution experiments, and sensors. This corpus of data offers enormous opportunities for extracting new knowledge and insights—if harnessed effectively. Machine learning (ML) techniques have demonstrated remarkable success in data analytics, thus offering a new paradigm for data-intense analyses and scientific investigations through combustion machine learning (CombML). While data-driven methods are utilized in various combustion areas, recent advances in algorithmic developments, the accessibility of open-source software libraries, the availability of computational resources, and the abundance of data have together rendered ML techniques ubiquitous in scientific analysis and engineering. This article examines ML techniques for applications in combustion science and engineering. Starting with a review of sources of data, data-driven techniques, and concepts, we examine supervised, unsupervised, and semi-supervised ML methods. Various combustion examples are considered to illustrate and to evaluate these methods. Next, we review past and recent applications of ML approaches to problems in combustion, spanning fundamental combustion investigations, propulsion and energy-conversion systems, and fire and explosion hazards. Challenges unique to CombML are discussed and further opportunities are identified, focusing on interpretability, uncertainty quantification, robustness, consistency, creation and curation of benchmark data, and the augmentation of ML methods with prior combustion-domain knowledge.

33 ADVANCED PROPULSION SYSTEMS↗

Analysis of multispectral data using an unsupervised classification technique: Application to VAS

A statistical classification method based on clustering of multidimensional histograms was applied to several channels of the VAS multispectral imagery. The method automatically discriminates and classifies atmospheric ground features such as cloud types, atmospheric moisture patterns, ocean, or ground. Such a clustering method has the advantage of forming natural data groupings, without a priori classification. Clusters are not limited by straight lines or plane surfaces as is the case in threshold methods. The method was applied to simultaneous full resolution images from channels 8 (11.2 micron), 10 (6.7 micron), and 12 (3.9 micron). Twenty image segments of 64 by 64, 12 image segments of 128 by 128, and 4 image segments of 254 by 254 picture elements were analyzed. In addition, normal VISSR mode images at 1800, 1830, and 2000 GMT were used to identify the classes. The gray levels measured along a scan line and the result of the classification scheme (dashed curves) for the three channels investigated are shown. Each point of the image is affected to a class. Each class is identified by a center of gravity that is represented by a vector in the three dimensional space of gray levels.

Szejwach, G.↗

Notes for the improvement of the spatial and spectral data classification method

This report examines the spatial and spectral clustering technique for the unsupervised automatic classification and mapping of earth resources satellite data, and makes theoretical analysis of the decision rules and tests in order to suggest how the method might best be applied to other flight data such as Skylab and Spacelab.

Dalton, C. C.↗

Physics constrained unsupervised deep learning for rapid, high resolution scanning coherent diffraction reconstruction

By circumventing the resolution limitations of optics, coherent diffractive imaging (CDI) and ptychography are making their way into scientific fields ranging from X-ray imaging to astronomy. Yet, the need for time consuming iterative phase recovery hampers real-time imaging. While supervised deep learning strategies have increased reconstruction speed, they sacrifice image quality. Furthermore, these methods’ demand for extensive labeled training data is experimentally burdensome. Here, we propose an unsupervised physics-informed neural network reconstruction method, PtychoPINN, that retains the factor of 100-to-1000 speedup of deep learning-based reconstruction while improving reconstruction quality by combining the diffraction forward map with real-space constraints from overlapping measurements. In particular, PtychoPINN gains a factor of 4 in linear resolution and an 8 dB improvement in PSNR while also accruing improvements in generalizability and robustness. This blend of performance and computational efficiency offers exciting prospects for high-resolution real-time imaging in high-throughput environments such as X-ray free electron lasers (XFELs) and diffraction-limited light sources.

97 MATHEMATICS AND COMPUTING↗

sciCAN: single-cell chromatin accessibility and gene expression data integration via cycle-consistent adversarial network

The boom in single-cell technologies has brought a surge of high dimensional data that come from different sources and represent cellular systems from different views. With advances in these single-cell technologies, integrating single-cell data across modalities arises as a new computational challenge. Here, we present an adversarial approach, sciCAN, to integrate single-cell chromatin accessibility and gene expression data in an unsupervised manner. We benchmarked sciCAN with 5 existing methods in 5 scATAC-seq/scRNA-seq datasets, and we demonstrated that our method dealt with data integration with consistent performance across datasets and better balance of mutual transferring between modalities than the other 5 existing methods. We further applied sciCAN to 10X Multiome data and confirmed that the integrated representation preserves biological relationships within the hematopoietic hierarchy. Finally, we investigated CRISPR-perturbed single-cell K562 ATAC-seq and RNA-seq data to identify cells with related responses to different perturbations in these different modalities.

59 BASIC BIOLOGICAL SCIENCES↗

Prioritizing Scientific Data for Transmission

A software system has been developed for prioritizing newly acquired geological data onboard a planetary rover. The system has been designed to enable efficient use of limited communication resources by transmitting the data likely to have the most scientific value. This software operates onboard a rover by analyzing collected data, identifying potential scientific targets, and then using that information to prioritize data for transmission to Earth. Currently, the system is focused on the analysis of acquired images, although the general techniques are applicable to a wide range of data modalities. Image prioritization is performed using two main steps. In the first step, the software detects features of interest from each image. In its current application, the system is focused on visual properties of rocks. Thus, rocks are located in each image and rock properties, such as shape, texture, and albedo, are extracted from the identified rocks. In the second step, the features extracted from a group of images are used to prioritize the images using three different methods: (1) identification of key target signature (finding specific rock features the scientist has identified as important), (2) novelty detection (finding rocks we haven t seen before), and (3) representative rock sampling (finding the most average sample of each rock type). These methods use techniques such as K-means unsupervised clustering and a discrimination-based kernel classifier to rank images based on their interest level.

Castano, Rebecca↗

Application of unsupervised deep learning to image segmentation and in-situ contact angle measurements in a CO 2 -water-rock system

Rock surface wettability is a critical property that regulates multiphase flows in porous media, which can be quantified using the surface contact angle (CA). X-ray micro-computed tomography (μCT) provides an effective approach to in-situ measurements of surface CAs. However, the CA measurement accuracy depends significantly on the quality of CT image segmentation, which is the clustering of CT pixels into separate phases. Inspired by this, we developed a deep learning (DL)-based CA measurement workflow. Motivated by the recent tremendous progress in unsupervised learning techniques and aiming to avoid expensive manual data annotations, an unsupervised DL pipeline for CT image segmentation was proposed and implemented, which includes unsupervised model training and post-processing. The unsupervised model training was driven by a novel loss function constrained with feature similarity and spatial continuity and implemented by iterative forward and backward paths; the former clustered the pixel-wise feature vectors extracted by convolution neural networks, whereas the latter updated the parameters using gradient descent. An over-segmentation strategy was adopted for model training. The post-processing steps based on agglomerative hierarchical clustering (AHC) were implemented to further merge the over-segmented model output to the desired cluster number, which is intended to improve the efficiency of image segmentation. The developed unsupervised DL pipeline was compared with other commonly-used image segmentation methods using pixel-wise and physics-based evaluation metrics on a synthetic raw-image dataset, which had a known ground truth. The unsupervised DL pipeline showed the best performance. Next, the segmented images were input to an automatic CA measurement tool, and the results were validated by comparisons with manual measurements. The CA values from the manual and automatic measurements showed similar distributions and statistical properties. The automatic measurement demonstrated a wider spectrum because of the much larger number of measurement data points. The primary novelty of the unsupervised DL pipeline developed in this study lies in the novel loss function and the over-segmentation strategy associated with AHC post-processing. Finally, the workflow has been proven an efficient tool for pore-scale wettability characterization, which has a wide range of applications in fundamental studies of multiphase flows in natural porous media, which have critical implications to geological carbon sequestration, hydrocarbon energy recovery, and contaminant transport in groundwater.

42 ENGINEERING↗

Unsupervised Spatio-Temporal Data Mining Framework for Burned Area Mapping

A method reduces processing time required to identify locations burned by fire by receiving a feature value for each pixel in an image, each pixel representing a sub-area of a location. Pixels are then grouped based on similarities of the feature values to form candidate burn events. For each candidate burn event, a probability that the candidate burn event is a true burn event is determined based on at least one further feature value for each pixel in the candidate burn event. Candidate burn events that have a probability below a threshold are removed from further consideration as burn events to produce a set of remaining candidate burn events.

Boriah, Shyam↗

FIB-ToF-SIMS characterization of irradiated U-10Zr

Post-irradiation examination (PIE) is critical for the performance assessment and qualification of nuclear fuels. Secondary ion mass spectrometry (SIMS) is a powerful materials characterization technique that allows for elemental and isotopic mapping with a depth resolution greater than EDS and EPMA. However, it has not yet been applied to PIE of metallic nuclear fuel. Here, in this work, we characterize an fast neutron spectrum irradiated U-10Zr fuel sample using a time-of-flight SIMS (ToF-SIMS) system connected to a FIB/SEM system, which allows for flexible sample analysis compared to a dedicated ToF-SIMS instrument. Analysis of the resulting hyperspectral micrograph data was aided by the development of an unsupervised machine learning (ML) algorithm that iterates on existing methods to segment the 3D micrographic datasets based on the similarity of mass spectra. The results showed that the FIB-ToF-SIMS instrument was potentially capable of spatially resolving closed fission gas bubbles in 3D by continued ion sputtering of the analyzed volume. Additionally, the ML algorithm proved useful in revealing the chemical segregation of light fission products (those with an atomic mass between approximately 85–105 amu, such as ruthenium and rhodium) plus matrix zirconium, heavy fission products (those with an atomic mass between approximately 135–150 amu, such as the lanthanides) and uranium. Future studies are planned to conduct FIB-ToF-SIMS analysis on more irradiated U-Zr samples to study the constituent redistribution.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Optimal Electrification Using Renewable Energies: Microgrid Installation Model with Combined Mixture k-Means Clustering Algorithm, Mixed Integer Linear Programming, and Onsset Method

Optimal planning and design of microgrids are priorities in the electrification of off-grid areas. Indeed, in one of the Sustainable Development Goals (SDG 7), the UN recommends universal access to electricity for all at the lowest cost. Several optimization methods with different strategies have been proposed in the literature as ways to achieve this goal. This paper proposes a microgrid installation and planning model based on a combination of several techniques. The programming language Python 3.10 was used in conjunction with machine learning techniques such as unsupervised learning based on K-means clustering and deterministic optimization methods based on mixed linear programming. These methods were complemented by the open-source spatial method for optimal electrification planning: onsset. Four levels of study were carried out. The first level consisted of simulating the model obtained with a cluster, which is considered based on the elbow and k-means clustering method as a case study. The second level involved sizing the microgrid with a capacity of 40 kW and optimizing all the resources available on site. The example of the different resources in the Togo case was considered. At the third level, the work consisted of proposing an optimal connection model for the microgrid based on voltage stability constraints and considering, above all, the capacity limit of the source substation. Finally, the fourth level involved a planning study of electrification strategies based mainly on microgrids according to the study scenario. The results of the first level of study enabled us to obtain an optimal location for the centroid of the cluster under consideration, according to the different load positions of this cluster. Then, the results of the second level of study were used to highlight the optimal resources obtained and proposed by the optimization model formulated based on the various technology costs, such as investment, maintenance, and operating costs, which were based on the technical limits of the various technologies. In these results, solar systems account for 80% of the maximum load considered, compared to 7.5% for wind systems and 12.5% for battery systems. Next, an optimal microgrid connection model was proposed based on the constraints of a voltage stability limit estimated to be 10% of the maximum voltage drop. The results obtained for the third level of study enabled us to present selective results for load nodes in relation to the source station node. Finally, the last results made it possible to plan electrification using different network technologies and systems in the short and long term. The case study of Togo was taken into account. The various results obtained from the different techniques provide the necessary leads for a feasibility study for optimal electrification of off-grid areas using microgrid systems.

24 POWER TRANSMISSION AND DISTRIBUTION↗