Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Unsupervised machine learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Uranium Oxide Synthetic Pathway Discernment through Unsupervised Morphological Analysis

We present a novel unsupervised machine learning method for quantitative representation of scanning electron micrographs and its applications and performance for nuclear forensic analysis of uranium ore concentrates. The method uses a vector quantizing variational autoencoder followed by a histogram operation to encode a micrograph into a single dimensional representation, called the latent vector. The method requires no extant labeling of the data and can be applied over large datasets of micrographs with minimal human interaction. The representations generated are broadly descriptive of each micrograph and the microstructure of the material imaged. In the case of uranium ore concentrate analysis, the representations were amenable to processing reagent and ore concentrate species classification with accuracy of 81:8%, which is competitive with state-of-the-art supervised networks. The representations were also used to classify previously unseen processing routes, were able to classify imaging parameters such as magnification (to 76:0% accuracy), were able to classify fine grained process parameters such as calcining temperature (to 74:4% accuracy), and their informatic properties indicate that they are generally descriptive of the image represented. This method can be applied across microstructure analysis fields to perform quantitative analysis without the need for labor intensive and possibly biased human analysis.

Scanning Electron Microscopy, Vector Quantizing Va↗

Machine learning to identify geologic factors associated with production in geothermal fields: A casestudy using 3D geologic data, Brady geothermal field, Nevada

In this paper, we present an analysis using unsupervised machine learning (ML) to identify the key geologic factors that contribute to the geothermal production in the Brady geothermal field. Brady is a hydrothermal system in northwestern Nevada that supports both electricity production and direct use of hydrothermal fluids. Transmissive fluid flow pathways are relatively rare in the subsurface but are critical components of hydrothermal systems like Brady and many other types of fluid flow systems in fractured rock. The ML method, non-negative matrix factorization with k-means clustering (NMFk), is applied to a library of fourteen 3D geologic characteristics hypothesized to control hydrothermal circulation in the Brady geothermal field. Our results indicate the macro-scale faults and a local step-over in the fault system preferentially occur along with production wells when compared to injection wells and non-productive wells. We infer that these are the key geologic characteristics that control the through-going hydrothermal transmission pathways at Brady. Our results demonstrate 1) the specific geologic controls on the Brady hydrothermal system and 2) the efficacy of pairing ML techniques with 3D geologic characterization to enhance the understanding of subsurface processes.

15 GEOTHERMAL ENERGY↗

Machine Learning to Identify Geologic Factors Associated with Production in Geothermal Fields: A Case-Study Using 3D Geologic Data from Brady Geothermal Field and NMFk

In this paper, we present an analysis using unsupervised machine learning (ML) to identify the key geologic factors that contribute to the geothermal production in Brady geothermal field. Brady is a hydrothermal system in northwestern Nevada that supports both electricity production and direct use of hydrothermal fluids. Transmissive fuid-fow pathways are relatively rare in the subsurface, but are critical components of hydrothermal systems like Brady and many other types of fuid-fow systems in fractured rock. Here, we analyze geologic data with ML methods to unravel the local geologic controls on these pathways. The ML method, non-negative matrix factorization with k-means clustering (NMFk), is applied to a library of 14 3D geologic characteristics hypothesized to control hydrothermal circulation in the Brady geothermal field. Our results indicate that macro-scale faults and a local step-over in the fault system preferentially occur along production wells when compared to injection wells and non-productive wells. We infer that these are the key geologic characteristics that control the through-going hydrothermal transmission pathways at Brady. Our results demonstrate: (1) the specific geologic controls on the Brady hydrothermal system and (2) the efficacy of pairing ML techniques with 3D geologic characterization to enhance the understanding of subsurface processes. This submission includes the published journal article detailing this work, the published 3D geologic map of the Brady Geothermal Area used as a basis to develop structural and geological variables that are hypothesized to control or effect permeability or connectivity, 3D well data, along which geologic data were sampled for PCA analyses, and associated metadata file. This work was done using the GeoThermalCloud framework, which is part of SmartTensors (both are linked below).

15 GEOTHERMAL ENERGY↗

Machine Learning Reveals Memory of the Parent Phases in Ferroelectric Relaxors Ba(Ti$_{1-x}$,Zr x )O 3

Machine learning has been establishing its potential in multiple areas of condensed matter physics and materials science. Here, in this work, an unsupervised machine learning workflow is developed and used within a framework of first-principles-based atomistic simulations to investigate phases, phase transitions, and their structural origins in ferroelectric relaxors, Ba(Ti 1-x ,Zr x )O 3 . The applicability of the workflow is first demonstrated to identify phases and phase transitions in the parent compound, a prototypical ferroelectric BaTiO 3 . Then the workflow is applied for Ba(Ti 1-x ,Zrx)O 3 with x ≤ 0.25 to reveal i) that some of the compounds bear a subtle memory of BaTiO 3 phases beyond the point of the pinched phase transition, which could contribute to their enhanced electromechanical response; ii) the existence of peculiar phases with delocalized precursors of nanodomains—likely candidates for the controversial polar nanoregions; and iii) nanodomain phases for the largest concentrations of x.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Machine-learning predictions of the shale wells’ performance

The ultra-low permeability nature of shale reservoirs leads to an extended linear flow and necessitates horizontal wells with multi-stage engineered fractures to efficiently extract hydrocarbons resources. These artificially-generated and naturally-occurring fractures form complex networks that create complex flow regimes which control oil production. These fractures are neither identical nor equally-spaced, which leads to a production profile with a masked onset of the boundary-dominated flow. The combination of the extended linear flow with the indeterminate onset of the boundary-dominated flow challenges the current deterministic analytic approaches to forecast the estimated ultimate recovery (EUR). In this work, we propose a novel machine-learning approach which overcomes these challenges and provides reliable EUR estimates based on field-wide analyses. We implement a novel unsupervised machine learning (ML) methodology, which allows for automatic identification of the optimal number of features (signals) present in the data based on non-negative matrix/tensor factorization coupled with k-means clustering incorporating regularization and physics constraints. In the presented analyses, the input data to the ML algorithm is the available (public) production history from the field collected at existing unconventional reservoirs. We validate our approach through hindcasting of the production data, where we achieved an excellent agreement. In addition, our approach is able to identify the poorly-performing wells, which could benefit from early refracing. Our approach provides fast and accurate estimations of the well performance without presumptions about the state of the well or the flow regime.

03 NATURAL GAS↗

Machine learning to identify geologic factors associated with production in geothermal fields: a case-study using 3D geologic data, Brady geothermal field, Nevada

Abstract In this paper, we present an analysis using unsupervised machine learning (ML) to identify the key geologic factors that contribute to the geothermal production in Brady geothermal field. Brady is a hydrothermal system in northwestern Nevada that supports both electricity production and direct use of hydrothermal fluids. Transmissive fluid-flow pathways are relatively rare in the subsurface, but are critical components of hydrothermal systems like Brady and many other types of fluid-flow systems in fractured rock. Here, we analyze geologic data with ML methods to unravel the local geologic controls on these pathways. The ML method, non-negative matrix factorization with k -means clustering (NMF k ), is applied to a library of 14 3D geologic characteristics hypothesized to control hydrothermal circulation in the Brady geothermal field. Our results indicate that macro-scale faults and a local step-over in the fault system preferentially occur along production wells when compared to injection wells and non-productive wells. We infer that these are the key geologic characteristics that control the through-going hydrothermal transmission pathways at Brady. Our results demonstrate: (1) the specific geologic controls on the Brady hydrothermal system and (2) the efficacy of pairing ML techniques with 3D geologic characterization to enhance the understanding of subsurface processes.

58 GEOSCIENCES↗

Hydrogen Bonding Inside Anionic Polymeric Brush Layer: Machine Learning-Driven Exploration of the Relative Roles of the Polymer Steric Effect, Charging, and Type of Screening Counterions

This paper employs a combination of all-atom molecular dynamics (MD) simulations and unsupervised machine learning (ML) for studying the water-water hydrogen bonds (HBs) inside the anionic poly-acrylic acid (PAA) brushes modeled using all-atom MD simulations. PAA brush layer with different charge fraction (f), namely f=0, f=0.25, and f=1, is considered. Water-water interactions, both inside and outside the brush layer, are represented through distinct clusters of tupules of variables representing distances associated with the interacting water molecules. While clusters representing the HBs are present for water inside and outside the brushes, several clusters representing the long-range water-water interactions are missing for the water molecules inside the highly charged (f=1) PAA brushes. More importantly, inside highly charged brushes, the edge of the clusters representing the water-water HBs is progressively shortened, as compared to that in the bulk. Both these results stem from the presence of the PAA brushes imparting the steric effect and the charge effect, or the effect associated with enhanced interactions of water molecules with PE charges and counterions, thereby disrupting the water connectivity. This water-charged-species interaction also increases the water-water HB angle, i.e., makes the water-water HBs less stable inside the highly charged PAA brush layer. The narrowing of the clusters representing the HBs and the alteration of the angle characterizing the HBs confirm that the conditions defining the water-water HBs change inside the PAA brush layer as a function of the charges on the PAA brush layer. Furthermore, we show that the use of the generic definition of HBs, as compared to using our simulation-motivated modified definition of water-water HBs, overpredict the number of water-water HBs inside the PAA brush layer. Finally, we employ this all-atom-MD-ML framework to quantify the effect of other types of screening counterions (Li + , Ca 2+ , and Y 3+ ions) in determining the water-water interactions and water-water HB properties inside the PAA brush layer. Furthermore, the findings of the present study, confirming the weakening of water-water HBs inside the PAA brush layer, points to the possibility that the water molecules will be more available for hydrating the brush layer and counterions, thereby leading to a more pronounced wetting of the PAA brush layer.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Machine learning uncovers independently regulated modules in the Bacillus subtilis transcriptome

The transcriptional regulatory network (TRN) of Bacillus subtilis coordinates cellular functions of fundamental interest, including metabolism, biofilm formation, and sporulation. Here, we use unsupervised machine learning to modularize the transcriptome and quantitatively describe regulatory activity under diverse conditions, creating an unbiased summary of gene expression. We obtain 83 independently modulated gene sets that explain most of the variance in expression and demonstrate that 76% of them represent the effects of known regulators. The TRN structure and its condition-dependent activity uncover putative or recently discovered roles for at least five regulons, such as a relationship between histidine utilization and quorum sensing. The TRN also facilitates quantification of population-level sporulation states. As this TRN covers the majority of the transcriptome and concisely characterizes the global expression state, it could inform research on nearly every aspect of transcriptional regulation in B. subtilis.

59 BASIC BIOLOGICAL SCIENCES↗

Objective Phenotyping of Root System Architecture Using Image Augmentation and Machine Learning in Alfalfa (Medicago sativa L.)

Active breeding programs specifically for root system architecture (RSA) phenotypes remain rare; however, breeding for branch and taproot types in the perennial crop alfalfa is ongoing. Phenotyping in this and other crops for active RSA breeding has mostly used visual scoring of specific traits or subjective classification into different root types. While image-based methods have been developed, translation to applied breeding is limited. This research is aimed at developing and comparing image-based RSA phenotyping methods using machine and deep learning algorithms for objective classification of 617 root images from mature alfalfa plants collected from the field to support the ongoing breeding efforts. Our results show that unsupervised machine learning tends to incorrectly classify roots into a normal distribution with most lines predicted as the intermediate root type. Encouragingly, random forest and TensorFlow-based neural networks can classify the root types into branch-type, taproot-type, and an intermediate taproot-branch type with 86% accuracy. With image augmentation, the prediction accuracy was improved to 97%. Coupling the predicted root type with its prediction probability will give breeders a confidence level for better decisions to advance the best and exclude the worst lines from their breeding program. This machine and deep learning approach enables accurate classification of the RSA phenotypes for genomic breeding of climate-resilient alfalfa.

59 BASIC BIOLOGICAL SCIENCES↗

Machine learning technique to identify grains in polycrystalline materials samples

A method of identifying grains in polycrystalline materials, the method including (a) identifying local crystal structure of the polycrystalline material based on neighbor coordination or pattern recognition machine learning, the local crystal structure including grains and grain boundaries, (b) pre-processing the grains and the grain boundaries using image processing techniques, (c) conducting grain identification using unsupervised machine learning; and (d) refining a resolution of the grain boundaries.

Sankaranarayanan, Subramanian↗

General-Purpose Unsupervised Cyber Anomaly Detection via Non-Negative Tensor Factorization

Distinguishing malicious anomalous activities from unusual but benign activities is a fundamental challenge for cyber defenders. Prior studies have shown that statistical user behavior analysis yields accurate detections by learning behavior profiles from observed user activity. These unsupervised models are able to generalize to unseen types of attacks by detecting deviations from normal behavior, without knowledge of specific attack signatures. However, approaches proposed to date based on probabilistic matrix factorization are limited by the information conveyed in a two-dimensional space. Non-negative tensor factorization, on the other hand, is a powerful unsupervised machine learning method that naturally models multi-dimensional data, capturing complex and multi-faceted details of behavior profiles. Herein, our new unsupervised statistical anomaly detection methodology matches or surpasses state-of-the-art supervised learning baselines across several challenging and diverse cyber application areas, including detection of compromised user credentials, botnets, spam e-mails, and fraudulent credit card transactions.

97 MATHEMATICS AND COMPUTING↗

Defect detection in atomic-resolution images via unsupervised learning with translational invariance

Abstract Crystallographic defects can now be routinely imaged at atomic resolution with aberration-corrected scanning transmission electron microscopy (STEM) at high speed, with the potential for vast volumes of data to be acquired in relatively short times or through autonomous experiments that can continue over very long periods. Automatic detection and classification of defects in the STEM images are needed in order to handle the data in an efficient way. However, like many other tasks related to object detection and identification in artificial intelligence, it is challenging to detect and identify defects from STEM images. Furthermore, it is difficult to deal with crystal structures that have many atoms and low symmetries. Previous methods used for defect detection and classification were based on supervised learning, which requires human-labeled data. In this work, we develop an approach for defect detection with unsupervised machine learning based on a one-class support vector machine (OCSVM). We introduce two schemes of image segmentation and data preprocessing, both of which involve taking the Patterson function of each segment as inputs. We demonstrate that this method can be applied to various defects, such as point and line defects in 2D materials and twin boundaries in 3D nanocrystals.

36 MATERIALS SCIENCE↗

Identifying Different Classes of Seismic Noise Signals Using Unsupervised Learning

Abstract Proper classification of nontectonic seismic signals is critical for detecting microearthquakes and developing an improved understanding of ongoing weak ground motions. We use unsupervised machine learning to label five classes of nonstationary seismic noise common in continuous waveforms. Temporal and spectral features describing the data are clustered to identify separable types of emergent and impulsive waveforms. The trained clustering model is used to classify every 1 s of continuous seismic records from a dense seismic array with 10–30 m station spacing. We show that dominate noise signals can be highly localized and vary on length scales of hundreds of meters. The methodology demonstrates the complexity of weak ground motions and improves the standard of analyzing seismic waveforms with a low signal‐to‐noise ratio. Application of this technique will improve the ability to detect genuine microseismic events in noisy environments where seismic sensors record earthquake‐like signals originating from nontectonic sources.

Johnson, Christopher W.↗

Unraveling Hydrogen Induced Geochemical Reaction Mechanisms through Coupled Geochemical Modeling and Machine Learning

Underground hydrogen storage (UHS) provides a promising large-scale, long-term energy storage solution. A reasonable recovery of stored hydrogen is critical for a successful storage scheme. However, in subsurface reservoirs hydrogen is subject to active geochemical reactions that might result in hydrogen loss. In this study, we implemented a geochemical modeling approach coupled with an unsupervised machine learning technique called non-negative matrix factorization (NMF) to unravel the complex brine-rock-H 2 geochemical processes responsible for hydrogen losses, with particular focus on sulfate reduction reactions. NMF is applied to modeled mineral evolution and fluid component profiles to retrieve profiles that can be interpreted to more easily assess competing processes. NMF decouples simulated competing equilibrium reactions. This facilitates separation of overlapping reaction profiles from redox processes, dissolution fronts, and secondary precipitation while considering the effects of simulation parameters such as salinity, temperature, and total H 2 pressure. NMF successfully discriminates these competing effects in nonlinear ways, allowing robust interpretation. In addition, NMF reveals subtle coupled mineral associations and reaction fronts that are invisible to conventional model analysis. This integrated approach strengthens the conceptual understanding of complex nonlinear hydrogen-brine-rock interactions and advances geochemical research on UHS systems to resolve complexities in modeled geochemical systems without the need for direct experiments or prior knowledge. Furthermore, this study highlights the efficacy of combining geochemical modeling with machine learning techniques to enhance the interpretability of the intricate geochemical simulation output through deciphering the overlapping reaction path that cannot be achieved only using conventional analysis of geochemical models alone.

08 HYDROGEN↗

Spread spectrum time domain reflectometry (SSTDR) and frequency domain reflectometry (FDR) cable inspection using machine learning

Cables are initially qualified for nuclear power plant use for 40 years. As plants extend their operating license to 60 and 80 years, justification for continued cable use must shift to a condition-based approach since it is cost prohibitive to completely replace cables that are likely still capable of performing their design function. The Pacific Northwest National Laboratory (PNNL) Accelerated and Real Time Experimental Nodal Analysis (ARENA) cable motor test bed was used to test the response of a commercial spread spectrum time domain reflectometry (SSTDR) system, a laboratory instrument software-controlled SSTDR, and a vector network analyzer-based frequency domain reflectometry (FDR) system to various cable anomalies. The three instrument systems were able to interrogate cables over a range of frequency bandwidths that can be helpful for human data analysis. Data were subjected to supervised and unsupervised machine learning (ML) analyses to distinguish normal undamaged cable responses from anomalous cable responses. Both supervised and unsupervised ML approaches produced encouraging results with an undamaged/anomalous prediction accuracy from 0.69% to 0.87%. Recommendations for further development and field implementation include increased and more balanced sample sets particularly including more training data.

SSTDR, FDR, Reflectometry, Machine Learning, ARENA↗

ACAT

A Physics-Informed Machine Learning (PIML) framework for better system vulnerability assessment and faster corrective action recommendation. This framework consists of deriving physics-informed priors and smart sampling algorithms to help reduce the data samples of grid models and the dimension of simulation outputs, yielding small yet representative subset of the complex system, and both supervised and unsupervised machine learning (ML) algorithms for designing corrective actions

Chen, Yousu↗

Comparison of Supervised and Un-Supervised Machine Learning Algorithms for Threat Detection and Scintillator Performance for Radiation Portal Monitoring

Following the events of September 11, 2001, international border crossing have been equipped with radiation portal monitors (RPMs) to identify illicit radioactive material. Polyvinyl toluene (PVT) scintillators are commonly used due to their low cost and reasonable maintainability, however they offer low spectral resolution. Despite the fact that over twenty years has transpired since this event, radioisotopes are still typically identified by hand-crafted classification algorithms, e.g., total counts or energy windowing, and exhibit relatively poor performance in detecting threats at the low false alarm rates required to support the stream of commerce. While some improvement to performance has been realized via the use of supervised machine learning, these classification algorithms typically utilize simulations in lieu of real data due to the sparsity of data for one or more classes. Accordingly, the performance of these algorithms is somewhat less than optimal when examining experiments or simulations with model mismatch. Consequently, in this work, we examine the application of a number of unsupervised machine learning, anomaly detection based algorithms, to circumvent the inverse crime when analyzing spectroscopy data for RPMs. We also compare anomaly detection results with those obtained via the use of supervised classification detection ML algorithms when model mismatch is introduced between the simulated threat items utilized for training/testing. Finally, we compared the performance of the PVT scintillators to those obtained with higher resolution detectors using both anomaly detection and supervised classification algorithms.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

Machine Learning-Based Anomaly Detection for PMT Data Quality Monitoring in the SBN and DUNE

Maintaining high-quality detector data is essential for achieving the scientific objectives of the Short-Baseline Neutrino (SBN) Program at Fermilab. Current data quality monitoring (DQM) procedures rely primarily on threshold-based metrics and manual inspection of detector monitoring plots, making the detection of subtle or gradually developing anomalies both time-consuming and dependent on expert interpretation. This project developed and evaluated a machine-learning workflow for automatically identifying anomalous photomultiplier tube (PMT) channels in the Short-Baseline Near Detector (SBND) using optical-hit amplitude data. A Python-based analysis program was developed to process ROOT files, extract statistical features describing individual PMT amplitude distributions, and generate feature vectors for anomaly detection. These features were used to train an Isolation Forest model using data representing normal detector operation. The trained model was subsequently applied to independent detector runs to identify channels exhibiting statistically unusual behavior relative to the learned reference response. To support expert interpretation, the workflow generated complementary diagnostic products, including anomaly score distributions, normalized amplitude comparisons, decision-tree visualizations, and principal component analysis (PCA) projections. This project demonstrated the feasibility of integrating unsupervised machine learning into detector data-quality monitoring and developed a complete workflow for automated PMT performance assessment to aid expert-driven review. Beyond its technical contributions, the VFP appointment fostered a research collaboration between Aurora University and Fermilab and provided direct workforce development benefits by training the visiting faculty member in detector-scale machine-learning methods that are now being incorporated into undergraduate coursework and research. The methodology developed here provides a foundation for future applications to ProtoDUNE and other liquid argon time projection chamber (LArTPC) detectors, contributing to ongoing efforts to improve detector reliability, reduce manual monitoring requirements, and enable scalable data quality monitoring for future large-scale neutrino experiments, including the Deep Underground Neutrino Experiment (DUNE).

Colón Santana, Juan A. [Unlisted, US, IL]↗