Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Unsupervised learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Physics-guided Machine Learning: from Supervised Deep Networks to Unsupervised Lightweight Models [Slides]

Machine learning yields great potentials in improving imaging performance (i.e., accuracy and efficiency). The incorporation of governing equation will improve generalization and alleviate label scarcity. Employ physical properties can reduce model complexity and significantly save training cost without compromising accuracy. Combining SciML imaging and edge computing would allow broader applications in energy, medicine, and other domains.

97 MATHEMATICS AND COMPUTING↗

Machine-learned interatomic potentials by active learning: amorphous and liquid hafnium dioxide

We propose an active learning scheme for automatically sampling a minimum number of uncorrelated configurations for fitting the Gaussian Approximation Potential (GAP). Our active learning scheme consists of an unsupervised machine learning (ML) scheme coupled with a Bayesian optimization technique that evaluates the GAP model. We apply this scheme to a Hafnium dioxide (HfO 2 ) dataset generated from a "melt-quench" ab initio molecular dynamics (AIMD) protocol. Our results show that the active learning scheme, with no prior knowledge of the dataset, is able to extract a configuration that reaches the required energy fit tolerance. Further, molecular dynamics (MD) simulations performed using this active learned GAP model on 6144 atom systems of amorphous and liquid state elucidate the structural properties of HfO 2 with near ab initio precision and quench rates (i.e., 1.0 K/ps) not accessible via AIMD. The melt and amorphous X-ray structural factors generated from our simulation are in good agreement with experiment. In addition, the calculated diffusion constants are in good agreement with previous ab initio studies.

36 MATERIALS SCIENCE↗

XRF-ROI Finder: Machine Learning to Guide Region-of-Interest Scanning for X-ray Fluorescence Microscopy

The ROI-finder software is being developed for use by several Microscopy Group beamlines at Argonne National Laboratory, including 2-ID microprobes and 9-ID-B Bionanoprobe which use multi-scale scanning fluorescence microscopy to acquire elemental maps (multi-modal image data). Microscopy experiments require scan of samples at a coarse resolution followed by ROI identification using feature detection based on domain expertise. Finer resolution scans are then conducted based on identified ROI. The decision-making process based on domain expertise will be difficult to perform for faster data rates and much larger sampling volumes anticipated after APS-U necessitating the need for the ROI-finder software. The ROI- finder detects regions of interest through a continuous learning process, starting with a unsupervised representation learning and improving its recommendations through supervised learning and an interactive tool for user annotation. The scope of ongoing development efforts includes the integration of image registration module to correlate optical and X-ray images, extraction of feature morphology as well as elemental signatures in the image space and incorporation of beamtime streaming data by the scanning probe via EPICS.

CHOWDHURY, M. ARSHAD ZAHANGIR↗

Uncovering acoustic signatures of pore formation in laser powder bed fusion

Abstract We present a machine learning workflow to discover signatures in acoustic measurements that can be utilized to create a low-dimensional model to accurately predict the location of keyhole pores formed during additive manufacturing processes. Acoustic measurements were sampled at 100 kHz during single-layer laser powder bed fusion (LPBF) experiments, and spatio-temporal registration of pore locations was obtained from post-build radiography. Power spectral density (PSD) estimates of the acoustic data were then decomposed using non-negative matrix factorization with custom $$\varvec{k}$$ k -means clustering (NMF $$\varvec{k}$$ k ) to learn the underlying spectral patterns associated with pore formation. NMF $$\varvec{k}$$ k returned a library of basis signals and matching coefficients to blindly construct a feature space based on the PSD estimates in an optimized fashion. Moreover, the NMF $$\varvec{k}$$ k decomposition led to the development of computationally inexpensive machine learning models which are capable of quickly and accurately identifying pore formation with classification accuracy of supervised and unsupervised label learning greater than 95% and 90%, respectively. The intrinsic data compression of NMF k , the relatively light computational cost of the machine learning workflow, and the high classification accuracy makes the proposed workflow an attractive candidate for edge computing toward in-situ keyhole pore prediction in LPBF.

36 MATERIALS SCIENCE↗

Entropy and Boundary Based Adversarial Learning for Large Scale Unsupervised Domain Adaptation

Supervised semantic segmentation methods provide state-of-the-art performance, but their performance is limited by the amount of quality labeled data they need for training. Scarcity of labeled data and non-transferablity of models, due to cross-domain discrepancy makes it a bigger challenge for remote sensing imagery analysis. In this work, we approach this problem through adversarial learning, driven by entropy and boundary of region-of-interest for unsupervised domain adaptation. This concept helps with better boundary prediction and encourages target domain entropy maps (probability/uncertainty maps) to be similar to source domains. In particular, we showed that deriving informative entropy through the adversarial learning is essential to enable the adaptation. We used a large scale cross country building extraction dataset to validate the framework. The experimental results show the usefulness of considering boundary and entropy driven adversarial learning for adaptation.

Makkar, Nikhil↗

Unsupervised Clustering and Supervised Regression Learning to Select High Temperature Oxidation-Resistant Materials

High temperature oxidation and corrosion degradation mechanisms dictate the lifetime of materials critical to energy production. The combination of modeling and experimental approaches such as machine learning (ML) and data analytics, with sufficient experimental data, can accelerate the development of new materials while limiting its cost. In the present work, ML will be applied to two high temperature oxidation data libraries (Oak Ridge National Laboratory and National Air and Space Administration) that comprised of about 5000 mass change sample datasheets for a variety of materials and temperatures in dry air and air + 10 % H2O. A python code was developed to prepare the data for machine learning by collecting and formatting oxidation rate constants, alloy compositions and environment of exposure into a single data frame. Scikit-learn library and Statistics and Machine Learning Toolbox within MathWorks were then used to perform unsupervised clustering and supervised regression learning. The impact of dataset distribution on the performance of the developed ML models was evaluated. Potential strategies to improve the predictions and enhance extrapolative capability of the previously trained model were investigated.

Romedenne, Marie [ORNL] (ORCID:0000000317936561)↗

From clutter to clarity: Emergent neural operators via questionnaire metrics

Real-world datasets in chemical engineering and bioengineering processes—such as those from catalytic reactors, multiphase flows, polymerization reactors, bioreactors, and clinical trials—can often be unlabeled or disorganized, rendering the training of existing supervised learning models ineffective at learning the underlying dynamics. To salvage these datasets for decision-making, we first seek to obtain clarity from the cluttered data. Here, we present a framework for developing “structural” generative models, discovering emergent equations, and constructing efficient emulators from scrambled datasets by integrating unsupervised organizational learning techniques (Questionnaires) with advanced deep learning architectures (Deep Hidden Physics Models and Deep Operator Networks). Our approach is demonstrated on two illustrative model systems: (a) a 1D advection–diffusion partial differential equation representing a winding underground pipe and (b) an ensemble of Stuart–Landau oscillators, an agent-based system of coupled ordinary differential equations. In both cases, we successfully reconstruct meaningful spatial, temporal, and parameter embeddings from scrambled data, enabling good predictions of system dynamics. As a result, we highlight the framework’s potential for broader applications, enabling data-driven system identification in fields with inherently disorganized or hidden parameter spaces.

42 ENGINEERING↗

Identifying Climate Patterns Using Clustering Autoencoder Techniques

Abstract The complexity of growing spatiotemporal resolution of climate simulations produces a variety of climate patterns under different projection scenarios. This paper proposes a new data-driven climate classification workflow via an unsupervised deep learning technique that can dimensionally reduce the vast volume of spatiotemporal numerical climate projection data into a compact representation. We aim to identify distinct zones that capture multiple climate variables as well as their future changes under different climate change scenarios. Our approach leverages convolutional autoencoders combined with k -means clustering (standard autoencoder) and online clustering based on the Sinkhorn–Knopp algorithm (clustering autoencoder) across the conterminous United States (CONUS) to capture unique climate patterns in a data-driven fashion from the Geophysical Fluid Dynamics Laboratory Earth System Model with GOLD component (GFDL-ESM2G). The developed approach compresses 70 years of GFDL-ESM2G simulation at 0.125° spatial resolution across the CONUS under multiple warming scenarios to a lower-dimensional space by a factor of 660 000 and then tested on 150 years of GFDL-ESM2G simulation data. The results show that five climate clusters capture physically reasonable and spatially stable climatological patterns matched to known climate classes defined by human experts. Results also show that using a clustering autoencoder can reduce the computational time for clustering by up to 9.2 times when compared to using a standard autoencoder. Our five unique climate patterns resulting from the deep learning–based clustering of the lower-dimensional space thereby enable us to provide insights on hydrometeorology and its spatial heterogeneity across the conterminous United States immediately without downloading large climate datasets. Significance Statement This paper presents a data-driven climate classification approach using unsupervised deep learning to dimensionally reduce climate model outputs and to identify distinct climate regions for their future changes. Our approach compresses climate information for 70 years of Geophysical Fluid Dynamics Laboratory Earth System Model data across the conterminous United States (CONUS) at 0.125° spatial resolution. The results reveal that five climate clusters capture reasonable and stable climatological patterns matched to known climate patterns. The embedded clustering process in deep learning provides ×9.2 times faster execution than the k -means clustering technique. These results give us insight about climate spatial patterns and heterogeneity of hydrological patterns across the conterminous United States without downloading large climate datasets.

Kurihana, Takuya↗

Leading-Order Analysis by Artificial Intelligence [Slides]

The following topics are addressed in this seminar presentation: The author's background; What is leading-order analysis?; What are supervised and unsupervised machine learning and what is artificial intelligence?; The definition of AI; and, Conclusions and outlook.

42 ENGINEERING↗

High temperature oxidation of corrosion resistant alloys from machine learning

Parabolic rate constants, k p , were collected from published reports and calculated from corrosion product data (sample mass gain or corrosion product thickness) and tabulated for 75 alloys exposed to temperatures between ~800 and 2000 K (~500–1700 °C; 900–3000°F). Data were collected for environments including lab air, ambient and supercritical carbon dioxide, supercritical water, and steam. Materials studied include low- and high-Cr ferritic and austenitic steels, nickel superalloys, and aluminide materials. A combination of Arrhenius analysis, simple linear regression, supervised and unsupervised machine learning methods were used to investigate the relations between composition and oxidation kinetics. The supervised machine learning techniques produced the lowest mean standard errors. The most significant elements controlling oxidation kinetics were Ni, Cr, Al, and Fe, with Mo and Co composition also found to be significant features. The activation energies produced from the machine learning analysis were in the correct distributions for the diffusion constants for the oxide scales expected to dominate in each class.

Materials Science↗

Fast Grain Mapping with Sub-Nanometer Resolution Using 4D-STEM with Grain Classification by Principal Component Analysis and Non-Negative Matrix Factorization

High-throughput grain mapping with sub-nanometer spatial resolution is demonstrated using scanning nanobeam electron diffraction (also known as 4D scanning transmission electron microscopy, or 4D-STEM) combined with high-speed direct-electron detection. An electron probe size down to 0.5 nm in diameter is used and the sample investigated is a gold–palladium nanoparticle catalyst. Computational analysis of the 4D-STEM data sets is performed using a disk registration algorithm to identify the diffraction peaks followed by feature learning to map the individual grains. Two unsupervised feature learning techniques are compared: principal component analysis (PCA) and non-negative matrix factorization (NNMF). The characteristics of the PCA versus NNMF output are compared and the potential of the 4D-STEM approach for statistical analysis of grain orientations at high spatial resolution is discussed.

47 OTHER INSTRUMENTATION↗

Automatic Point Cloud Building Envelope Segmentation (AutoCuBES)

The Auto-CuBES algorithm is based on unsupervised machine learning that automatically labels 3D point cloud data and reduces the time spent in manual segmentation. The algorithm can process high-resolution point clouds and generate a wire-frame building envelope model with a small set of calibration parameters. The algorithm inputs a 3D point cloud generated by commonly available surveying equipment and outputs a wire-frame model of the building envelope. Unsupervised machine learning methods were used to identify facades, windows, and doors while minimizing the number of calibration parameters.

Puente, BryanMaldonado↗

Condition monitoring and anomaly detection in cyber-physical systems

The modern industrial environment is equipping myriads of smart manufacturing machines where the state of each device can be monitored continuously. Such monitoring can help identify possible future failures and develop a cost-effective maintenance plan. However, it is a daunting task to perform early detection with low false positives and negatives from the huge volume of collected data. This requires developing a holistic machine learning framework to address the issues in condition monitoring of high priority components and develop efficient techniques to detect anomalies that can detect and possibly localize the faulty components. This paper presents a comparative analysis of recent machine learning approaches for robust, cost-effective anomaly detection in cyber-physical systems. While detection has been extensively studied, very few researchers have analyzed the localization of the anomalies. We show that supervised learning outperforms unsupervised algorithms. For supervised cases, we achieve near-perfect accuracy of 98% (specifically for tree-based algorithms). In contrast, the best-case accuracy in the unsupervised cases was 63%—the area under the receiver operating characteristic curve (AUC) exhibits similar outcomes as an additional metric.

Marfo, William↗

OmicsMLMentor: A Web Application for Guided Machine Learning Analysis of Omics Data

Expression-based omics technologies (e.g. proteomics, metabolomics, transcriptomics, etc.) increasingly rely on supervised and unsupervised machine learning (ML) models to find key biomolecules distinguishing conditions, identify natural groupings in biological data, or generate predictions for outcomes of interest. Fitting ML models to omics data presents several challenges, including handling missing data, selecting a normalization method, choosing a valid model, and optimizing hyperparameters, all requiring statistical programming skills to address these challenges. Thus, the open-source web application SLOPE was designed to lower the barrier to ML modeling for omics data. SLOPE supports the fitting of 15 ML models (10 supervised and 5 unsupervised) tailored to omics datasets, such as proteomics, metabolomics, lipidomics, and transcriptomics. SLOPE offers several omics-specific features, including methods for handling missingness (imputation, conversion, removal), normalization tests, ranking of models based on the structure of a user’s data and user input, and optimal hyperparameter selections using cross-validation splits. By streamlining ML workflows for omics analysis, SLOPE address critical gaps in existing online web tools, facilitating a broader adoption of these models for omics research. Here, SLOPE is applied to data from a lignin exposure study to highlight the workflow for fitting both supervised and unsupervised models to data.

lipidomics↗

Source-agnostic gravitational-wave detection with recurrent autoencoders

Abstract We present an application of anomaly detection techniques based on deep recurrent autoencoders (AEs) to the problem of detecting gravitational wave (GW) signals in laser interferometers. Trained on noise data, this class of algorithms could detect signals using an unsupervised strategy, i.e. without targeting a specific kind of source. We develop a custom architecture to analyze the data from two interferometers. We compare the obtained performance to that obtained with other AE architectures and with a convolutional classifier. The unsupervised nature of the proposed strategy comes with a cost in terms of accuracy, when compared to more traditional supervised techniques. On the other hand, there is a qualitative gain in generalizing the experimental sensitivity beyond the ensemble of pre-computed signal templates. The recurrent AE outperforms other AEs based on different architectures. The class of recurrent AEs presented in this paper could complement the search strategy employed for GW detection and extend the discovery reach of the ongoing detection campaigns.

47 OTHER INSTRUMENTATION↗

Learning Stochastic Parametric Differentiable Predictive Control Policies

We present a scalable unsupervised learning-based method for obtaining explicit control policies for model predictive control problems for stochastic linear systems with additive uncertainties subject to nonlinear chance constraints. We call the proposed method stochastic parametric differentiable predictive control (SP-DPC), which extends the recently proposed deterministic DPC policy optimization algorithm. We formulate the SP-DPC as a deterministic approximation to the stochastic parametric constrained optimal control problem via independent sampling of the problem's parameters and uncertainties. This formulation allows us to directly compute the policy gradients via automatic differentiation of the problem's value function, evaluated over sampled parameters and uncertainties. In particular, the computed expectation of the problem's value function is backpropagated through the finite-time closed-loop system rollouts parametrized by a known nominal system dynamics model and neural control policy. We also provide theoretical probabilistic guarantees on closed-loop stability and chance constraints satisfaction for systems controlled by learned neural policies. We demonstrate the computational efficiency and scalability of the proposed policy optimization algorithm in three numerical examples, including systems with a large number of states or subject to nonlinear constraints.

Drgona, Jan↗

Unsupervised Process Anomaly Detection and Identification Using the Leave-One-Variable-Out Approach

Automated anomaly detection and identification can signal equipment issues and pinpoint causes in large-scale industrial systems. For systems with limited failure history, unsupervised machine learning methods can be utilized as they do not require past failures. This study introduces the leave-one-variable-out (LOVO) model, which masks one variable at a time to predict the others, learning underlying process correlations. Detection performance was assessed with synthetic and experimental data, while identification performance used only synthetic data due to its ability to generate labeled anomaly types. For detection using synthetic data, the LOVO model generally outperformed comparative models; while using experimental data, the comparative methods outperformed the LOVO model. However, the comparative methods required selecting a latent size, and these conclusions pertain to using the optimal size. In practice, it would not be feasible to always select the optimal value, and incorrect selections impacted performance. In contrast, the LOVO model does not require a latent space. For identification using synthetic data, the LOVO model was slightly outperformed in interpretability and repeatability but still demonstrated impressive results. These outcomes suggest that the LOVO model is an effective model and may be more easily implemented without the challenging tuning process of selecting a latent size.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Uranium Oxide Synthetic Pathway Discernment through Unsupervised Morphological Analysis

We present a novel unsupervised machine learning method for quantitative representation of scanning electron micrographs and its applications and performance for nuclear forensic analysis of uranium ore concentrates. The method uses a vector quantizing variational autoencoder followed by a histogram operation to encode a micrograph into a single dimensional representation, called the latent vector. The method requires no extant labeling of the data and can be applied over large datasets of micrographs with minimal human interaction. The representations generated are broadly descriptive of each micrograph and the microstructure of the material imaged. In the case of uranium ore concentrate analysis, the representations were amenable to processing reagent and ore concentrate species classification with accuracy of 81:8%, which is competitive with state-of-the-art supervised networks. The representations were also used to classify previously unseen processing routes, were able to classify imaging parameters such as magnification (to 76:0% accuracy), were able to classify fine grained process parameters such as calcining temperature (to 74:4% accuracy), and their informatic properties indicate that they are generally descriptive of the image represented. This method can be applied across microstructure analysis fields to perform quantitative analysis without the need for labor intensive and possibly biased human analysis.

Scanning Electron Microscopy, Vector Quantizing Va↗