Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “unsupervised method”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Automated Identification of Characteristic Droplet Size Distributions in Stratocumulus Clouds Utilizing a Data Clustering Algorithm

Abstract Droplet-level interactions in clouds are often parameterized by a modified gamma fitted to a “global” droplet size distribution. Do “local” droplet size distributions of relevance to microphysical processes look like these average distributions? This paper describes an algorithm to search and classify characteristic size distributions within a cloud. The approach combines hypothesis testing, specifically, the Kolmogorov–Smirnov (KS) test, and a widely used class of machine learning algorithms for identifying clusters of samples with similar properties: density-based spatial clustering of applications with noise (DBSCAN) is used as the specific example for illustration. The two-sample KS test does not presume any specific distribution, is parameter free, and avoids biases from binning. Importantly, the number of clusters is not an input parameter of the DBSCAN-type algorithms but is independently determined in an unsupervised fashion. As implemented, it works on an abstract space from the KS test results, and hence spatial correlation is not required for a cluster. The method is explored using data obtained from the Holographic Detector for Clouds (HOLODEC) deployed during the Aerosol and Cloud Experiments in the Eastern North Atlantic (ACE-ENA) field campaign. The algorithm identifies evidence of the existence of clusters of nearly identical local size distributions. It is found that cloud segments have as few as one and as many as seven characteristic size distributions. To validate the algorithm’s robustness, it is tested on a synthetic dataset and successfully identifies the predefined distributions at plausible noise levels. The algorithm is general and is expected to be useful in other applications, such as remote sensing of cloud and rain properties. Significance Statement A typical cloud can have billions of drops spread over tens or hundreds of kilometers in space. Keeping track of the sizes, positions, and interactions of all of these droplets is impractical, and, as such, information about the relative abundance of large and small drops is typically quantified with a “size distribution.” Droplets in a cloud interact locally, however, so this work is motivated by the question of whether the cloud droplet size distribution is different in different parts of a cloud. A new method, based on hypothesis testing and machine learning, determines how many different size distributions are contained in a given cloud. This is important because the size distribution describes processes such as cloud droplet growth and light transmission through clouds.

54 ENVIRONMENTAL SCIENCES↗

Neural Density Estimation and Uncertainty Quantification for ChemCam Spectra [Slides]

The ChemCam instrument of Curiosity uses laser-induced breakdown spectroscopy (LIBS). It fires a laser at target and vaporizes rock surfaces, creating a plasma. Three spectrographs divide the plasma light into wavelengths for chemical analysis: ultraviolet, violet, and visible near-infrared. Regression methods (SVR, PCR, CNN) have been employed for calibration (prediction of the elemental composition of samples); however, labeled ChemCam samples are limited. Here, we focus on unsupervised learning and employ generative models from ChemCam analysis. Further, we use labels (supervised) in combination to the generative model to compute uncertainties related to predictions. We report generative modeling can be successfully applied to model real-world data. Normalizing flow models can be efficiently constructed on latent spaces for fast downstream inference. Unsupervised and supervised learning can be combined to form an uncertainty quantification framework.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

AI-NERD: Elucidation of relaxation dynamics beyond equilibrium through AI-informed X-ray photon correlation spectroscopy

Abstract Understanding and interpreting dynamics of functional materials in situ is a grand challenge in physics and materials science due to the difficulty of experimentally probing materials at varied length and time scales. X-ray photon correlation spectroscopy (XPCS) is uniquely well-suited for characterizing materials dynamics over wide-ranging time scales. However, spatial and temporal heterogeneity in material behavior can make interpretation of experimental XPCS data difficult. In this work, we have developed an unsupervised deep learning (DL) framework for automated classification of relaxation dynamics from experimental data without requiring any prior physical knowledge of the system. We demonstrate how this method can be used to accelerate exploration of large datasets to identify samples of interest, and we apply this approach to directly correlate microscopic dynamics with macroscopic properties of a model system. Importantly, this DL framework is material and process agnostic, marking a concrete step towards autonomous materials discovery.

36 MATERIALS SCIENCE↗

Semi-Supervised, Non-Intrusive Disaggregation of Nodal Load Profiles With Significant Behind-the-Meter Solar Generation

It is of imperative interests for regional transmission organizations (RTOs) to effectively extract actual load profiles at transmission nodes with significant behind-the-meter solar generation, which remains a gap in the existing technology paradigm. This paper proposes an explicit yet efficient linear estimator to disaggregate actual load profiles at transmission buses with significant behind-the-meter (BTM) solar generations. The proposed estimator is based on disaggregating (i.e., extracting) at locations close to transmission buses under consideration. Further, to overcome the lack of “ground truth” and validate the performance of the proposed algorithms, we first propose semi-supervised mechanisms with parameter tuning as well as unsupervised clustering and leverage the unique characteristics of zero-crossing points in BTM solar peaking behaviors, which we refer to as “Zone-to-Node (Z2N)” methods. Next, we further propose a bi-level Node-to-Node (N2N) framework that improves the overall disaggregation performances compared to Z2N. Numerical results are presented using real-world data at PJM Interconnection.

14 SOLAR ENERGY↗

PixelLearn

PixelLearn is an integrated user-interface computer program for classifying pixels in scientific images. Heretofore, training a machine-learning algorithm to classify pixels in images has been tedious and difficult. PixelLearn provides a graphical user interface that makes it faster and more intuitive, leading to more interactive exploration of image data sets. PixelLearn also provides image-enhancement controls to make it easier to see subtle details in images. PixelLearn opens images or sets of images in a variety of common scientific file formats and enables the user to interact with several supervised or unsupervised machine-learning pixel-classifying algorithms while the user continues to browse through the images. The machinelearning algorithms in PixelLearn use advanced clustering and classification methods that enable accuracy much higher than is achievable by most other software previously available for this purpose. PixelLearn is written in portable C++ and runs natively on computers running Linux, Windows, or Mac OS X.

Mazzoni, Dominic↗

W2VPCA: A Machine Learning Method for Measuring Attitudes With Natural Language

Company strategy influences many decisions in freight transportation. Behavioral models of company decision-making therefore could benefit from including strategy variables. However, strategy is difficult to observe and quantify. Attitudinal surveys of company executives can be used to collect measurements of latent strategy to use in quantitative models. However, surveys are costly and burdensome. Text mining methods to collect measurements overcome these issues somewhat, but typically require manual intervention and ignore the context of words, which can be problematic. This study introduces a new machine learning method to generate strategy measurement data from existing big text data. The new method, called W2VPCA, combines Natural Language Processing and Principal Components Analysis. W2VPCA produces measurement data that serve as quantitative indicators of latent strategy in behavioral models. W2VPCA is unsupervised, data-driven, and uses information on word context. We apply W2VPCA to generate measurements of latent strategies using readily available, large-scale text data: annual company reports. The empirical measurements are used successfully to associate two latent strategies, one focusing on distribution and the other on products, with truck fleet and distribution center outsourcing decisions. The main empirical outcome is that the W2VPCA measurements outperform Bag-of-Words measurements in a psychometric analysis of latent firm strategies. While this study focuses on freight behavioral models, W2VPCA may also have applications in behavioral modeling in other domains.

97 MATHEMATICS AND COMPUTING↗

Creating ground truth for nanocrystal morphology: a fully automated pipeline for unbiased transmission electron microscopy analysis

Control over colloidal nanocrystal morphology (size, size distribution, and shape) is important for tailoring the functionality of individual nanocrystals and their ensemble behavior. Despite this, traditional methods to quantify nanocrystal morphology are laborious. New developments in automated morphology classification will accelerate these analyses but the assessment of machine learning models is limited by human accuracy for ground truth, causing even unsupervised machine learning models to have inherent bias. Herein, we introduce synthetic image rendering to solve the ground truth problem of nanocrystal morphology classification. By simulating 2D images of nanocrystal shapes via a function of high-dimensional parameter space, we trained a convolutional neural network to link unique morphologies to their simulated parameters, defining nanocrystal morphology quantitatively rather than qualitatively. An automated pipeline then processes, quantitatively defines, and classifies nanocrystal morphology from experimental transmission electron microscopy (TEM) images. Using improved computer vision techniques, 42,650 nanocrystals were identified, assessed, and labeled with quantitative parameters, offering a 600-fold improvement in efficiency over best-practice manual measurements. Further, a classification algorithm was trained with a prediction accuracy of 99.5%, which can successfully analyze a range of concave, convex, and irregular nanocrystal shapes. The resulting pipeline was applied to differentiating two syntheses of nominally cuboidal CsPbBr 3 nanocrystals and uniquely classifying binary nickel sulfide nanocrystal phase based on morphology. This pipeline provides a simple, efficient, and unbiased method to quantify nanocrystal morphology and represents a practical route to construct large datasets with an absolute ground truth for training unbiased morphology-based machine learning algorithms.

77 NANOSCIENCE AND NANOTECHNOLOGY↗

Pulsed Thermal Tomography Nondestructive Examination of Additively Manufactured Reactor Materials and Components (Final Technical Report)

Metal Additive Manufacturing (AM) is a promising method for cost-efficient fabrication of complex shape structures for applications in harsh environment, such as in a nuclear reactor. However, internal defects (pores) occur in high-strength AM alloys, which are manufactured with Laser Powder Bed Fusion (LPBF) AM method. Pulsed Infrared Thermography (PIT) is an efficient nondestructive evaluation (NDE) method to examine actual structures, because this method offers one-sided non-contact measurements, and fast processing of large sample areas. However, imaging of material defects, particularly defects with sizes at microscopic level, is challenging. In this report, we benchmark the performance of several Unsupervised Learning (UL) algorithms designed to enhance imaging of microscopic defects in metals with PIT. UL aims to learn the latent principal patterns (dictionaries) in PIT data to detect defects with minimal human supervision. Performance of Independent Component Analysis (ICA), Sparse Coding (SC), Principal Component Analysis (PCA) and Exploratory Factor Analysis (EFA) was compared using F-score, UL model training time and defects reconstruction time. We obtained the average F-score of 0.75, and a highest F-score of 0.89 for the EFA algorithm. Overall, EFA outperforms other UL algorithms considered in this study. In another approach, we investigate Thermal Tomography (TT), which is a computational method for reconstruction of depth profile of internal material defects from PIT nondestructive evaluation (NDE). TT algorithm obtains depth reconstructions of thermal effusivity, which has been shown to provide visualization of subsurface internals defects in metals. In many applications, one needs to determine the defect shape and orientation from reconstructed effusivity images. Interpretation of TT images is non-trivial because of blurring, which increases with depth due to heat diffusion-based nature of image formation. We have developed a deep learning convolutional neural network (CNN) to classify size and orientation of subsurface material defects in TT images. CNN was trained with TT images produced with computer simulations of 2D metallic structures (thin plates) containing elliptical subsurface voids. Performance of CNN was investigated using test TT images developed with computer simulations of plates containing elliptical defects, and defects with shape imported from scanning electron microscopy (SEM) images. CNN demonstrated the ability to classify radii and angular orientation of elliptical defects in previously unseen test TT images. We have also demonstrated that CNN trained on TT images of elliptical defects is capable of classifying shape and orientation of irregular defects. Training the CNN on irregular defect shapes instead of on elliptical shapes would make the resulting classifications more descriptive of actual defect shapes. However, this requires a much higher volume of SEM images of material defects, which are difficult to obtain because of random occurrence of defects in LPBF. To address this challenge, we developed a generative adversarial network (GAN) to augment the existing dataset of SEM defect images. The GAN model is demonstrated to create novel yet realistic defect shapes that can be used as input for simulated PTT images to train CNN. We also investigate several approaches based on Gaussian Random Circle and Bezier Curves for constructing parametric models of irregular-shape defects.

36 MATERIALS SCIENCE↗

Enhancement of Tropical Land Cover Mapping with Wavelet-Based Fusion and Unsupervised Clustering of SAR and Landsat Image Data

The characterization and the mapping of land cover/land use of forest areas, such as the Central African rainforest, is a very complex task. This complexity is mainly due to the extent of such areas and, as a consequence, to the lack of full and continuous cloud-free coverage of those large regions by one single remote sensing instrument, In order to provide improved vegetation maps of Central Africa and to develop forest monitoring techniques for applications at the local and regional scales, we propose to utilize multi-sensor remote sensing observations coupled with in-situ data. Fusion and clustering of multi-sensor data are the first steps towards the development of such a forest monitoring system. In this paper, we will describe some preliminary experiments involving the fusion of SAR and Landsat image data of the Lope Reserve in Gabon. Similarly to previous fusion studies, our fusion method is wavelet-based. The fusion provides a new image data set which contains more detailed texture features and preserves the large homogeneous regions that are observed by the Thematic Mapper sensor. The fusion step is followed by unsupervised clustering and provides a vegetation map of the area.

LeMoigne, Jacqueline↗

Fast and efficient identification of anomalous galaxy spectra with neural density estimation

ABSTRACT Current large-scale astrophysical experiments produce unprecedented amounts of rich and diverse data. This creates a growing need for fast and flexible automated data inspection methods. Deep learning algorithms can capture and pick up subtle variations in rich data sets and are fast to apply once trained. Here, we study the applicability of an unsupervised and probabilistic deep learning framework, the probabilistic auto-encoder, to the detection of peculiar objects in galaxy spectra from the SDSS survey. Different to supervised algorithms, this algorithm is not trained to detect a specific feature or type of anomaly, instead it learns the complex and diverse distribution of galaxy spectra from training data and identifies outliers with respect to the learned distribution. We find that the algorithm assigns consistently lower probabilities (higher anomaly score) to spectra that exhibit unusual features. For example, the majority of outliers among quiescent galaxies are E+A galaxies, whose spectra combine features from old and young stellar population. Other identified outliers include LINERs, supernovae, and overlapping objects. Conditional modelling further allows us to incorporate additional information. Namely, we evaluate the probability of an object being anomalous given a certain spectral class, but other information such as metrics of data quality or estimated redshift could be incorporated as well. We make our code publicly available.

Böhm, Vanessa↗

Machine learning enabled quantification of the hydrogen bonds inside the polyelectrolyte brush layer probed using all-atom molecular dynamics simulations

The configuration of densely grafted charged polyelectrolyte (PE) brushes is strongly dictated by the properties and behavior of the counterions that screen the PE brush charges and the solvent molecules (typically water) that solvate the brush molecules and these screening counterions. Only recently, efforts have been made to study the PE brushes atomistically, thereby shedding light on the properties of brush-supported ions and water molecules. However, even for such efforts, there are limitations associated with using a generic definition to estimate certain properties of water and ions inside the brush layer. For example, water–water hydrogen bonds (HBs) will behave differently for locations outside and inside the brush layer, given the fact that the densely closely grafted PE brush molecules create a soft nanoconfinement where the water connectivity becomes highly disrupted: therefore, using the same definition to quantify the HBs inside and outside the brush layer will be unwise. In this paper, we address this limitation by employing an unsupervised machine learning (ML) approach to predict the water–water hydrogen bonding inside a cationic PE brush layer modeled using all-atom molecular dynamics (MD) simulations. Here, the ML method, which relies on a clustering approach and uses the equilibrium coordinates of the water molecules (obtained from the all-atom MD simulations) as the input, is capable of identifying the structural modification of water–water HBs (revealed through appropriate clustering of the data) inside the PE brush layer induced soft nanoconfinement. Such capabilities would not have been possible by using a generic definition of the HBs. Our calculations lead to four key findings: (1) the clusters formed inside and outside the brush layer are structurally similar; (2) the margin of the cluster is shorter inside the PE brush layer confirming the possible disruption of the HBs inside the PE brush layer; (3) the average “hydrogen–acceptor-oxygen–donor-oxygen” angle that defines the HB is reduced for the HBs formed inside the brush layer; (4) the use of the generic definition (definition usable for characterizing the HBs in brush-free bulk) leads to an overprediction of the number of HBs formed inside the PE brush layer.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Extraction and classification of objects in multispectral images

Presented here is an algorithm that partitions a digitized multispectral image into parts that correspond to objects in the scene being sensed. The algorithm partitions an image into successively smaller rectangles and produces a partition that tends to minimize a criterion function. Supervised and unsupervised classification techniques can be applied to partitioned images. This partition-then-classify approach is used to process images sensed from aircraft and the ERTS-1 satellite, and the method is shown to give relatively accurate results in classifying agricultural areas and extracting urban areas.

Robertson, T. V.↗

Comparing storm resolving models and climates via unsupervised machine learning

Global storm-resolving models (GSRMs) have gained widespread interest because of the unprecedented detail with which they resolve the global climate. However, it remains difficult to quantify objective differences in how GSRMs resolve complex atmospheric formations. This lack of comprehensive tools for comparing model similarities is a problem in many disparate fields that involve simulation tools for complex data. To address this challenge we develop methods to estimate distributional distances based on both nonlinear dimensionality reduction and vector quantization. Our approach automatically learns physically meaningful notions of similarity from low-dimensional latent data representations that the different models produce. This enables an intercomparison of nine GSRMs based on their high-dimensional simulation data (2D vertical velocity snapshots) and reveals that only six are similar in their representation of atmospheric dynamics. Furthermore, we uncover signatures of the convective response to global warming in a fully unsupervised way. Our study provides a path toward evaluating future high-resolution simulation data more objectively.

54 ENVIRONMENTAL SCIENCES↗

Anomalously high elastic modulus of a poly(ethylene oxide)-based composite electrolyte

The practical use of lithium metal anodes in solid-state batteries requires a polymer membrane with high lithium-ion conductivity, thermal/electrochemical stability, and mechanical strength. The primary challenge is to effectively decouple the ionic conductivity and mechanical strength of the polymer electrolytes. We report a remarkably facile single step synthetic strategy based on in-situ crosslinking of poly(ethylene oxide) (xPEO) in the presence of a woven glass fiber (GF). Such a simple method yields composite polymer electrolytes (CPE) of anomalously high elastic modulus up to 2.5 GPa over a broad temperature range (20 °C – 245 °C) that has never been previously documented. An unsupervised machine learning algorithm, K-mean clustering analysis, was implemented on the hyperspectral Raman mapping at the xPEO/GF interface. Using such a unique means, we show for the first time that the promoted mechanical strength originates from xPEO and GF interactions through dynamic hydrogen and ionic bonding. High ionic conductivity is achieved by the addition plasticizer (e.g. tetraglyme), where trifluoromethanesulfonate anions are tethered to the xPEO matrix and Li + cations are favorably transported through coordination with the plasticizer. Further, stringent galvanostatic cycling tests indicates the CPE can be stably cycled for >3000 h in a Li-metal symmetric cell at a moderate temperature (nearly 1500 Coulombs/cm 2 Li equivalents), outperforming most of the PEO-based electrolytes. The GF reinforced CPE reported here has multifunctional uses, such as solid electrolytes for all solid-state batteries and membranes for redox-flow batteries. Although the focus of this study is on lithium-based batteries, the results are equally promising for other alkali metal based batteries such as sodium and potassium.

25 ENERGY STORAGE↗

TomoEncoders: 3D Autoencoders for feature extraction in X-ray tomography

Real-time steering of time-resolved or in-situ X-ray tomography requires capturing changes in morphological descriptors in a sample (e.g., porosity, particle size, and crack width) during continuous data acquisition. Image segmentation (2D or3D) followed by quantitative measurement is the conventional method for tracking changes in these descriptors with respect to a previous time-step or a 3D search in a volume. However, image segmentation is expensive. As a faster and unsupervised alternative, a feature-extraction approach using a convolutional autoencoders was developed, where the latent space of the encoder responds to relative changes in morphology with-out prior knowledge of the morphological descriptors.

TEKAWADE, ANIKET↗

An Investigation of State-Space Model Fidelity for SSME Data

In previous studies, a variety of unsupervised anomaly detection techniques for anomaly detection were applied to SSME (Space Shuttle Main Engine) data. The observed results indicated that the identification of certain anomalies were specific to the algorithmic method under consideration. This is the reason why one of the follow-on goals of these previous investigations was to build an architecture to support the best capabilities of all algorithms. We appeal to that goal here by investigating a cascade, serial architecture for the best performing and most suitable candidates from previous studies. As a precursor to a formal ROC (Receiver Operating Characteristic) curve analysis for validation of resulting anomaly detection algorithms, our primary focus here is to investigate the model fidelity as measured by variants of the AIC (Akaike Information Criterion) for state-space based models. We show that placing constraints on a state-space model during or after the training of the model introduces a modest level of suboptimality. Furthermore, we compare the fidelity of all candidate models including those embodying the cascade, serial architecture. We make recommendations on the most suitable candidates for application to subsequent anomaly detection studies as measured by AIC-based criteria.

Martin, Rodney Alexander↗

Exploring Continuous Seismic Data at an Industry Facility Using Unsupervised Machine Learning

Seismic data recorded at industrial sites contain valuable information on anthropogenic activities. With advances in machine learning and computing power, new opportunities have emerged to explore the seismic wavefield in these complex environments. We applied two unsupervised machine learning algorithms to analyze continuous seismic data collected from an industrial facility in Texas, United States. The Uniform Manifold Approximation and Projection for Dimension Reduction algorithm was used to reduce the dimensionality of the data and generate 2D embeddings. Then, the Hierarchical Density-Based Spatial Clustering of Applications with Noise method was employed to automatically group these embeddings into distinct signal clusters. Our analysis of over 1400 hr (around 59 days) of continuous seismic data revealed five and seven signal clusters at two separate stations. At both stations, we identified clusters associated with background noise and vehicle traffic, with the latter’s temporal patterns aligning closely with the facility’s work schedule. Furthermore, the algorithms detected signal clusters from unknown sources and underline the ability of unsupervised machine learning for uncovering previously unrecognized patterns. Our analysis demonstrates the effectiveness of unsupervised approaches in examining continuous seismic data without requiring prior knowledge or pre-existing labels.

58 GEOSCIENCES↗

Objective Phenotyping of Root System Architecture Using Image Augmentation and Machine Learning in Alfalfa (Medicago sativa L.)

Active breeding programs specifically for root system architecture (RSA) phenotypes remain rare; however, breeding for branch and taproot types in the perennial crop alfalfa is ongoing. Phenotyping in this and other crops for active RSA breeding has mostly used visual scoring of specific traits or subjective classification into different root types. While image-based methods have been developed, translation to applied breeding is limited. This research is aimed at developing and comparing image-based RSA phenotyping methods using machine and deep learning algorithms for objective classification of 617 root images from mature alfalfa plants collected from the field to support the ongoing breeding efforts. Our results show that unsupervised machine learning tends to incorrectly classify roots into a normal distribution with most lines predicted as the intermediate root type. Encouragingly, random forest and TensorFlow-based neural networks can classify the root types into branch-type, taproot-type, and an intermediate taproot-branch type with 86% accuracy. With image augmentation, the prediction accuracy was improved to 97%. Coupling the predicted root type with its prediction probability will give breeders a confidence level for better decisions to advance the best and exclude the worst lines from their breeding program. This machine and deep learning approach enables accurate classification of the RSA phenotypes for genomic breeding of climate-resilient alfalfa.

59 BASIC BIOLOGICAL SCIENCES↗