Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “unsupervised machine learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Hunting for Polluted White Dwarfs and Other Treasures with Gaia XP Spectra and Unsupervised Machine Learning

White dwarfs (WDs) polluted by exoplanetary material provide the unprecedented opportunity to directly observe the interiors of exoplanets. However, spectroscopic surveys are often limited by brightness constraints, and WDs tend to be very faint, making detections of large populations of polluted WDs difficult. In this paper, we aim to increase considerably the number of WDs with multiple metals in their atmospheres. Using 96,134 WDs with Gaia DR3 BP/RP (XP) spectra, we constructed a 2D map using an unsupervised machine-learning technique called Uniform Manifold Approximation and Projection (UMAP) to organize the WDs into identifiable spectral regions. The polluted WDs are among the distinct spectral groups identified in our map. We have shown that this selection method could potentially increase the number of known WDs with five or more metal species in their atmospheres by an order of magnitude. Such systems are essential for characterizing exoplanet diversity and geology.

79 ASTRONOMY AND ASTROPHYSICS↗

4D-STEM Coupled with Unsupervised Machine Learning to Reveal at Large-Scale the Microstructural Evolution in Li- and Mn-Rich Cathodes

Li- and Mn-rich (LMR) layered oxides are known to exhibit a thin surface reconstruction layer, which grows during electrochemical cycling in a manner that depends on exposed crystallographic facets, cycling conditions, and electrolyte chemistry. Direct characterization of this layer has traditionally relied on high-resolution electron microscopy, which is inherently limited to small fields of view. Here, we employ four-dimensional scanning transmission electron microscopy (4D-STEM) combined with unsupervised machine-learning clustering to quantitatively map phase distributions over large areas and track their evolution in LMR cathodes during electrochemical aging. Our results show that the surface reconstruction layer consists predominantly of a rocksalt phase, whose thickness varies across different facets following activation cycling and becomes substantially thicker and more uniform during calendar aging. In contrast, a spinel-like phase is observed within the particle bulk. Large-area phase mapping and correlative high-resolution imaging reveal that this spinel-like phase preferentially nucleates at bulk crystallographic defects, including boundaries between 60°-rotated layered domains and associated mixed-phase regions, rather than exclusively at the particle surface. Our findings establish a mechanistic distinction between surface-driven rocksalt formation and bulk-defect-mediated spinel nucleation while demonstrating the unique capability of 4D-STEM to provide statistically robust, mesoscale insight into complex phase-evolution processes in LMR cathodes.

4D-STEM↗

Informed unsupervised machine learning analysis of dislocation microstructure from high-resolution differential aperture X-ray structural microscopy data

This study leverages high-resolution differential-aperture X-ray structural microscopy (DAXM) to probe the local dislocation structure in deformed 304L-stainless steel at small strain, by measuring the lattice rotation and deviatoric elastic strain with a sub-micron resolution. For a single grain in a polycrystalline specimen, the measured lattice rotation field over the measured volume exhibited a multimodal distribution while the deviatoric elastic strain showed a single-mode distribution. An unsupervised Cauchy mixture machine learning model was developed to resolve the multimodal distribution of the lattice rotation. By mapping the lattice rotation data associated with each Cauchy peak in the model back onto the measured volume, we identify contiguous regions of the crystal rotated near the average values corresponding to the peaks of the overall rotation distribution. These regions represent the grain subdivision in the microstructure. Finally, the dislocation density tensor was also computed and its norm was laid over the rotation field to detect the subgrain boundaries. This step provided a validation of the Cauchy mixture model for the analysis of the lattice rotation distribution. The current study highlights the integration of advanced X-ray microscopy techniques with data-driven analysis methods to uncover detailed microstructure scales in deformed crystals.

Machine learning; Lattice rotation; High-energy X-↗

Search for New Phenomena in Two-Body Invariant Mass Distributions Using Unsupervised Machine Learning for Anomaly Detection at s = 13 TeV with the ATLAS Detector

Searches for new resonances are performed using an unsupervised anomaly-detection technique. Events with at least one electron or muon are selected from 140 fb − 1 of p p collisions at s = 13 TeV recorded by ATLAS at the Large Hadron Collider. The approach involves training an autoencoder on data, and subsequently defining anomalous regions based on the reconstruction loss of the decoder. Studies focus on nine invariant mass spectra that contain pairs of objects consisting of one light jet or b jet and either one lepton ( e , μ ) , photon, or second light jet or b jet in the anomalous regions. No significant deviations from the background hypotheses are observed. Limits on contributions from generic Gaussian signals with various widths of the resonance mass are obtained for nine invariant masses in the anomalous regions. © 2024 CERN, for the ATLAS Collaboration 2024 CERN

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Sensor Anomaly Detection for Nuclear Reactor Systems Utilizing Linear Regression and K-Means Unsupervised Machine Learning

Nuclear reactors and related systems are becoming increasingly complex due to advancing technologies in next-generation power reactors. This increased complexity necessitates enhanced automation and data management capabilities. To successfully realize autonomous systems, methods must be developed to handle vast volumes of data and effectively distinguish anomalous data from noise and expected data. While impressive models utilizing digital twins and similar approaches are under development, here we propose a simplified model for analyzing fundamental methods and techniques. Initially, we created a general dataset by using initial data from PCTRAN in order to represent ideal steady-state conditions. We then inserted anomalies based on prevalent sensor anomaly types (e.g., point anomalies, linear drift, and downward deviations), along with unusual anomalies such as exponential drift and upward deviations. To detect anomalies, we developed a program that employs data partitioning and linear regression to preprocess and filter the anomalous data. A K-Means machine learning (ML) method was then applied to separate and count the data within the anomalous partition. The results from all datasets—apart from exponential growth—demonstrated positive outcomes, with each returning multiple instances of greaterthan-95% accuracy. We conducted further investigations using Idaho National Laboratory’s RAVEN software to perform a sensitivity analysis on the input variables (R 2 Tolerance, Slope Tolerance, and Window Size) and found that the output variables (Accuracy and Time) were most sensitive to the Window Size. Despite the promising results published, further development is required to effectively apply these methods to nuclear systems. Nevertheless, the strengths of this approach are evident and hold promise for future applications in the field.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Presentation: Sensor Anomaly Detection for Nuclear Reactor Systems Utilizing Linear Regression and K-Means Unsupervised Machine Learning: An overview of methods and results

This presentation is a culmination of work which has occurred over the course of a 10-week internship. Anomaly detection methods must be both robust enough to detect subtle anomalies yet not so sensitive to report false positives, which would result significant loss of revenue. Methods currently being developed for autonomous systems are often pursuing a Digital Twin method, which will look at the entire system and model it as a whole. This presentation, however, focuses less on direct application to an NPP, rather acting as a proof of concept for the methods developed. For the project, we look to develop methods to analyze steady-state data and report anomalies.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Photovoltaic Inverter Failure Mechanism Estimation Using Unsupervised Machine Learning and Reliability Assessment

This article introduces a data-driven approach to assessing failure mechanisms and reliability degradation in outdoor photovoltaic (PV) string inverters. The manufacturer's stated PV inverter lifetime can vary due to the impact of operating site conditions. To address limitations in degradation estimation through accelerated testing, condition monitoring, or degradation modeling, we propose a machine learning (ML) oriented approach. Utilizing data from a 1.4 MW PV power plant operational since 2016, with 46 string PV inverters tied to the grid, we employ the unsupervised one-class support vector machine ML technique to analyze inverter and sensor data, capable of classifying humidity cycling and temperature fluctuations as dominant failure mechanisms. Utilizing the anomaly alert relationship and alert details specific to the inverter, the level of PV inverter output is considered as its availability or available reliability. Subsequently, a continuous Markov model is applied to six-month alert data, revealing an average stated reliability of 20% after 20 years of continuous operation. These results support recommendations for time-bound preventive measures to enhance PV inverter reliability under diverse outdoor conditions. Furthermore, the approach provides a nondestructive, top–down, and generalized method for analyzing any commercial PV inverter exposed to outdoor conditions, contingent on the availability of relevant data.

14 SOLAR ENERGY↗

Unsupervised Machine Learning for Exploratory Data Analysis of Exoplanet Transmission Spectra

Abstract Transit spectroscopy is a powerful tool for decoding the chemical compositions of the atmospheres of extrasolar planets. In this paper, we focus on unsupervised techniques for analyzing spectral data from transiting exoplanets. After cleaning and validating the data, we demonstrate methods for: (i) initial exploratory data analysis, based on summary statistics (estimates of location and variability); (ii) exploring and quantifying the existing correlations in the data; (iii) preprocessing and linearly transforming the data to its principal components; (iv) dimensionality reduction and manifold learning; (v) clustering and anomaly detection; and (vi) visualization and interpretation of the data. To illustrate the proposed unsupervised methodology, we use a well-known public benchmark data set of synthetic transit spectra. We show that there is a high degree of correlation in the spectral data, which calls for appropriate low-dimensional representations. We explore a number of different techniques for such dimensionality reduction and identify several suitable options in terms of summary statistics, principal components, etc. We uncover interesting structures in the principal component basis, namely well-defined branches corresponding to different chemical regimes of the underlying atmospheres. We demonstrate that those branches can be successfully recovered with a K-means clustering algorithm in a fully unsupervised fashion. We advocate for lower-dimensional representations of the spectroscopic data in terms of the main principal components, in order to reveal the existing structure in the data and quickly characterize the chemical class of a planet.

Matchev, Konstantin T. (ORCID:0000000341829096)↗

Measuring Galactic dark matter through unsupervised machine learning

ABSTRACT Measuring the density profile of dark matter in the Solar neighbourhood has important implications for both dark matter theory and experiment. In this work, we apply autoregressive flows to stars from a realistic simulation of a Milky Way-type galaxy to learn – in an unsupervised way – the stellar phase space density and its derivatives. With these as inputs, and under the assumption of dynamic equilibrium, the gravitational acceleration field and mass density can be calculated directly from the Boltzmann equation without the need to assume either cylindrical symmetry or specific functional forms for the galaxy’s mass density. We demonstrate our approach can accurately reconstruct the mass density and acceleration profiles of the simulated galaxy, even in the presence of Gaia-like errors in the kinematic measurements.

Astronomy & Astrophysics↗

Seeking regularity from irregularity: unveiling the synthesis–nanomorphology relationships of heterogeneous nanomaterials using unsupervised machine learning

Nanoscale morphology of functional materials determines their chemical and physical properties. However, despite increasing use of transmission electron microscopy (TEM) to directly image nanomorphology, it remains challenging to quantify the information embedded in TEM data sets, and to use nanomorphology to link synthesis and processing conditions to properties. We develop an automated, descriptor-free analysis workflow for TEM data that utilizes convolutional neural networks and unsupervised learning to quantify and classify nanomorphology, and thereby reveal synthesis–nanomorphology relationships in three different systems. While TEM records nanomorphology readily in two-dimensional (2D) images or three-dimensional (3D) tomograms, we advance the analysis of these images by identifying and applying a universal shape fingerprint function to characterize nanomorphology. After dimensionality reduction through principal component analysis, this function then serves as the input for morphology grouping through unsupervised learning. We demonstrate the wide applicability of our workflow to both 2D and 3D TEM data sets, and to both inorganic and organic nanomaterials, including tetrahedral gold nanoparticles mixed with irregularly shaped impurities, hybrid polymer-patched gold nanoprisms, and polyamide membranes with irregular and heterogeneous 3D crumple structures. In each of these systems, unsupervised nanomorphology grouping identifies both the diversity and the similarity of the nanomaterial across different synthesis conditions, revealing how synthetic parameters guide nanomorphology development. Our work opens possibilities for enhancing synthesis of nanomaterials through artificial intelligence and for understanding and controlling complex nanomorphology, both for 2D systems and in the far less explored case of 3D structures, such as those with embedded voids or hidden interfaces.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Comparing storm resolving models and climates via unsupervised machine learning

Global storm-resolving models (GSRMs) have gained widespread interest because of the unprecedented detail with which they resolve the global climate. However, it remains difficult to quantify objective differences in how GSRMs resolve complex atmospheric formations. This lack of comprehensive tools for comparing model similarities is a problem in many disparate fields that involve simulation tools for complex data. To address this challenge we develop methods to estimate distributional distances based on both nonlinear dimensionality reduction and vector quantization. Our approach automatically learns physically meaningful notions of similarity from low-dimensional latent data representations that the different models produce. This enables an intercomparison of nine GSRMs based on their high-dimensional simulation data (2D vertical velocity snapshots) and reveals that only six are similar in their representation of atmospheric dynamics. Furthermore, we uncover signatures of the convective response to global warming in a fully unsupervised way. Our study provides a path toward evaluating future high-resolution simulation data more objectively.

54 ENVIRONMENTAL SCIENCES↗

Predicting the oxidation states of Mn ions in the oxygen-evolving complex of photosystem II using supervised and unsupervised machine learning

Abstract Serial Femtosecond Crystallography at the X-ray Free Electron Laser (XFEL) sources enabled the imaging of the catalytic intermediates of the oxygen evolution reaction of Photosystem II (PSII). However, due to the incoherent transition of the S-states, the resolved structures are a convolution from different catalytic states. Here, we train Decision Tree Classifier and K-means clustering models on Mn compounds obtained from the Cambridge Crystallographic Database to predict the S-state of the X-ray, XFEL, and CryoEM structures by predicting the Mn’s oxidation states in the oxygen-evolving complex. The model agrees mostly with the XFEL structures in the dark S 1 state. However, significant discrepancies are observed for the excited XFEL states (S 2 , S 3, and S 0 ) and the dark states of the X-ray and CryoEM structures. Furthermore, there is a mismatch between the predicted S-states within the two monomers of the same dimer, mainly in the excited states. We validated our model against other metalloenzymes, the valence bond model and the Mn spin densities calculated using density functional theory for two of the mismatched predictions of PSII. The model suggests designing a more optimized sample delivery and illumiation systems are crucial to precisely resolve the geometry of the advanced S-states to overcome the noncoherent S-state transition. In addition, significant radiation damage is observed in X-ray and CryoEM structures, particularly at the dangler Mn center (Mn4). Our model represents a valuable tool for investigating the electronic structure of the catalytic metal cluster of PSII to understand the water splitting mechanism.

Plant Sciences↗

An unsupervised machine learning based approach to identify efficient spin-orbit torque materials

Materials with large spin–orbit torque (SOT) hold considerable significance for many spintronic applications because of their potential for energy-efficient magnetization switching. Unfortunately, most of the existing materials exhibit an SOT efficiency factor that is much less than unity, requiring a large current for magnetization switching. The search for new materials that can exhibit an SOT efficiency much greater than unity is a topic of active research, and only a few such materials have been identified using conventional approaches. In this paper, we present a machine learning-based approach using a word embedding model that can identify new results by deciphering non-trivial correlations among various items in a specialized scientific text corpus. We show that such a model can be used to identify materials likely to exhibit high SOT and rank them according to their expected SOT strengths. The model captured the essential spintronics knowledge embedded in scientific abstracts within various materials science, physics, and engineering journals and identified 97 new materials to exhibit high SOT. Among them, 16 candidate materials are expected to exhibit an SOT efficiency greater than unity, and one of them has recently been confirmed with experiments with quantitative agreement with the model prediction.

Sayed, Shehrin↗

Disentangling ferroelectric domain wall geometries and pathways in dynamic piezoresponse force microscopy via unsupervised machine learning

In this work, domain switching pathways in ferroelectric materials visualized by dynamic piezoresponse force microscopy (PFM) are explored via variational autoencoder, which simplifies the elements of the observed domain structure, crucially allowing for rotational invariance, thereby reducing the variability of local polarization distributions to a small number of latent variables. For small sampling window sizes the latent space is degenerate, and variability is observed only in the direction of a single latent variable that can be identified with the presence of domain wall. For larger window sizes, the latent space is 2D, and the disentangled latent variables can be generally interpreted as the degree of switching and complexity of domain structure. Applied to multiple consecutive PFM images acquired while monitoring domain switching, the polarization switching mechanism can thus be visualized in the latent space, providing insight into domain evolution mechanisms and their correlation with the microstructure.

77 NANOSCIENCE AND NANOTECHNOLOGY↗

via machinae : Searching for stellar streams using unsupervised machine learning

ABSTRACT We develop a new machine learning algorithm, via machinae, to identify cold stellar streams in data from the Gaia telescope. via machinae is based on ANODE, a general method that uses conditional density estimation and sideband interpolation to detect local overdensities in the data in a model agnostic way. By applying ANODE to the positions, proper motions, and photometry of stars observed by Gaia, via machinae obtains a collection of those stars deemed most likely to belong to a stellar stream. We further apply an automated line-finding method based on the Hough transform to search for line-like features in patches of the sky. In this paper, we describe the via machinae algorithm in detail and demonstrate our approach on the prominent stream GD-1. Though some parts of the algorithm are tuned to increase sensitivity to cold streams, the via machinae technique itself does not rely on astrophysical assumptions, such as the potential of the Milky Way or stellar isochrones. This flexibility suggests that it may have further applications in identifying other anomalous structures within the Gaia data set, for example debris flow and globular clusters.

79 ASTRONOMY AND ASTROPHYSICS↗

Quantifying and Zoning Urban Heat Island Effects Using Unsupervised Machine Learning

This work explores the Urban Heat Island (UHI) effects in Maricopa County, Arizona, employing a simulation-based approach that combines large-scale building energy modeling with advanced spatial analysis. Utilizing the Automatic Building Energy Modeling (AutoBEM) software suite, we simulated the energy consumption for approximately 1.35 million buildings based on the Model America version 1.0 (MAv1) dataset. Our methodology incorporated spatial analysis at multiple scales, including individual buildings, clusters of zones determined by K-means clustering, and geographical level evaluation based on Zip codes. The results revealed significant variations in energy consumption and heat emissions across different building types and urban zones. High-emission hotspots identified through clustering pointed to areas most contributing to the UHI effects. Zip code-based area analysis further contextualized these findings, offering an urban context-based perspective on emission distribution and informing potential urban energy policies for mitigating UHI effects.

Chowdhury, Shovan [ORNL]↗