Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “unsupervised method”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Internship Final Report on the unsupervised learning sensor fusion (ULSF) approach

This paper describes a summer internship project undertaken at Sandia National Labs (SNL), both current status and future work. The project was to explore various machine learning approaches for use on turbulent flow data. Specifically, unsupervised classification of turbulent flow data was explored. First, the usage of models in this field is discussed, and several issues in the common usage of the models are identified. Solutions to these issues are then proposed, in the form of a Bayesian filtering approach which probabilistically incorporates multiple sources of data to improve confidence in a result. Several types of sensors are suggested for this method, the incorporation of which range from semi-supervised learning approaches to fully unsupervised. These approaches are then tested on several turbulent flow cases.

97 MATHEMATICS AND COMPUTING↗

Extracting galactic structure parameters from multivariated density estimation

Multivariate statistical analysis, including includes cluster analysis (unsupervised classification), discriminant analysis (supervised classification) and principle component analysis (dimensionlity reduction method), and nonparameter density estimation have been successfully used to search for meaningful associations in the 5-dimensional space of observables between observed points and the sets of simulated points generated from a synthetic approach of galaxy modelling. These methodologies can be applied as the new tools to obtain information about hidden structure otherwise unrecognizable, and place important constraints on the space distribution of various stellar populations in the Milky Way. In this paper, we concentrate on illustrating how to use nonparameter density estimation to substitute for the true densities in both of the simulating sample and real sample in the five-dimensional space. In order to fit model predicted densities to reality, we derive a set of equations which include n lines (where n is the total number of observed points) and m (where m: the numbers of predefined groups) unknown parameters. A least-square estimation will allow us to determine the density law of different groups and components in the Galaxy. The output from our software, which can be used in many research fields, will also give out the systematic error between the model and the observation by a Bayes rule.

Chen, B.↗

Solar forecasting using machine learned cloudiness classification

Methods and systems for predicting irradiance include learning a classification model using unsupervised learning based on historical irradiance data. The classification model is updated using supervised learning based on an association between known cloudiness states and historical weather data. A cloudiness state is predicted based on forecasted weather data. An irradiance is predicted using a regression model associated with the cloudiness state.

Hamann, Hendrik F.↗

Automated integration gate selection for Gaussian mixture model pulse shape discrimination

Pulse shapes differ between neutron and gamma particles when measured with detector devices employing pulse shape discriminating (PSD) scintillators. Digitized waveforms can be used in detection systems to perform pulse shape discrimination for this application. Prior Gaussian Mixture Model (GMM) methods require access to the pulse full-waveform. Reducing the waveform to a smaller set of combined samples reduces computational cost while affecting PSD performance. In this work, we develop a method for selecting the best performing combination of integration gates, or contiguous summed segments of the digitized pulse for PSD. The method uses a discrimination score based on the GMM PSD approach. Furthermore, the final selection is performed using Bayesian Optimization. PSD detection results are compared with varying numbers of selected gates on time-of-flight (TOF) data. This method can be used to fully automate the selection of gates in an unsupervised (without ground truth labels) setting.

42 ENGINEERING↗

AICCA: AI-Driven Cloud Classification Atlas

Clouds play an important role in the Earth’s energy budget, and their behavior is one of the largest uncertainties in future climate projections. Satellite observations should help in understanding cloud responses, but decades and petabytes of multispectral cloud imagery have to date received only limited use. This study describes a new analysis approach that reduces the dimensionality of satellite cloud observations by grouping them via a novel automated, unsupervised cloud classification technique based on a convolutional autoencoder, an artificial intelligence (AI) method good at identifying patterns in spatial data. Our technique combines a rotation-invariant autoencoder and hierarchical agglomerative clustering to generate cloud clusters that capture meaningful distinctions among cloud textures, using only raw multispectral imagery as input. Cloud classes are therefore defined based on spectral properties and spatial textures without reliance on location, time/season, derived physical properties, or pre-designated class definitions. We use this approach to generate a unique new cloud dataset, the AI-driven cloud classification atlas (AICCA), which clusters 22 years of ocean images from the Moderate Resolution Imaging Spectroradiometer (MODIS) on NASA’s Aqua and Terra instruments—198 million patches, each roughly 100 km × 100 km (128 × 128 pixels)—into 42 AI-generated cloud classes, a number determined via a newly-developed stability protocol that we use to maximize richness of information while ensuring stable groupings of patches. AICCA thereby translates 801 TB of satellite images into 54.2 GB of class labels and cloud top and optical properties, a reduction by a factor of 15,000. The 42 AICCA classes produce meaningful spatio-temporal and physical distinctions and capture a greater variety of cloud types than do the nine International Satellite Cloud Climatology Project (ISCCP) categories—for example, multiple textures in the stratocumulus decks along the West coasts of North and South America. We conclude that our methodology has explanatory power, capturing regionally unique cloud classes and providing rich but tractable information for global analysis. AICCA delivers the information from multi-spectral images in a compact form, enables data-driven diagnosis of patterns of cloud organization, provides insight into cloud evolution on timescales of hours to decades, and helps democratize climate research by facilitating access to core data.

97 MATHEMATICS AND COMPUTING↗

Unsupervised learning-enabled pulsed infrared thermographic microscopy of subsurface defects in stainless steel

Metallic structures produced with laser powder bed fusion (LPBF) additive manufacturing method (AM) frequently contain microscopic porosity defects, with typical approximate size distribution from one to 100 microns. Presence of such defects could lead to premature failure of the structure. In principle, structural integrity assessment of LPBF metals can be accomplished with nondestructive evaluation (NDE). Pulsed infrared thermography (PIT) is a non-contact, one-sided NDE method that allows for imaging of internal defects in arbitrary size and shape metallic structures using heat transfer. PIT imaging is performed using compact instrumentation consisting of a flash lamp for deposition of a heat pulse, and a fast frame infrared (IR) camera for measuring surface temperature transients. However, limitations of imaging resolution with PIT include blurring due to heat diffusion, sensitivity limit of the IR camera. We demonstrate enhancement of PIT imaging capability with unsupervised learning (UL), which enables PIT microscopy of subsurface defects in high strength corrosion resistant stainless steel 316 alloy. PIT images were processed with UL spatial–temporal separation-based clustering segmentation (STSCS) algorithm, refined by morphology image processing methods to enhance visibility of defects. The STSCS algorithm starts with wavelet decomposition to spatially de-noise thermograms, followed by UL principal component analysis (PCA), fine-tuning optimization, and neural learning-based independent component analysis (ICA) algorithms to temporally compress de-noised thermograms. The compressed thermograms were further processed with UL-based graph thresholding K-means clustering algorithm for defects segmentation. The STSCS algorithm also includes online learning feature for efficient re-training of the model with new data. For this study, metallic specimens with calibrated microscopic flat bottom hole defects, with diameters in the range from 203 to 76 µm, were produced using electro discharge machining (EDM) drilling. While the raw thermograms do not show any material defects, using STSCS algorithm to process PIT images reveals defects as small as 101 µm in diameter. To the best of our knowledge, this is the smallest reported size of a sub-surface defect in a metal imaged with PIT, which demonstrates the PIT capability of detecting defects in the size range relevant to quality control requirements of LPBF-printed high-strength metals.

36 MATERIALS SCIENCE↗

Lithium-Ion Battery Diagnostics Using Electrochemical Impedance via Machine-Learning

Diagnosing battery states such as health, state-of-charge, or temperature is crucial for ensuring the safety and reliability of electrochemical energy storage systems. While some states, such as temperature, may be measured using cheap sensors, accurate diagnosis of battery health metrics usually requires time-consuming performance measurements, making them infeasible for use in real-world operation. These health metrics can be measured during lab-testing and then estimated on-line using predictive life models or via state observer algorithms such as Kalman filters, but these predictive methods should be supplemented by actual measurement of battery health whenever possible to ensure reliability. Rapid measurement of battery health may be done by various types of fast diagnostic techniques such as electrochemical impedance spectroscopy (EIS), which can be performed in only a few minutes and require only a fraction of the energy and power needed for a full charge and discharge measurement. But there is a substantial challenge for estimating battery health using EIS data, as EIS is sensitive to cell temperature, state-of-charge, current, and resting time in addition to health. Thus, utilizing EIS data to predict battery capacity requires correcting for all these additional variables, a task that is extremely difficult to handle analytically. This talk utilizes machine-learning methods to estimate the effectiveness of battery capacity prediction from EIS data, leveraging a data set of hundreds of EIS measurements recorded at varying temperature and state-of-charge throughout a 500-day aging study of 32 commercial, large-format NMC-Graphite lithium-ion batteries. Using EIS as input to machine-learning models is complicated by the nonlinear response of impedance to battery health, temperature, and state-of-charge, as well as the collinearity between the impedance response at neighboring frequencies, which can easily lead to overfit models. To train robust models, features from EIS data need to be extracted from the data or some subset of critical frequencies selected. Many approaches for extracting and selecting features from EIS data from electrochemical analysis and machine-learning fields were identified for analysis: using the entire raw spectra; selection of one, two, or many frequencies from the entire spectra; selecting interesting points from the EIS measurement using domain knowledge; fitting EIS with an equivalent-circuit model; calculating statistics on the raw impedance values; and reducing the dimensionality of the data using unsupervised linear (principal component analysis) and non-linear (uniform manifold approximation and projection) methods. These approaches were rigorously compared using a machine-learning pipeline approach, training linear, Gaussian process, and random forest regression models and quantifying performance using cross-validation as well as a held-out test set. An artificial neural network model trained on the raw spectra was also tested. Promising pipelines were fine-tuned via Bayesian hyperparameter optimization using cross-validation loss and training with class-specific weights to counter data set imbalance. The most reliable method for utilizing impedance in this work was the selection of two optimal frequencies through an exhaustive search, resulting in about 2% mean absolute error on test data for both Gaussian process and random forest model architectures. Interrogation of a variety of models reveals critical frequencies of 100 Hz and 103 Hz for this data set, though the optimal set of frequencies is not necessarily intuitive, i.e., the best performing models are not simply those that use impedance at frequencies that have the highest correlation to the relative discharge capacity. The best performing model is an ensemble model, which is able to predict battery capacity with 1.9% mean absolute error for unseen cells using impedance recorded at a variety of temperatures and states-of-charge.

battery↗

End-to-end learning of multiple sequence alignments with differentiable Smith–Waterman

Abstract Motivation Multiple sequence alignments (MSAs) of homologous sequences contain information on structural and functional constraints and their evolutionary histories. Despite their importance for many downstream tasks, such as structure prediction, MSA generation is often treated as a separate pre-processing step, without any guidance from the application it will be used for. Results Here, we implement a smooth and differentiable version of the Smith–Waterman pairwise alignment algorithm that enables jointly learning an MSA and a downstream machine learning system in an end-to-end fashion. To demonstrate its utility, we introduce SMURF (Smooth Markov Unaligned Random Field), a new method that jointly learns an alignment and the parameters of a Markov Random Field for unsupervised contact prediction. We find that SMURF learns MSAs that mildly improve contact prediction on a diverse set of protein and RNA families. As a proof of concept, we demonstrate that by connecting our differentiable alignment module to AlphaFold2 and maximizing predicted confidence, we can learn MSAs that improve structure predictions over the initial MSAs. Interestingly, the alignments that improve AlphaFold predictions are self-inconsistent and can be viewed as adversarial. This work highlights the potential of differentiable dynamic programming to improve neural network pipelines that rely on an alignment and the potential dangers of optimizing predictions of protein sequences with methods that are not fully understood. Availability and implementation Our code and examples are available at: https://github.com/spetti/SMURF. Supplementary information Supplementary data are available at Bioinformatics online.

59 BASIC BIOLOGICAL SCIENCES↗

Unsupervised Machine Learning for Exploratory Data Analysis of Exoplanet Transmission Spectra

Abstract Transit spectroscopy is a powerful tool for decoding the chemical compositions of the atmospheres of extrasolar planets. In this paper, we focus on unsupervised techniques for analyzing spectral data from transiting exoplanets. After cleaning and validating the data, we demonstrate methods for: (i) initial exploratory data analysis, based on summary statistics (estimates of location and variability); (ii) exploring and quantifying the existing correlations in the data; (iii) preprocessing and linearly transforming the data to its principal components; (iv) dimensionality reduction and manifold learning; (v) clustering and anomaly detection; and (vi) visualization and interpretation of the data. To illustrate the proposed unsupervised methodology, we use a well-known public benchmark data set of synthetic transit spectra. We show that there is a high degree of correlation in the spectral data, which calls for appropriate low-dimensional representations. We explore a number of different techniques for such dimensionality reduction and identify several suitable options in terms of summary statistics, principal components, etc. We uncover interesting structures in the principal component basis, namely well-defined branches corresponding to different chemical regimes of the underlying atmospheres. We demonstrate that those branches can be successfully recovered with a K-means clustering algorithm in a fully unsupervised fashion. We advocate for lower-dimensional representations of the spectroscopic data in terms of the main principal components, in order to reveal the existing structure in the data and quickly characterize the chemical class of a planet.

Matchev, Konstantin T. (ORCID:0000000341829096)↗

Annual Report for Structure-Aware Unsupervised, Transformational Machine Learning for Drug Discovery

The major goal of this project is to develop machine learning (ML) methods to enable improved predictive power on real drug discovery for novel targets. More specifically, we plan to demonstrate the capability and effectiveness of ML tools utilizing unlabeled large-volume protein-ligand datasets. We also plan to demonstrate the capability and effectiveness of the developed methods by testing on a realistic drug discovery task to identify pan-coronavirus protease inhibitors such as SARS-CoV-2. While the overall goals and milestones remain consistent with the original proposal, certain technical details have been modified, which we will describe in this report.

97 MATHEMATICS AND COMPUTING↗

The use of unsupervised clustering as a classifier for LACIE MSS data

The author has identified the following significant results. This classification method appears to give accurate field center results and to give practical, statistically consistent and accurate estimates of crop proportions. The accuracy of this method is attributable to certain qualities of the particular clustering algorithm. These qualities are freedom from assumptions about Gaussian data, and the continual updating of distribution estimates, including updating the number of modes. This method is relatively tolerant of errors in the determination of crop type, as crop identity is used only for identifying clusters, and not for computing signatures.

Pentland, A. P.↗

Comparative Analysis of TRGBs (CATs) from Unsupervised, Multi-halo-field Measurements: Contrast is Key

The tip of the red giant branch (TRGB) is an apparent discontinuity of the luminosity function (LF) due to the end of the red giant evolutionary phase and is used to measure distances in the local universe. In practice, tip localization via edge detection response (EDR) relies on several methods applied on a case-by-case basis. It is hard to evaluate how individual choices affect a distance estimation using only a single host field while also avoiding confirmation bias. To devise a standardized approach, we compare unsupervised, algorithmic analyses of the TRGB in multiple halo fields per galaxy. We first optimize methods for the lowest field-to-field dispersion, including spatial filtering, smoothing, and weighting of LF, color band selection, and tip selection based on the number of likely RGB stars and the ratio of stars below versus above the tip (R). We find R, which we call the tip contrast, to be the most important indicator of the quality of EDR measurements; higher R selection can decrease field-to-field dispersion. Further, since R is found to correlate with the age or metallicity of the stellar population based on theoretical modeling, it might result in a displacement of the detected tip magnitude. We find a tip-contrast relation with a slope of -0.023 ± 0.0046 mag/ratio, an ~5σ result that can be used to correct these variations in the detections. When using TRGB to establish a distance ladder, consistent TRGB standardization using tip-contrast relation across rungs is vital to make robust cosmological measurements.

79 ASTRONOMY AND ASTROPHYSICS↗

VoroClust

SAND2025-11465O VoroClust, also known as Voronoi Clustering, is a fast, density-based unsupervised clustering algorithm applicable to high-resolution and high-dimensional data. It operates as quickly as distance-based clustering methods while effectively capturing complex regional geometries, matching the performance of current density-based methods. VoroClust employs a data-centered sphere cover to reduce computational demands while preserving data topology. It propagates clusters outward from local density peaks. Although supervised machine learning is powerful for applications like image classification and segmentation, it requires comprehensive, consistent datasets, which many applications lack. Unsupervised clustering algorithms analyze the structure of each dataset rather than relying on similarities with other examples, making them well-suited for practical applications with insufficient or inappropriate data for supervised learning. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Ebeida, Mohamed [Sandia National Lab. (SNL-CA), Li↗

Can Selforganizing Maps Accurately Predict Photometric Redshifts?

We present an unsupervised machine-learning approach that can be employed for estimating photometric redshifts. The proposed method is based on a vector quantization called the self-organizing-map (SOM) approach. A variety of photometrically derived input values were utilized from the Sloan Digital Sky Survey's main galaxy sample, luminous red galaxy, and quasar samples, along with the PHAT0 data set from the Photo-z Accuracy Testing project. Regression results obtained with this new approach were evaluated in terms of root-mean-square error (RMSE) to estimate the accuracy of the photometric redshift estimates. The results demonstrate competitive RMSE and outlier percentages when compared with several other popular approaches, such as artificial neural networks and Gaussian process regression. SOM RMSE results (using delta(z) = z(sub phot) - z(sub spec)) are 0.023 for the main galaxy sample, 0.027 for the luminous red galaxy sample, 0.418 for quasars, and 0.022 for PHAT0 synthetic data. The results demonstrate that there are nonunique solutions for estimating SOM RMSEs. Further research is needed in order to find more robust estimation techniques using SOMs, but the results herein are a positive indication of their capabilities when compared with other well-known methods

Way, Michael J.↗

Missing Wedge Completion via Unsupervised Learning with Coordinate Networks

Cryogenic electron tomography (cryoET) is a powerful tool in structural biology, enabling detailed 3D imaging of biological specimens at a resolution of nanometers. Despite its potential, cryoET faces challenges such as the missing wedge problem, which limits reconstruction quality due to incomplete data collection angles. Recently, supervised deep learning methods leveraging convolutional neural networks (CNNs) have considerably addressed this issue; however, their pretraining requirements render them susceptible to inaccuracies and artifacts, particularly when representative training data is scarce. To overcome these limitations, we introduce a proof-of-concept unsupervised learning approach using coordinate networks (CNs) that optimizes network weights directly against input projections. This eliminates the need for pretraining, reducing reconstruction runtime by 3–20× compared to supervised methods. Our in silico results show improved shape completion and reduction of missing wedge artifacts, assessed through several voxel-based image quality metrics in real space and a novel directional Fourier Shell Correlation (FSC) metric. Our study illuminates benefits and considerations of both supervised and unsupervised approaches, guiding the development of improved reconstruction strategies.

42 ENGINEERING↗

AutoPhaseNN: unsupervised physics-aware deep learning of 3D nanoscale Bragg coherent diffraction imaging

Abstract The problem of phase retrieval underlies various imaging methods from astronomy to nanoscale imaging. Traditional phase retrieval methods are iterative and are therefore computationally expensive. Deep learning (DL) models have been developed to either provide learned priors or completely replace phase retrieval. However, such models require vast amounts of labeled data, which can only be obtained through simulation or performing computationally prohibitive phase retrieval on experimental datasets. Using 3D X-ray Bragg coherent diffraction imaging (BCDI) as a representative technique, we demonstrate AutoPhaseNN, a DL-based approach which learns to solve the phase problem without labeled data. By incorporating the imaging physics into the DL model during training, AutoPhaseNN learns to invert 3D BCDI data in a single shot without ever being shown real space images. Once trained, AutoPhaseNN can be effectively used in the 3D BCDI data inversion about 100× faster than iterative phase retrieval methods while providing comparable image quality.

36 MATERIALS SCIENCE↗

Scaling Building Energy Audits through Machine Learning Methods on Novel Drone Image Data

Building energy audits are time-consuming and labor-intensive. This paper describes a new method using machine learning (ML) techniques on novel data sources (drone images) to improve the identification of building characteristics and retrofit opportunities, and thereby reduce the effort for audits. The new ML method includes: (1) Building footprint extraction using line extraction, polygonization, and polygon-merging, (2) Building envelope extraction using PIX4d modeling software to reconstruct a building 3D model, (3) Visualization tool for viewing images from the 3D model, (4) Window-to-wall ratio (WWR) using state-of-art deep neural network semantic segmentation, (5) Envelope thermal anomaly detection using an unsupervised machine learning clustering algorithm, and (6) Rooftop energy equipment detection based on an object detection algorithm. The testing of this method involved a comparison of additional ML-generated information overlaid on current ‘state-of-practice’ audit and remote assessment baselines using evaluation metrics: labor time and associated cost, marginal benefits of using ML-generated information in workflows for audits and remote assessments, integration potential with existing processes and tools, and replicability/scalability of the method. In two test buildings in California that had comprehensive drawings and meter data available, the ML method effectively generated a building footprint, envelope, rooftop equipment, WWR, and locations of envelope thermal anomalies. Projected target segments of the ML method are sites with minimal drawings and energy data, and underserved sectors such as multistoried housing, disadvantaged communities, and schools for which the ML method can enable identification of building asset characteristics and prioritization of envelope retrofits and decentralized energy equipment retrofits.

Singh, Reshma↗

Enhancing transfer learning in angle-resolved photoemission spectroscopy (ARPES) with spatially-aware representations via graph convolution

A recent application of machine learning has been to spatially-resolved angle-resolved photoemission spectroscopy (ARPES). Here we advance the state-of-the-art by applying representational learning to transform ARPES data into an embedding space of a pre-trained self-supervised learning model, thus enhancing the pipeline that improves the bandstructure classification and domain assignment/segmentation performance compared to a k-means clustering method. In the current iteration, the real-space information is entered into the domain assignment through the graph convolution method, which improves the transfer learning performance of the original self-supervised model. Lastly, an unsupervised automated tool is developed that incorporates these techniques to enable automatic domain assignment.

ARPES↗