Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “unsupervised method”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

DNABERT-S: pioneering species differentiation with species-aware DNA embeddings

SUMMARY: We introduce DNABERT-S, a tailored genome model that develops species-aware embeddings to naturally cluster and segregate DNA sequences of different species in the embedding space. Differentiating species from genomic sequences (i.e. DNA and RNA) is vital yet challenging, since many real-world species remain uncharacterized, lacking known genomes for reference. Embedding-based methods are therefore used to differentiate species in an unsupervised manner. DNABERT-S builds upon a pre-trained genome foundation model named DNABERT-2. To encourage effective embeddings to error-prone long-read DNA sequences, we introduce Manifold Instance Mixup (MI-Mix), a contrastive objective that mixes the hidden representations of DNA sequences at randomly selected layers and trains the model to recognize and differentiate these mixed proportions at the output layer. We further enhance it with the proposed Curriculum Contrastive Learning (C2LR) strategy. Empirical results on 28 diverse datasets show DNABERT-S's effectiveness, especially in realistic label-scarce scenarios. For example, it identifies twice more species from a mixture of unlabeled genomic sequences, doubles the Adjusted Rand Index (ARI) in species clustering, and outperforms the top baseline's performance in 10-shot species classification with just a 2-shot training. AVAILABILITY AND IMPLEMENTATION: Model, codes, and data are publically available at https://github.com/MAGICS-LAB/DNABERT_S.

Zhou, Zhihan↗

Combining variational autoencoders and physical bias for improved microscopy data analysis *

Electron and scanning probe microscopy produce vast amounts of data in the form of images or hyperspectral data, such as electron energy loss spectroscopy or 4D scanning transmission electron microscope, that contain information on a wide range of structural, physical, and chemical properties of materials. To extract valuable insights from these data, it is crucial to identify physically separate regions in the data, such as phases, ferroic variants, and boundaries between them. In order to derive an easily interpretable feature analysis, combining with well-defined boundaries in a principled and unsupervised manner, here we present a physics augmented machine learning method which combines the capability of variational autoencoders to disentangle factors of variability within the data and the physics driven loss function that seeks to minimize the total length of the discontinuities in images corresponding to latent representations. Our method is applied to various materials, including NiO-LSMO, BiFeO 3 , and graphene. The results demonstrate the effectiveness of our approach in extracting meaningful information from large volumes of imaging data. The customized codes of the required functions and classes to develop phyVAE is available at https://github.com/arpanbiswas52/phy-VAE.

97 MATHEMATICS AND COMPUTING↗

Towards deep computer vision for in-line defect detection in polymer electrolyte membrane fuel cell materials

Polymer Electrolyte Membrane (PEM) fuel cells are a promising source of alternative energy. However, their production is limited by a lack of well-established methods for quality control of their constituent materials like the membrane-electrode assembly during roll-to-roll manufacturing. One potential solution is the implementation of deep learning methods to detect unwanted defects through their detection in scanned images. Here we explore the detection of defects like scratches, pinholes, and scuffs in a sample dataset of PEM optical images using two deep learning algorithms: Patch Distribution Modeling (PaDiM) for unsupervised anomaly detection and Faster-RCNN for supervised object detection. Both methods achieve scores on performance metrics (ROC-AUC and PRO-AUC for PaDiM and AP for Faster-RCNN) that are comparable to their scores on benchmark datasets. These methods also have the potential to detect a wider range of defects compared to IR thermography and previous optical inspection methods. Overall, deep learning shows promise at detecting relevant defects of interest and has the potential to achieve real-time defect detection.

30 DIRECT ENERGY CONVERSION↗

Performance evaluation of the JPL interim digital SAR processor

The performance of the Interim Digital SAR Processor (IDP) was evaluated. The IDP processor was originally developed for experimental processing of digital SEASAT SAR data. One phase of the system upgrade which features parallel processing in three peripheral array processors, automated estimation for Doppler parameters, and unsupervised image pixel location determination and registration was executed. The method to compensate for the target range curvature effect was improved. A four point interpolation scheme is implemented to replace the nearest neighbor scheme used in the original IDP. The processor still maintains its fast throughput speed. The current performance and capability of the processing modes now available on the IDP system are updated.

Wu, C.↗

A Dynamic PCA and Machine Learning Tool for Automated Identification of Solar Wind Disturbances Impacting Earth’s Magnetosphere

Earth’s magnetosphere is continuously impacted by solar wind and interplanetary magnetic field (IMF) disturbances, such as shocks, discontinuities, magnetic clouds and more. Understanding how such disturbances propagate from the Sun and what is their impact on the different magnetospheric domains is key to understanding and forecasting energy transfer from the solar wind to Earth. The large number of overlapping solar wind and magnetospheric missions carrying magnetometers and the recent advances in communications and data storage technologies have enabled an unprecedented quantity of high-fidelity magnetic field data captured by in-situ spacecraft to be available at the click of a button. However, this massive quantity of available data can prove unwieldy for researchers, limiting the identification of interesting phenomena and disturbances to a relatively small percentage of the total dataset. Several techniques have been previously developed for automated identification of specific types of magnetic anomalies, but these methods are typically mission-specific and can be difficult to generalize. We present initial results for a generic method of automated anomaly detection in magnetic field measurements based on dimensionality reduction and unsupervised clustering via machine learning. The benefit of our technique is its high degree of generalizability and flexibility which make it a most useful data survey tool for a wide range of magnetic field datasets. This method can also be applied simultaneously to other observed time-series properties like plasma density, pressure, and velocity for more accurate event identification. Additionally, the application of this method to data captured by multiple spacecraft enables the simultaneous identification of disturbances and the determination of their propagation characteristics. Initial evaluation of this technique has been performed using data from Magnetospheric MultiScale (MMS) and THEMIS-ARTEMIS missions, providing a testbed scenario for the future Heliophysics Environmental and Radiation Measurement Experiment Suite (HERMES) platform instruments that will measure solar wind and IMF properties from lunar orbit onboard the Gateway station.

Miguel Martinez-Ledesma↗

A Dynamic PCA and Machine Learning Tool for Automated Identification of Solar Wind Disturbances Impacting Earth’s Magnetosphere

Earth’s magnetosphere is continuously impacted by solar wind and interplanetary magnetic field (IMF) disturbances, such as shocks, discontinuities, magnetic clouds and more. Understanding how such disturbances propagate from the Sun and what is their impact on the different magnetospheric domains is key to understanding and forecasting energy transfer from the solar wind to Earth. The large number of overlapping solar wind and magnetospheric missions carrying magnetometers and the recent advances in communications and data storage technologies have enabled an unprecedented quantity of high-fidelity magnetic field data captured by in-situ spacecraft to be available at the click of a button. However, this massive quantity of available data can prove unwieldy for researchers, limiting the identification of interesting phenomena and disturbances to a relatively small percentage of the total dataset. Several techniques have been previously developed for automated identification of specific types of magnetic anomalies, but these methods are typically mission-specific and can be difficult to generalize. We present initial results for a generic method of automated anomaly detection in magnetic field measurements based on dimensionality reduction and unsupervised clustering via machine learning. The benefit of our technique is its high degree of generalizability and flexibility which make it a most useful data survey tool for a wide range of magnetic field datasets. This method can also be applied simultaneously to other observed time-series properties like plasma density, pressure, and velocity for more accurate event identification. Additionally, the application of this method to data captured by multiple spacecraft enables the simultaneous identification of disturbances and the determination of their propagation characteristics. Initial evaluation of this technique has been performed using data from Magnetospheric MultiScale (MMS) and THEMIS-ARTEMIS missions, providing a testbed scenario for the future Heliophysics Environmental and Radiation Measurement Experiment Suite (HERMES) platform instruments that will measure solar wind and IMF properties from lunar orbit onboard the Gateway station.

Miguel Martinez-Ledesma↗

Microcrack Quantification in Composite Materials by a Neural Network Analysis of Ultrasound Spectral Data

Intra-ply microcracking in unlined composite pressure vessels can be very troublesome to detect and when linked through the thickness can provide leak paths that may hinder mission success. The leaks may lead to loss of pressure/propellant, increased risk of explosion and possible cryo-pumping into air pockets within the laminate. Ultrasonic techniques have been shown capable of detecting the presence of microcracking and in this work they are used to quantify the level of microcracking. Resonance ultrasound methods are utilized with artificial neural networks to build a microcrack prediction/measurement tool. Two networks are presented, one unsupervised to provide a qualitative measure of microcracking and one supervised which provides a quantitative assessment of the level of microcracking. The resonant ultrasound spectroscopic method is made sensitive to microcracking by tuning the input spectrum to the higher frequency (shorter wavelength) components allowing more significant interaction with the defects. This interaction causes the spectral characteristics to shift toward lower amplitudes at the higher frequencies. As the density of the defects increases more interactions occur and more drastic amplitude changes are observed. Preliminary experiments to quantify the level of microcracking induced in graphite/epoxy composite samples through a combination of tensile loading and cryogenic temperatures are presented. Both unsupervised (Kohonen) and supervised (radial basis function) artificial neural networks are presented to determine the measurable effect on the resonance spectrum of the ultrasonic data taken from the samples.

Walker, James L.↗

An operational method for estimating signal to noise ratios from data acquired with imaging spectrometers

A method, using the concept of local means and local standard deviations of small imaging blocks and using a box counting procedure, has been developed for unsupervised estimation of the average signal to noise ratios of images in which the noise is additive. The method has been applied to simulated images with Gaussian noise, to images acquired with the Airborne Visible Infrared Imaging Spectrometer and with the Geophysical and Environmental Research Imaging Spectrometer. The method is compared with other techniques for estimating signal to noise ratios from imaging data. It is believed that the method is generally applicable to images having many small homogeneous blocks. Because the method does not take into account interband effects, it is not appropriate to use the method to estimate signal to noise ratios from images having nonnegligible interband radiometric calibration errors at spatial scales less than or equal to the size of the small imaging blocks used in the noise estimation.

Gao, Bo-Cai↗

Unsupervised Image-Based Classification of Corrosion Severity in Automobile Engine Connecting Rods

Corrosion in engine connecting rods is a critical issue in the automotive industry, potentially leading to catastrophic engine failure, monetary losses, and safety hazards. The labor shortage in the industry further emphasizes the need for fast, accurate, and automated corrosion detection methods to ensure appropriate surface treatments can be applied to restore component integrity. We present an unsupervised image-based framework for classifying corrosion severity in automobile engine connecting rods using short-wave infrared (SWIR) and telecentric grayscale imaging. We employ the structural similarity index measure (SSIM) as a dissimilarity metric and the k-medians clustering algorithm for classification. Our algorithm achieves an overall accuracy of 80.64% for SWIR images, with 100% accuracy in classifying highly corroded samples. For grayscale images, the method attains an overall accuracy of 77.42%, with 90.91% accuracy for highly corroded samples. The method’s ability to work with different imaging modalities and its high accuracy in identifying severe corrosion cases make it a promising tool for automated corrosion assessment in the automotive industry, potentially improving efficiency and safety in engine component maintenance.

42 ENGINEERING↗

Unsupervised multimodal fusion of in-process sensor data for advanced manufacturing process monitoring

Effective monitoring of manufacturing processes is crucial for maintaining product quality and operational efficiency. Modern manufacturing environments often generate vast amounts of complementary multimodal data, including visual imagery from various perspectives and resolutions, hyperspectral data, and machine health monitoring information such as actuator positions, accelerometer readings, and temperature measurements. However, fusing and interpreting this complex, high-dimensional data presents significant challenges, particularly when labeled datasets are unavailable or impractical to obtain. This paper presents a novel approach to multimodal sensor data fusion in manufacturing processes, inspired by the Contrastive Language-Image Pre-training (CLIP) model. We leverage contrastive learning techniques to correlate different data modalities without the need for labeled data, overcoming limitations of traditional supervised machine learning methods in manufacturing contexts. Our proposed method demonstrates the ability to handle and learn encoders for five distinct modalities: visual imagery, audio signals, laser position (x and y coordinates), and laser power measurements. By compressing these high-dimensional datasets into low-dimensional representational spaces, our approach facilitates downstream tasks such as process control, anomaly detection, and quality assurance. The unsupervised nature of our method makes it broadly applicable across various manufacturing domains, where large volumes of unlabeled sensor data are common. We evaluate the effectiveness of our approach through a series of experiments, demonstrating its potential to enhance process monitoring capabilities in advanced manufacturing systems. This research contributes to the field of smart manufacturing by providing a flexible, scalable framework for multimodal data fusion that can adapt to diverse manufacturing environments and sensor configurations. The proposed method paves the way for more robust, data-driven decision-making in complex manufacturing processes.

Contrastive Learning↗

Confidentiality-preserving machine learning algorithms for soft-failure detection in optical communication networks

Automated fault management is at the forefront of next-generation optical communication networks. The increase in complexity of modern networks has triggered the need for programmable and software-driven architectures to support the operation of agile and self-managed systems. In these scenarios, the European Telecommunications Standards Institute zero-touch network and service management approach is imperative. The need for machine learning algorithms to process the large volume of telemetry data brings safety concerns as distributed cloud-computing solutions become the preferred approach for deploying reliable communication network automation. This paper’s contribution is twofold. First, we propose a simple yet effective method to guarantee the confidentiality of the telemetry data based on feature scrambling. The method allows the operation of third-party computational services without direct access to the full content of the collected data. Additionally, the effectiveness of four unsupervised machine learning algorithms for soft-failure detection is evaluated when applied to the scrambled telemetry data. The methods are based on factor analysis, principal component analysis, nonlinear principal component analysis, and singular value decomposition. Most dimensionality reduction algorithms have the common property that they can maintain similar levels of fault classification performance while hiding the data structure from unauthorized access. Evaluations of the proposed algorithms demonstrate this capability.

97 MATHEMATICS AND COMPUTING↗

Development of Multimodal Few-Shot Analytics for Electron Micrographs

Recent advances in materials data analytics have provided new avenues for determining process-structure-property (PSP) linkages in a variety of materials. Machine learning techniques including few-shot learning have increased the efficiency of classifying microscopy images for the purposes of material characterization. Attempts at creating a multimodal approach can provide further improvements to current models and help extract more salient features from data. In this vein, raw spectrum data was taken to provide an additional modality to our current pyCHIP classifier. Modifications in segmentation also show potential in improving the accuracy of the pyCHIP classifier. Classifier output was analyzed using network graphs and unsupervised clustering algorithms such as spectral clustering to detect better segmentation methods than the current “chipping” approach. We suggest that the chip selection process can be automated in the future using a combination of these techniques to enable high-throughput analyses.

36 MATERIALS SCIENCE↗

Knowledge Distillation for Anomaly Detection

Unsupervised deep learning techniques are widely used to identify anomalous behaviour. The performance of such methods is a product of the amount of training data and the model size. However, the size is often a limiting factor for the deployment on resource-constrained devices. Here, we present a novel procedure based on knowledge distillation for compressing an unsupervised anomaly detection model into a supervised deployable one and we suggest a set of techniques to improve the detection sensitivity. Compressed models perform comparably to their larger counterparts while significantly reducing the size and memory footprint.

Pol, Adrian Alan↗

Internship Final Report on the unsupervised learning sensor fusion (ULSF) approach

This paper describes a summer internship project undertaken at Sandia National Labs (SNL), both current status and future work. The project was to explore various machine learning approaches for use on turbulent flow data. Specifically, unsupervised classification of turbulent flow data was explored. First, the usage of models in this field is discussed, and several issues in the common usage of the models are identified. Solutions to these issues are then proposed, in the form of a Bayesian filtering approach which probabilistically incorporates multiple sources of data to improve confidence in a result. Several types of sensors are suggested for this method, the incorporation of which range from semi-supervised learning approaches to fully unsupervised. These approaches are then tested on several turbulent flow cases.

97 MATHEMATICS AND COMPUTING↗

Extracting galactic structure parameters from multivariated density estimation

Multivariate statistical analysis, including includes cluster analysis (unsupervised classification), discriminant analysis (supervised classification) and principle component analysis (dimensionlity reduction method), and nonparameter density estimation have been successfully used to search for meaningful associations in the 5-dimensional space of observables between observed points and the sets of simulated points generated from a synthetic approach of galaxy modelling. These methodologies can be applied as the new tools to obtain information about hidden structure otherwise unrecognizable, and place important constraints on the space distribution of various stellar populations in the Milky Way. In this paper, we concentrate on illustrating how to use nonparameter density estimation to substitute for the true densities in both of the simulating sample and real sample in the five-dimensional space. In order to fit model predicted densities to reality, we derive a set of equations which include n lines (where n is the total number of observed points) and m (where m: the numbers of predefined groups) unknown parameters. A least-square estimation will allow us to determine the density law of different groups and components in the Galaxy. The output from our software, which can be used in many research fields, will also give out the systematic error between the model and the observation by a Bayes rule.

Chen, B.↗

Solar forecasting using machine learned cloudiness classification

Methods and systems for predicting irradiance include learning a classification model using unsupervised learning based on historical irradiance data. The classification model is updated using supervised learning based on an association between known cloudiness states and historical weather data. A cloudiness state is predicted based on forecasted weather data. An irradiance is predicted using a regression model associated with the cloudiness state.

Hamann, Hendrik F.↗

Automated integration gate selection for Gaussian mixture model pulse shape discrimination

Pulse shapes differ between neutron and gamma particles when measured with detector devices employing pulse shape discriminating (PSD) scintillators. Digitized waveforms can be used in detection systems to perform pulse shape discrimination for this application. Prior Gaussian Mixture Model (GMM) methods require access to the pulse full-waveform. Reducing the waveform to a smaller set of combined samples reduces computational cost while affecting PSD performance. In this work, we develop a method for selecting the best performing combination of integration gates, or contiguous summed segments of the digitized pulse for PSD. The method uses a discrimination score based on the GMM PSD approach. Furthermore, the final selection is performed using Bayesian Optimization. PSD detection results are compared with varying numbers of selected gates on time-of-flight (TOF) data. This method can be used to fully automate the selection of gates in an unsupervised (without ground truth labels) setting.

42 ENGINEERING↗

AICCA: AI-Driven Cloud Classification Atlas

Clouds play an important role in the Earth’s energy budget, and their behavior is one of the largest uncertainties in future climate projections. Satellite observations should help in understanding cloud responses, but decades and petabytes of multispectral cloud imagery have to date received only limited use. This study describes a new analysis approach that reduces the dimensionality of satellite cloud observations by grouping them via a novel automated, unsupervised cloud classification technique based on a convolutional autoencoder, an artificial intelligence (AI) method good at identifying patterns in spatial data. Our technique combines a rotation-invariant autoencoder and hierarchical agglomerative clustering to generate cloud clusters that capture meaningful distinctions among cloud textures, using only raw multispectral imagery as input. Cloud classes are therefore defined based on spectral properties and spatial textures without reliance on location, time/season, derived physical properties, or pre-designated class definitions. We use this approach to generate a unique new cloud dataset, the AI-driven cloud classification atlas (AICCA), which clusters 22 years of ocean images from the Moderate Resolution Imaging Spectroradiometer (MODIS) on NASA’s Aqua and Terra instruments—198 million patches, each roughly 100 km × 100 km (128 × 128 pixels)—into 42 AI-generated cloud classes, a number determined via a newly-developed stability protocol that we use to maximize richness of information while ensuring stable groupings of patches. AICCA thereby translates 801 TB of satellite images into 54.2 GB of class labels and cloud top and optical properties, a reduction by a factor of 15,000. The 42 AICCA classes produce meaningful spatio-temporal and physical distinctions and capture a greater variety of cloud types than do the nine International Satellite Cloud Climatology Project (ISCCP) categories—for example, multiple textures in the stratocumulus decks along the West coasts of North and South America. We conclude that our methodology has explanatory power, capturing regionally unique cloud classes and providing rich but tractable information for global analysis. AICCA delivers the information from multi-spectral images in a compact form, enables data-driven diagnosis of patterns of cloud organization, provides insight into cloud evolution on timescales of hours to decades, and helps democratize climate research by facilitating access to core data.

97 MATHEMATICS AND COMPUTING↗