Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “unsupervised method”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

NMF-Based Anomaly Detection in CMS 2D Tracking Occupancy Histograms

The CMS experiment relies on Data Quality Monitoring (DQM) to ensure that recorded collision data are suitable for physics analysis. During LHC Run 3, each run contains many lumisections and tracking monitoring elements, making offline inspection challenging, especially for localized detector effects that may appear only for short periods of time. This poster presents an unsupervised machine-learning approach to identify anomalous lumisections in CMS tracking occupancy histograms using Non-Negative Matrix Factorization (NMF). The workflow uses offline CMS DQMIO tracking histograms retrieved with the CMS DIALS API and organized as two-dimensional occupancy maps for each lumisection. After selecting stable lumisections, the occupancy maps are normalized and arranged into a non-negative data matrix. The NMF model learns a compact set of basis patterns describing normal tracking occupancy. Each lumisection is then reconstructed from these learned components, and the reconstruction error is used as an anomaly score. Large residuals indicate occupancy patterns that deviate from normal detector behavior and are flagged for further inspection. This NMF-based approach provides a fast and interpretable way to flag lumisections whose tracking occupancy patterns differ from normal detector behavior. Preliminary studies show sensitivity to known tracking anomalies, and ongoing work is focused on validating the method across additional Run 3 Pixel and Strip detector issues.

Rodríguez Ramos, Iliomar [Puerto Rico U., Mayaguez↗

Wind Turbine Gearbox Failure Detection Through Cumulative Sum of Multivariate Time Series Data

The wind energy industry is continuously improving their operational and maintenance practice for reducing the levelized costs of energy. Anticipating failures in wind turbines enables early warnings and timely intervention, so that the costly corrective maintenance can be prevented to the largest extent possible. It also avoids production loss owing to prolonged unavailability. One critical element allowing early warning is the ability to accumulate small-magnitude symptoms resulting from the gradual degradation of wind turbine systems. Inspired by the cumulative sum control chart method, this study reports the development of a wind turbine failure detection method with such early warning capability. Specifically, the following key questions are addressed: what fault signals to accumulate, how long to accumulate, what offset to use, and how to set the alarm-triggering control limit. We apply the proposed approach to 2 years’ worth of Supervisory Control and Data Acquisition data recorded from five wind turbines. We focus our analysis on gearbox failure detection, in which the proposed approach demonstrates its ability to anticipate failure events with a good lead time.

17 WIND ENERGY↗

Model and remote-sensing-guided experimental design and hypothesis generation for monitoring snow-soil–plant interactions

In this study, we develop a machine-learning (ML)-enabled strategy for selecting hillslope-scale ecohydrological monitoring sites within snow-dominated mountainous watersheds, with a particular focus on snow-soil–plant interactions. Data layers rely on spatial data layers from both remote sensing and hydrological model simulations. Specifically, a Landsat-based foresummer drought sensitivity index is used to define the dependency of the annual peak plant productivity on the Palmer drought severity index in the early growing season. Hydrological simulations provide the spatiotemporal dynamics of near-surface soil moisture and snow depth. In this framework, a regression analysis identifies the key hydrological variables relevant to the spatial heterogeneity of drought sensitivity. We then apply unsupervised clustering to these key variables, using the Gaussian mixture model, to group hillslopes into several zones that have divergent relationships regarding soil moisture, snow dynamics, and drought sensitivity. Using the datasets collected in the East River Watershed (Crested Butte, Colorado, United States), results show that drought sensitivity is significantly correlated with model-derived soil moisture and snow-free timing over space and time. The relationship is, however, non-linear, such that the correlation decreases above a threshold elevation and in a heavy snow year due to large snowpacks, lateral flow, and soil storage limitations. Clustering is then able to define the zones that have high or low sensitivity to drought, as well as the mid-elevation regions where sensitivity is associated with the topographic aspect and net potential radiation. In addition, the algorithm identifies the most representative hillslopes with road/trail access within each zone for installing monitoring sites. Our method also aims to significantly increase the use of ML and model-simulation results to guide critical zone and watershed monitoring activities.

54 ENVIRONMENTAL SCIENCES↗

Analyzing acoustic emission data to identify cracking modes in cement paste using an artificial neural network

This research is focused on the identification of cracking mechanisms for cement paste using acoustic emission data, recorded from compression and notched four-point bending tests. A procedure is developed for analyzing the data by employing an agglomerative hierarchical clustering method, an artificial neural network, and a ray-tracing source location algorithm. An agglomerative hierarchical clustering method is utilized to cluster the AE data from a compression test using frequency-dependent features. A neural network is trained using the compression test data and applied to the AE data emitted during the four-point bending test. The clustered data from the four-point bending test is localized using a ray-tracing algorithm. Based on the occurrence and locations of the clustered events and signal feature analyses, potential cracking mechanisms are identified and assigned.

36 MATERIALS SCIENCE↗

Deep learning to estimate permeability using geophysical data

Time-lapse electrical resistivity tomography (ERT) is a popular geophysical method to estimate three-dimensional (3D) permeability fields from electrical potential difference measurements. Traditional inversion and data assimilation methods are used to ingest this ERT data into hydrogeophysical models to estimate permeability. Due to ill-posedness and the curse of dimensionality, existing inversion strategies provide poor estimates and low resolution of the 3D permeability field. Recent advances in deep learning provide us with powerful algorithms to overcome this challenge. This paper presents a deep learning (DL) framework to estimate the 3D subsurface permeability from time-lapse ERT data. To test the feasibility of the proposed framework, we train DL-enabled inverse models on simulation data. Each measurement in both synthetic and field data is standardized by removing the mean and scaling the time-series to unit variance. This pre-processing step is necessary to bring simulation data closer to field observations. Subsurface process models based on hydrogeophysics are used to generate this synthetic data. Training performed on limited simulation data resulted in the DL model over-fitting. An advanced data augmentation based on mixup is implemented to generate additional training samples to overcome this issue. This mixup technique creates weakly labeled (low-fidelity) samples from strongly labeled (high-fidelity) data. The weakly labeled training data is then used to develop DL-enabled inverse models and reduce over-fitting. As both time-lapse ERT (1133048 features/realization) and 3D permeability (585453 features/realization) data samples are from a high-dimensional space, principal component analysis (PCA) is employed to reduce dimensionality. Encoded ERT and encoded permeability are generated using the trained PCA estimators. A deep neural network is then trained to map the encoded ERT to encoded permeability. This mixup training and unsupervised learning allowed us to build a fast and reasonably accurate DL-based inverse model under limited simulation data. Results show that proposed weak supervised learning can capture salient spatial features in the 3D permeability field. Quantitatively, the average mean squared error (in terms of the natural log) on the strongly labeled training, validation, and test datasets is less than 0.5. The R 2 -score (global metric) is greater than 0.75, and the percent error in each cell (local metric) is less than 10%. Finally, an added benefit in terms of computational cost is that the proposed DL-based inverse model is at least O(10 4 ) times faster than running a forward model once it is trained. Data generation, DL model training, and hyperparameter tuning to identify optimal neural network architectures utilized high-performance computing resources while the DL inference is performed on a standard laptop. Approximately, O(10 5 ) processor hours are used for generating data and DL tuning and training. We acknowledge that the data generation and DL model development are expensive. But once a DL model is trained, it can be re-used for inversion rapidly for the given system, with set physics and domain. Note that traditional inversion may require multiple forward model simulations (e.g., in the order of 10 to 1000), which are very expensive. This computational savings ≈ O(10 5 ) – O(10 7 )) makes the proposed DL-based inverse model attractive for subsurface imaging and real-time ERT monitoring applications due to fast and yet reasonably accurate estimations of permeability field.

58 GEOSCIENCES↗

Use of machine learning to analyze chemistry card sort tasks

Education researchers are deeply interested in understanding the way students organize their knowledge. Card sort tasks, which require students to group concepts, are one mechanism to infer a student’s organizational strategy. However, the limited resolution of card sort tasks means they necessarily miss some of the nuance in a student’s strategy. Here in this work, we propose new machine learning strategies that leverage a potentially richer source of student thinking: free-form written language justifications associated with student sorts. Using data from a university chemistry card sort task, we use vectorized representations of language and unsupervised learning techniques to generate qualitatively interpretable clusters, which can provide unique insight in how students organize their knowledge. We compared these to machine learning analysis of the students’ sorts themselves. Machine learning-generated clusters revealed different organizational strategies than those built into the task; for example, sorts by difficulty or even discipline. There were also many more categories generated by machine learning for what we would identify as more novice-like sorts and justifications than originally built into the task, suggesting students’ organizational strategies converge when they become more expert-like. Finally, we learned that categories generated by machine learning for students’ justifications did not always match the categories for their sorts, and these cases highlight the need for future research on students’ organizational strategies, both manually and aided by machine learning. In sum, the use of machine learning to analyze results from a card sort task has helped us gain a more nuanced understanding of students’ expertise, and demonstrates a promising tool to add to existing analytic methods for card sorts.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

TransformerG2G: Adaptive time-stepping for learning temporal graph embeddings using transformers

Dynamic graph embedding has emerged as a very effective technique for addressing diverse temporal graph analytic tasks (i.e., link prediction, node classification, recommender systems, anomaly detection, and graph generation) in various applications. Such temporal graphs exhibit heterogeneous transient dynamics, varying time intervals, and highly evolving node features throughout their evolution. Hence, incorporating long-range dependencies from the historical graph context plays a crucial role in accurately learning their temporal dynamics. In this paper, we develop a graph embedding model with uncertainty quantification, TransformerG2G, by exploiting the advanced transformer encoder to first learn intermediate node representations from its current state (t) and previous context (over timestamps [t–1,t–l], l is the length of context). Moreover, we employ two projection layers to generate lower-dimensional multivariate Gaussian distributions as each node's latent embedding at timestamp t. We consider diverse benchmarks with varying levels of "novelty" as measured by the TEA (Temporal Edge Appearance) plots. Here, our experiments demonstrate that the proposed TransformerG2G model outperforms conventional multi-step methods and our prior work (DynG2G) in terms of both link prediction accuracy and computational efficiency, especially for high degree of novelty. Furthermore, the learned time-dependent attention weights across multiple graph snapshots reveal the development of an automatic adaptive time stepping enabled by the transformer. Importantly, by examining the attention weights, we can uncover temporal dependencies, identify influential elements, and gain insights into the complex interactions within the graph structure. For example, we identified a strong correlation between attention weights and node degree at the various stages of the graph topology evolution.

97 MATHEMATICS AND COMPUTING↗

Deep Learning for Spectroscopic X-ray Nano-Imaging Denoising

Synchrotron transmission X-ray microscopy with absorption near edge structure (TXM-XANES) is a powerful tool for investigating the structure and composition of materials at nano- to meso-scales. It is, however, often challenged by high levels of noise that obscure critical details at the single-pixel level. To address this issue, a deep learning-based algorithm is developed for suppressing the image noise, grounded in self-supervised learning principles. In contrast to traditional image denoising methods, this approach successfully enhances the visibility of fine details while significantly reducing the noise in the X-ray images. Through this advancement, the potential of the approach for improving the accuracy and interpretability of the TXM-XANES data is demonstrated, thereby enabling more precise detection of nanoscale phenomena such as inhomogeneous cation redox and metal segregation in battery cathode materials. This technique offers an effective new avenue for harnessing the full potential of synchrotron TXM-XANES imaging, paving the way for a range of exciting new studies in materials science and beyond.

36 MATERIALS SCIENCE↗

Learning Canonical Embeddings for Unsupervised Shape Correspondence With Locally Linear Transformations

We present a new approach to unsupervised shape correspondence learning between pairs of point clouds. We make the first attempt to adapt the classical locally linear embedding algorithm (LLE)-originally designed for nonlinear dimensionality reduction-for shape correspondence. The key idea is to find dense correspondences between shapes by first obtaining high-dimensional neighborhood-preserving embeddings of low-dimensional point clouds and subsequently aligning the source and target embeddings using locally linear transformations. We demonstrate that learning the embedding using a new LLE-inspired point cloud reconstruction objective results in accurate shape correspondences. More specifically, the approach comprises an end-to-end learnable framework of extracting high-dimensional neighborhood-preserving embeddings, estimating locally linear transformations in the embedding space, and reconstructing shapes via divergence measure-based alignment of probability density functions built over reconstructed and target shapes. Our approach enforces embeddings of shapes in correspondence to lie in the same universal/canonical embedding space, which eventually helps regularize the learning process and leads to a simple nearest neighbors approach between shape embeddings for finding reliable correspondences. Comprehensive experiments show that the new method makes noticeable improvements over state-of-the-art approaches on standard shape correspondence benchmark datasets covering both human and nonhuman shapes.

deformation↗

A Computational Theory for the Emergence of Grammatical Categories in Cortical Dynamics

A general agreement in psycholinguistics claims that syntax and meaning are unified precisely and very quickly during online sentence processing. Although several theories have advanced arguments regarding the neurocomputational bases of this phenomenon, we argue that these theories could potentially benefit by including neurophysiological data concerning cortical dynamics constraints in brain tissue. In addition, some theories promote the integration of complex optimization methods in neural tissue. In this paper we attempt to fill these gaps introducing a computational model inspired in the dynamics of cortical tissue. In our modeling approach, proximal afferent dendrites produce stochastic cellular activations, while distal dendritic branches–on the other hand–contribute independently to somatic depolarization by means of dendritic spikes, and finally, prediction failures produce massive firing events preventing formation of sparse distributed representations. The model presented in this paper combines semantic and coarse-grained syntactic constraints for each word in a sentence context until grammatically related word function discrimination emerges spontaneously by the sole correlation of lexical information from different sources without applying complex optimization methods. By means of support vector machine techniques, we show that the sparse activation features returned by our approach are well suited—bootstrapping from the features returned by Word Embedding mechanisms—to accomplish grammatical function classification of individual words in a sentence. In this way we develop a biologically guided computational explanation for linguistically relevant unification processes in cortex which connects psycholinguistics to neurobiological accounts of language. We also claim that the computational hypotheses established in this research could foster future work on biologically-inspired learning algorithms for natural language processing applications.

59 BASIC BIOLOGICAL SCIENCES↗

Mapping structural heterogeneity at the nanoscale with scanning nano-structure electron microscopy (SNEM)

Here, in this work, we explore the use of scanning electron diffraction (also known as 4D-STEM) coupled with electron atomic pair distribution function analysis (ePDF) to understand the local order (structure and chemistry) as a function of position in a complex multicomponent system, a hot rolled, Ni-encapsulated, Zr 65 Cu 17.5 Ni 10 Al 7.5 bulk metallic glass (BMG), with a spatial resolution of 3 nm. We show that it is possible to gain insight into the chemistry and chemical clustering/ordering tendency in different regions of the sample, including in the vicinity of nano-scale crystallites that are identified from virtual dark field images and in heavily deformed regions at the edge of the BMG. In addition to simpler analysis, unsupervised machine learning was used to extract partial PDFs from the material, modeled as a quasi-binary alloy, and map them in space. These maps allowed key insights not only into the local average composition, as validated by EELS, but also a unique insight into chemical short-range ordering tendencies in different regions of the sample during formation. The experiments are straightforward and rapid and, unlike spectroscopic measurements, don’t require energy filters on the instrument. We spatially map different quantities of interest (QoI’s), defined as scalars that can be computed directly from positions and widths of ePDF peaks or parameters refined from fits to the patterns. We developed a flexible and rapid data reduction and analysis software framework that allows experimenters to rapidly explore images of the sample on the basis of different QoI’s. The power and flexibility of this approach are explored and described in detail. Because of the fact that we are getting spatially resolved images of the nanoscale structure obtained from ePDFs we call this approach scanning nano-structure electron microscopy (SNEM), and we believe that it will be powerful and useful extension of current 4D-STEM methods.

36 MATERIALS SCIENCE↗

Optimal dimensionality selection for independent component analysis of transcriptomic data

Independent component analysis is an unsupervised machine learning algorithm that separates a set of mixed signals into a set of statistically independent source signals. Applied to high-quality gene expression datasets, independent component analysis effectively reveals both the source signals of the transcriptome as co-regulated gene sets, and the activity levels of the underlying regulators across diverse experimental conditions. Two major variables that affect the final gene sets are the diversity of the expression profiles contained in the underlying data, and the user-defined number of independent components, or dimensionality, to compute. Availability of high-quality transcriptomic datasets has grown exponentially as high-throughput technologies have advanced; however, optimal dimensionality selection remains an open question. We computed independent components across a range of dimensionalities for four gene expression datasets with varying dimensions (both in terms of number of genes and number of samples). We computed the correlation between independent components across different dimensionalities to understand how the overall structure evolves as the number of user-defined components increases. We then measured how well the resulting gene clusters reflected known regulatory mechanisms, and developed a set of metrics to assess the accuracy of the decomposition at a given dimension. We found that over-decomposition results in many independent components dominated by a single gene, whereas under-decomposition results in independent components that poorly capture the known regulatory structure. From these results, we developed a new method, called OptICA, for finding the optimal dimensionality that controls for both over- and under-decomposition. Specifically, OptICA selects the highest dimension that produces a low number of components that are dominated by a single gene. We show that OptICA outperforms two previously proposed methods for selecting the number of independent components across four transcriptomic databases of varying sizes. OptICA avoids both over-decomposition and under-decomposition of transcriptomic datasets resulting in the best representation of the organism’s underlying transcriptional regulatory network.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Detecting thermodynamic phase transition via explainable machine learning of photoemission spectroscopy

Identifying thermodynamic signatures of electronic phases, such as superconductivity, is challenging in low-dimensional materials due to strong fluctuations and low probing volume. Spectroscopic methods are often used to identify new bulk phases, but their main measurable quantity—electronic energy gaps—is no longer an effective order parameter in low-dimensional and fluctuating systems. Combining angle-resolved photoemission with a domain-adversarial neural network, we report a data-driven method to identify thermodynamic phase transitions solely based on single-particle spectra. We demonstrate 97.6% accuracy in cuprate superconductor Bi 2 Sr 2 CaCu 2 O 8+δ with strong superconducting fluctuations. This model notably compensates for the scarcity of experimental data by leveraging virtually inexhaustible simulated data. Further, its explainability reveals the crucial role of in-gap spectral weight in detecting phase fluctuations and thermodynamic transitions. Our work pinpoints the spectroscopic signatures of fluctuating orders and enables using spectroscopy for machine-learning-assisted material discovery for low-dimensional and strong coupling systems.

2D materials↗

Dense autoencoders, clustering techniques, and semi-supervised learning for HPGe $γ$-spectra

Classifying high-resolution gamma spectra by their isotopic content is an essential task in nuclear forensics and other applications. Traditional analysis methods are often time-intensive, but machine learning (ML) may help analysts quickly process many spectra. Such methods tend to rely on abundant, well-labeled data for training. Historical gamma data exists in various fields but is not uniformly useful for supervised ML due to inconsistent labeling. Here, to address some of these challenges, we present a method to classify and organize unlabeled data from high-purity germanium detectors using an autoencoding neural network (autoencoder). We trained dense autoencoders to compress gamma data into latent representations that enable efficient data characterization. By clustering the encoded spectra or lower-dimensional mappings of them, we identified and removed portions of over-abundant data categories, resulting in a more balanced dataset and improved autoencoder performance. This encoding and clustering pipeline also enabled the organization of spectra into self-consistent categories. Finally, we found that encoded representations showed potential as inputs for semi-supervised learning of nuclide identification (NID) labels, achieving an average F1 score of 0.85 ± 0.03 when mapping encodings to a set of 65 isotope labels.

Autoencoders↗

Assessing the Application of a Genomic Network Analysis in Population Ecology: Inferring Patterns of Dispersal and Geographic Structure in the Emerging Pathogen, Coccidioides

A challenge in population ecology studies is identifying how to best group individuals into populations, especially when individual origin is unknown. Machine learning has improved upon traditional methods of identifying population structure and is more efficient at handling large, complex datasets. We demonstrate the applicability of a machine learning method to identify hierarchical population structure in an emerging pathogen, Coccidioides spp., the causative agent of Valley fever. We compared the network clusters to structure identified by traditional tools as a validation of the network performance. We used publicly available whole-genome data for 48 C. immitis and 102 C. posadasii, resulting in 168,211 genome-wide SNPs among the two species. The network analysis grouped samples into populations comparable to the literature for these species but also identified fine-scale geographic structure and travel-associated cases not reported thus far. Exploring different resolutions in the network made it easy to identify unique genotypes specific to California and possibly Nevada, as well as Phoenix- and Tucson-acquired infections in non-endemic areas, regardless of reported travel history. The present study provides a promising example of how a ML-based network analysis can improve our ability to understand pathogen ecology, group cases into populations and infer travel-associated infections.

59 BASIC BIOLOGICAL SCIENCES↗

“Thought I’d Share First” and Other Conspiracy Theory Tweets from the COVID-19 Infodemic: Exploratory Study

Background: The COVID-19 outbreak has left many people isolated within their homes; these people are turning to social media for news and social connection, which leaves them vulnerable to believing and sharing misinformation. Health-related misinformation threatens adherence to public health messaging, and monitoring its spread on social media is critical to understanding the evolution of ideas that have potentially negative public health impacts. Objective: The aim of this study is to use Twitter data to explore methods to characterize and classify four COVID-19 conspiracy theories and to provide context for each of these conspiracy theories through the first 5 months of the pandemic. Methods: We began with a corpus of COVID-19 tweets (approximately 120 million) spanning late January to early May 2020. We first filtered tweets using regular expressions (n=1.8 million) and used random forest classification models to identify tweets related to four conspiracy theories. Our classified data sets were then used in downstream sentiment analysis and dynamic topic modeling to characterize the linguistic features of COVID-19 conspiracy theories as they evolve over time. Results: Analysis using model-labeled data was beneficial for increasing the proportion of data matching misinformation indicators. Random forest classifier metrics varied across the four conspiracy theories considered (F1 scores between 0.347 and 0.857); this performance increased as the given conspiracy theory was more narrowly defined. We showed that misinformation tweets demonstrate more negative sentiment when compared to non-misinformation tweets and that theories evolve over time, incorporating details from unrelated conspiracy theories as well as real-world events. Conclusions: Although we focus here on health-related misinformation, this combination of approaches is not specific to public health and is valuable for characterizing misinformation in general, which is an important first step in creating targeted messaging to counteract its spread. Initial messaging should aim to preempt generalized misinformation before it becomes widespread, while later messaging will

5g↗

Galactic ArchaeoLogIcaL ExcavatiOns (GALILEO) II. t-SNE portrait of local fossil relics and structures

Based on high-quality Apache Point Observatory Galactic Evolution Experiment (APOGEE) DR17 and Gaia DR3 data for 1742 red giants stars within 5 kpc of the Sun and not rotating with the Galactic disk (V φ < 100 km s -1 ), we used the nonlinear technique of unsupervised analysis t-Distributed Stochastic Neighbor Embedding (t-SNE) to detect coherent structures in the space of ten chemical-abundance ratios: [Fe/H], [O/Fe], [Mg/Fe], [Si/Fe], [Ca/Fe], [C/Fe], [N/Fe], [Al/Fe], [Mn/Fe], and [Ni/Fe]. Additionally, we obtained orbital parameters for each star using the nonaxisymmetric gravitational potential GravPot16. Seven structures are detected, including Splash, Gaia-Sausage-Enceladus (GSE), the high-α heated-disk population, N-C-O peculiar stars, and inner disk-like stars, plus two other groups that did not match anything previously reported in the literature, here named Galileo 5 and Galileo 6 (G5 and G6). These two groups overlap with Splash in [Fe/H], with G5 having a lower metallicity than G6, and they are both between GSE and Splash in the [Mg/Mn] versus [Al/Fe] plane, with G5 being in the α-rich in situ locus and G6 on the border of the α-poor in situ one. Nonetheless, their low [Ni/Fe] hints at a possible ex situ origin. Their orbital energy distributions are between Splash and GSE, with G5 being slightly more energetic than G6. We verified the robustness of all the obtained groups by exploring a large range of t-SNE parameters, applying it to various subsets of data, and also measuring the effect of abundance errors through Monte Carlo tests.

79 ASTRONOMY AND ASTROPHYSICS↗