Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Unsupervised machine learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

Natural Language Processing Methods for Air Traffic Management Text and Speech Data

This presentation discusses two efforts of the NARI AI/ML Intern team during the Fall 2021 OSTEM Internship term. For Letters of Agreement (LoA), we have studied how LoAs are structured and explored the question ‘What is an LoA constraint?’ To do this, our approach is data-driven, iterative, and assisted by machine learning when available. In this presentation, we will walk through our tasks of manually scanning through documents, performing a preliminary entity labelling task, and our unsupervised analysis on LoA procedures sections. After this research phase, we define the smallest constraint unit in an LoA, and start to perform entity extraction. Looking towards constraint extraction, we are also exploring the use of a one-class support vector machine (OneClassSVM) model to identify patterns within the data. The second effort of our team this term is focused on Air Traffic Control System Command Center (ATCSCC) advisory meetings, and the subsequent advisory documents that get published from their content. These advisory documents are important to give readily accessible summaries of daily operations, so that data centers, airline officials, and other stakeholders can easily understand the context of these meetings in real time. In applying machine learning to this scenario, two natural language processing tasks are used. First is developing machine learning models to convert the meeting speech data into text. With this text, use of extractive and abstractive text summarization models are used to automatically generate preliminary versions of the advisory documents.

Natural Language Processing↗

Characterizing 4-string contact interaction using machine learning

Abstract The geometry of 4-string contact interaction of closed string field theory is characterized using machine learning. We obtain Strebel quadratic differentials on 4-punctured spheres as a neural network by performing unsupervised learning with a custom-built loss function. This allows us to solve for local coordinates and compute their associated mapping radii numerically. We also train a neural network distinguishing vertex from Feynman region. As a check, 4-tachyon contact term in the tachyon potential is computed and a good agreement with the results in the literature is observed. We argue that our algorithm is manifestly independent of number of punctures and scaling it to characterize the geometry ofn-string contact interaction is feasible.

Physics↗

Model-agnostic search for dijet resonances with anomalous jet substructure in proton–proton collisions at $\sqrt{s}$ = 13 TeV

This paper presents a model-agnostic search for narrow resonances in the dijet final state in the mass range 1.8-6 TeV. The signal is assumed to produce jets with substructure atypical of jets initiated by light quarks or gluons, with minimal additional assumptions. Search regions are obtained by utilizing multivariate machine-learning methods to select jets with anomalous substructure. A collection of complementary anomaly detection methods - based on unsupervised, weakly supervised, and semisupervised algorithms - are used in order to maximize the sensitivity to unknown new physics signatures. These algorithms are applied to data corresponding to an integrated luminosity of 138 fb -1 , recorded by the CMS experiment at the LHC, at a center-of-mass energy of 13 TeV. No significant excesses above background expectations are seen. Exclusion limits are derived on the production cross section of benchmark signal models varying in resonance mass, jet mass, and jet substructure. Many of these signatures have not been previously sought, making several of the limits reported on the corresponding benchmark models the first ever. When compared to benchmark inclusive and substructure-based search strategies, the anomaly detection methods are found to significantly enhance the sensitivity to a variety of models.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Parametric Analysis of a Hover Test Vehicle using Advanced Test Generation and Data Analysis

Large complex aerospace systems are generally validated in regions local to anticipated operating points rather than through characterization of the entire feasible operational envelope of the system. This is due to the large parameter space, and complex, highly coupled nonlinear nature of the different systems that contribute to the performance of the aerospace system. We have addressed the factors deterring such an analysis by applying a combination of technologies to the area of flight envelop assessment. We utilize n-factor (2,3) combinatorial parameter variations to limit the number of cases, but still explore important interactions in the parameter space in a systematic fashion. The data generated is automatically analyzed through a combination of unsupervised learning using a Bayesian multivariate clustering technique (AutoBayes) and supervised learning of critical parameter ranges using the machine-learning tool TAR3, a treatment learner. Covariance analysis with scatter plots and likelihood contours are used to visualize correlations between simulation parameters and simulation results, a task that requires tool support, especially for large and complex models. We present results of simulation experiments for a cold-gas-powered hover test vehicle.

Gundy-Burlet, Karen↗

Machine-learning-based automatic small-angle measurement between planar surfaces in interferometer images: A 2D multilayer Laue lenses case

Here, we report a new machine-learning-based approach to automatically measure the small angle between multiple planar surfaces characterized by white light interferometers. By applying an unsupervised clustering algorithm, DBSCAN (Density-Based Spatial Clustering of Applications with Noise), the multiple surfaces in an interferometer image are automatically identified as distinct surfaces. The angles between every two surfaces are then calculated through the surface fitting. This method can be applied to multiple surfaces regardless of their shapes and locations and significantly simplifies the angle measurement procedure. Using the developed method, we have demonstrated a quick and precise angle measurement for the alignment of 2D Multilayer Laue Lenses (MLLs) for the development of high-resolution x-ray microscopy. This automatic, accurate, and robust small-angle measurement method is compatible with widely used white light interferometers and can be further applied to other metrology applications of interferometer results.

36 MATERIALS SCIENCE↗

Exploratory analysis of machine learning techniques in the Nevada geothermal play fairway analysis

Play fairway analysis (PFA) is commonly used to generate geothermal potential maps and guide exploration studies, with a particular focus on locating and characterizing blind geothermal systems. This study evaluates the application of machine learning techniques to PFA in the Great Basin region of Nevada. Following the evaluation of various techniques, we identified two approaches to PFA that produced promising results, 1) supervised Bayesian probabilistic neural networks to generate geothermal potential maps with confidence intervals, and 2) unsupervised principal component analysis paired with k-means clustering to generate both cluster maps to help identify spatial patterns, as well as new combined feature inputs. We applied these techniques to perform a comparative analysis between two principal sets of geological and geophysical features related to permeability and heat and a set of positive (known geothermal resources) and negative training sites (known drill sites with unsuitable geothermal conditions). We found that these methods constrain previously unrecognized feature controls on geothermal favorability, many of which are spatially organized within the extent of cluster groups and the major structural-hydrologic domains of the study area. Furthermore, we utilized exploratory unsupervised modeling to highlight spatial relationships between input data and predictive output results of our supervised modeling. As a result, we demonstrate how our models compare to the previous Nevada PFA and how the rapid insights these machine learning techniques offer may support future assessments of both known and undiscovered blind geothermal systems in the Great Basin region of Nevada and beyond.

15 GEOTHERMAL ENERGY↗

Detecting thermodynamic phase transition via explainable machine learning of photoemission spectroscopy

Identifying thermodynamic signatures of electronic phases, such as superconductivity, is challenging in low-dimensional materials due to strong fluctuations and low probing volume. Spectroscopic methods are often used to identify new bulk phases, but their main measurable quantity—electronic energy gaps—is no longer an effective order parameter in low-dimensional and fluctuating systems. Combining angle-resolved photoemission with a domain-adversarial neural network, we report a data-driven method to identify thermodynamic phase transitions solely based on single-particle spectra. We demonstrate 97.6% accuracy in cuprate superconductor Bi 2 Sr 2 CaCu 2 O 8+δ with strong superconducting fluctuations. This model notably compensates for the scarcity of experimental data by leveraging virtually inexhaustible simulated data. Further, its explainability reveals the crucial role of in-gap spectral weight in detecting phase fluctuations and thermodynamic transitions. Our work pinpoints the spectroscopic signatures of fluctuating orders and enables using spectroscopy for machine-learning-assisted material discovery for low-dimensional and strong coupling systems.

2D materials↗

3-D Geologic Controls of Hydrothermal Fluid Flow at Brady Geothermal Field, Nevada using PCA

In many hydrothermal systems, fracture permeability along faults provides pathways for groundwater to transport heat from depth. Faulting generates a range of deformation styles that cross-cut heterogeneous geology, resulting in complex patterns of permeability, porosity, and hydraulic conductivity. Vertical connectivity (a through going network of permeable areas that allows advection of heat from depth to the shallow subsurface) is rare and is confined to relatively small volumes that have highly variable spatial distribution. This local compartmentalization of connectivity represents a significant challenge to understanding hydrothermal circulation and for exploring, developing, and managing hydrothermal resources. Here, we present an evaluation of the geologic characteristics that control this compartmentalization in hydrothermal systems through 3-D analysis of the Brady geothermal field in western Nevada. A published 3-D geologic map of the Brady area is used as a basis to develop structural and geological variables that are hypothesized to control or effect permeability or connectivity. The 3-D distribution of these variables is compared to the distribution of productive and non-productive fluid flow intervals along production wells and non-productive wells via principal component analysis (PCA). This comparison elucidates which geologic and structural variables are most closely associated with productive fluid flow intervals. Results indicate that production intervals at Brady are located: (1) within or near to known and stress-loaded macro-scale faults, and (2) in areas of high fault and fracture density. This submission includes the published journal article detailing this work, the published 3-D geologic map of the Brady Geothermal Area used as a basis to develop structural and geological variables that are hypothesized to control or effect permeability or connectivity, 3-D well data, along which geologic data were sampled for PCA analyses, and associated metadata file. This work was done using existing R programs.

15 GEOTHERMAL ENERGY↗

Superconducting Radio-frequency Cavity Fault Classification Using Machine Learning at Jefferson Laboratory

We report on the development of machine learning models for classifying C100 superconducting radiofrequency (SRF) cavity faults in the Continuous Electron Beam Accelerator Facility (CEBAF) at Jefferson Lab. Of the 418 SRF cavities in CEBAF, 96 are designed with a digital low-level RF system configured such that a cavity fault triggers recordings of RF signals for each of eight cavities in the cryomodule. Subject matter experts analyze the collected time-series data and identify which of the eight cavities faulted first and classify the type of fault. This information is used to find trends and strategically deploy mitigations to problematic cryomodules. However, manually labeling the data is laborious and time-consuming. By leveraging machine learning, near real-time - rather than postmortem - identification of the offending cavity and classification of the fault type has been implemented. We discuss the performance of the machine learning models during a recent physics run. We also discuss efforts for further insights into fault types through unsupervised learning techniques and present preliminary work on cavity and fault prediction using data collected prior to a failure event.

Tennant, C. D.↗

Monitoring Fracture Hydromechanical Evolution in the Lab and Field Using Unsupervised Metric Learning

Fractures evolve in time through thermal‐hydraulic‐mechanical‐chemical (THMC) processes that alter their long‐range hydraulic transport properties and modify subsurface behavior and activities. The location of subsurface fractures makes it necessary to use remote sensing techniques such as passive or active seismic monitoring for fracture characterization. In this paper, we develop a machine learning approach to monitor the evolution of fracture properties using passive seismic sources in a laboratory setting and using active seismic monitoring from the Sanford Underground Research Facility in Lead, South Dakota, at a depth of 1.25 km in amphibolite rock during stimulation of natural fractures as well as during induced fracturing. The unsupervised metric learning technique applies tandem neural networks (twin (Siamese) or triplet) with contrastive loss and adaptive margins to track slowly varying systems for which class or similarity labels are not available. The approach adopts locality‐sensitive hashing to divide time‐ordered contiguous data into an arbitrary number of pseudo‐classes. Contrastive‐loss training with many hash bins generates an evolving latent‐space trajectory. This approach enables unsupervised metric learning for seismic data stacks under the condition of contiguous state sampling and slowly varying fracture properties. The displacement discontinuity theory provides a mechanistic foundation for the fracture‐dependent trajectories that are related to relaxation of fractures with time‐dependent specific stiffness responding to changes in stress or fluid saturation.

02 PETROLEUM↗

Characterization of Acoustic Emissions From Analogue Rocks Using Sparse Regression‐DMDc

Abstract Moisture loss in rock is known to generate acoustic emissions (AE). Phenomena that result in AE during drying are related to the movement of fluids through the pores and induced‐cracks that arise from differential mineral shrinkage, especially in clay‐bearing rock. AE from the movement of fluids occurs from the reconfiguration of fluid interfaces during drying, while AE from mineral shrinkage involves the debonding within or between minerals. Here, analogue rock samples were used to examine the differences in the AE signatures when one or both AE source‐types are present. An unsupervised sparse regression model, Dynamic Mode Decomposition with control, that extends Dynamic Mode Decomposition is used to characterize the AE signals recorded during the drying of porous analogue rock samples fabricated with ordinary Portland cement, with and without clay. This method can effectively and accurately reconstruct acoustic signals emitted from samples that only experience moisture loss without cracking. However, the method struggles to reconstruct signals from samples with intricate crack networks that formed during drying because AE generating mechanisms can emit contemporaneously, and the resulting waves propagate through drying‐induced cracks that can lead to multiple internal reflections. Thus, the differential reconstruction accuracy of time series generated by different underlying physical processes provides a robust filter for reducing large data catalogs. In general, both dynamics and sparse initiating events are learned directly from data and this method exposes a data hierarchy based on the complexity of the intrinsic dynamics.

58 GEOSCIENCES↗

Plug & play directed evolution of proteins with gradient-based discrete MCMC

Abstract A long-standing goal of machine-learning-based protein engineering is to accelerate the discovery of novel mutations that improve the function of a known protein. We introduce a sampling framework for evolving proteins in silico that supports mixing and matching a variety of unsupervised models, such as protein language models, and supervised models that predict protein function from sequence. By composing these models, we aim to improve our ability to evaluate unseen mutations and constrain search to regions of sequence space likely to contain functional proteins. Our framework achieves this without any model fine-tuning or re-training by constructing a product of experts distribution directly in discrete protein space. Instead of resorting to brute force search or random sampling, which is typical of classic directed evolution, we introduce a fast Markov chain Monte Carlo sampler that uses gradients to propose promising mutations. We conduct in silico directed evolution experiments on wide fitness landscapes and across a range of different pre-trained unsupervised models, including a 650 M parameter protein language model. Our results demonstrate an ability to efficiently discover variants with high evolutionary likelihood as well as estimated activity multiple mutations away from a wild type protein, suggesting our sampler provides a practical and effective new paradigm for machine-learning-based protein engineering.

59 BASIC BIOLOGICAL SCIENCES↗

Deep generative learning of magnetic frustration in artificial spin ice from magnetic force microscopy images

Increasingly large datasets of microscopic images with nanoscale resolution facilitate the development of machine learning methods to identify and analyze subtle physical phenomena embedded within the images. In this work, microscopic images of honeycomb lattice spin-ice samples serve as datasets from which we automate the calculation of net magnetic moments and directional orientations of spin-ice configurations. In the first stage of our workflow, machine learning models are trained to accurately predict magnetic moments and directions within spin-ice structures. Variational Autoencoders (VAEs), an emergent unsupervised deep learning technique, are employed to generate high-quality synthetic magnetic force microscopy (MFM) images and extract latent feature representations, thereby reducing experimental and segmentation errors. The second stage of proposed methodology enables precise identification and prediction of frustrated vertices and nanomagnetic segments, effectively correlating structural and functional aspects of microscopic images. This facilitates the design of optimized spin-ice configurations with controlled frustration patterns, enabling potential on-demand synthesis.

36 MATERIALS SCIENCE↗

Anomaly Detection in Seismic Data with Deep Learning: Application for Instrument Failure Detection and Forecasting

Seismic data quality assessment (QA) is the first and one of the most important steps before conducting any further data analysis. Traditional methods involve checking various metrics, such as spike detection and power spectral density, by setting strict thresholds or comparing data against synthetic benchmarks. However, these approaches often rely on pre-existing knowledge and assumptions about data anomalies, leading to potential misclassification of unusual cases. Here, in this study, we propose a deep autoencoder model, an unsupervised learning approach that evaluates data quality without making assumptions about normal and anomalous data, which can be used to identify deviations in recorded data that may indicate nascent instrument failure. We test the model with the U.S. International Monitoring System (IMS) seismic stations and demonstrate the capability of detecting anomalies on a monthly scale. This could prompt station operators to examine potential problems early, allowing sufficient time for instrument maintenance to prevent data outages. In addition, we use a new manually selected testing dataset to compare our model performance against two supervised machine learning (ML) approaches and a standard QA package, as baseline models. When applied to the dataset containing known data anomalies, performance of the supervised and unsupervised ML approaches is similar, with an accuracy of 88.1% for our model compared to ∼90% for the supervised ML approach and 78.2% for the standard QA package. Our model outperforms the baseline models when applied to new stations, where new types of data anomalies can be station-specific and not included in the training dataset. Finally, we show model transferability by training the model with data from the Global Seismograph Network only and applying it to the IMS network data. The results suggest that our model is generalizable and can be applied to new stations with good accuracy.

Lin, Jiun-Ting [Lawrence Livermore National Labora↗

Phasor-Measurement-Unit-Based Data Analytics Using Digital Twin and PhasorAnalytics Software

A major objective of this project was to apply GE’s commercial machine learning and data analytics toolsets to large-scale, real-world, anonymized Phasor Measurement Unit (PMU) datasets in order to extract signatures, correlated and/or causal factors, and precursor patterns associated with significant power system phenomena. The project had a particular emphasis on extraction of insights relevant to asset health monitoring, real-time load modeling and cybersecurity monitoring. Additionally, the team was directed to undertake a comprehensive data quality analysis for the provided datasets and encouraged to estimate the ‘machine-learning readiness’ of the datasets by documenting any major obstacles to the application of commercial machine learning algorithms. To accomplish the aforementioned objectives, the project team’s work centered around the identification of key event signatures and application of the identified event signatures for event detection and event classification. The industry-validated, semi-supervised machine learning strategy employed for event signature identification involved several major tasks, including data-preprocessing, generation of an overabundance of features, normal data identification, normality modeling, and event signature identification through a methodical, quantitative ranking of features in order of relevance to each studied event type. Throughout the project, data quality issues and mitigation techniques were investigated. In this report, insights are provided regarding the readiness of the provided synchrophasor datasets for application of machine learning and data analytics. The methodologies employed for this technical strategy are summarized in this report. With regards to data preprocessing and feature generation, the provided Training and Test Datasets were ingested into GE’s big data environment. Subsequently, the team applied bad data cleansing and data imputation scripts, event detection scripts, and application programming interfaces (APIs) to the datasets for convenient data access. The project team completed development and validation of dozens of physics-based, statistics-based and transformation-based feature functions used for the extraction of over 60 synchrophasor features. Using a new parallel feature generation technology developed on this project, over 60 features have been rapidly generated for the full two years’ worth of Training and Test Dataset data associated with both the Eastern and Western interconnects. Even accommodating for temporal down-sampling inherent to the feature extraction procedure, this parallel feature generation activity resulted in a massive feature set with a storage requirement approximately equal to that of the raw training dataset itself. With regards to normal data identification and normality modeling, a normality model was built using the feature data extracted from the Training Dataset and iteratively refined subsequent to incremental adjustments and expansions of the Training Dataset feature data. With respect to event characterization and signature identification, an event signature identification pipeline was developed and used in conjunction with the normality model to identify over 15 event signatures for key event categories within the Training Dataset. The identified event signatures were used to characterize hundreds of key events in terms of relative severity, duration, and location of the event. An investigation was undertaken to identify correlated and causal factors involved in transformer events. A separate investigation into temporal trends in ring-down analysis results was undertaken to determine possible associations between system dynamics and various other factors such as loading, season or year. To validate the identified event signatures, additional work was undertaken to develop signature-based anomaly detection and classification tools suitable for convenient application to the synchrophasor datasets. The anomaly detection and classification tools, suitable for online application, were then applied to the entirety of the Eastern Interconnect Training and Test Datasets. Performance of the event detection and classification tools was evaluated upon receipt of the Test Dataset event logs (i.e., the labels for events contained in the Test Dataset), and promising results were obtained despite several challenges (documented herein) associated with application of supervised or semi-supervised machine learning methods to large-scale, anonymized datasets. Finally, the detection and classification tools were used to detect, classify, and characterize thousands of new events not included in the original event logs provided by the DOE within both the Training and Test Datasets.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Characterizing Signatures of Geothermal Exploration Data with Machine Learning Techniques: An Application to the Nevada Play Fairway Analysis

We are introducing machine learning methods to the play fairway analysis to generate geothermal potential maps to support the evaluation of geothermal resource potential and the exploration for undiscovered blind geothermal systems in the Nevada Great Basin region. Our project aims to identify new ways to combine the play fairway data and empirically organize relationships between feature weights and labels in an improved workflow. As a means of doing this, we introduce machine learning methods to evaluate the influence of certain geological and geophysical features/feature sets in predicting geothermal favorability. This report highlights promising approaches based on supervised and unsupervised learning methods. First, we demonstrate a filter method applied to supervised classification modeling. The supervised filter method is based on permutation analysis to evaluate every possible feature combination/drop out scenario and rank feature influence based on the performance variance of supervised classification models. Additionally, we present an unsupervised factor analysis based on principal component analysis coupled with a semi-supervised kmeans clustering algorithm. This analysis allows us to identify the optimal number of groups/clusters for training sites and structural settings to identify feature patterns including correlation, variance, and latent and dominant feature relationships. The results from these methods offer a promising avenue for identifying favorable sources of predictive information to identify the locations of blind geothermal systems and furthering our understanding of complex geothermal feature and label relationships in the Great Basin region and beyond.

15 GEOTHERMAL ENERGY↗

Recent Advances toward Efficient Calculation of Higher Nuclear Derivatives in Quantum Chemistry

In this article, we provide an overview of state-of-the-art techniques that are being developed for efficient calculation of second and higher nuclear derivatives of quantum mechanical (QM) energy. Calculations of nuclear Hessians and anharmonic terms incur high costs and memory and scale poorly with system size. Three emerging classes of methods—machine learning (ML), automatic differentiation (AD), and matrix completion (MC)—have demonstrated promise in overcoming these challenges. We illustrate studies that employ unsupervised ML methods to reduce the need for multiple Hessian calculations in dynamics simulations and those that utilize supervised ML to construct approximate potential energy surfaces and estimate Hessians and anharmonic terms at reduced cost. By extension, if electronic structure operations could be written in a manner similar to functions underlying ML methods, rapid differentiation or AD routines can be employed to inexpensively calculate higher arbitrary-order derivatives. While ML approaches are typically black-box, we describe methods such as compressed sensing (CS) and MC, which explicitly leverage problem-specific mathematical properties of higher derivatives such as sparsity and low-rank, to complete higher derivative information using only a small, incomplete sample. The three classes of methods facilitate reliable predictions of observables ranging from infrared spectra to thermal conductivity and constitute a promising way forward in accurately capturing otherwise intractable higher-order responses of QM energy to nuclear perturbations.

38 RADIATION CHEMISTRY, RADIOCHEMISTRY, AND NUCLEA↗