Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “unsupervised machine learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Using Machine Learning to Identify Novel Hydroclimate States

Anthropogenic climate change is expected to alter drought risk in the future. However, droughts are not uncommon or unprecedented, as documented in tree-ring-based reconstructions of the summer average Palmer drought severity index (PDSI). Using an unsupervised machine-learning method trained on these reconstructions of pre-industrial climate, we identify outliers: years in which the spatial pattern of PDSI is unusual relative to ‘normal' variability. We show that in many regions, outliers are more frequently identified in the twentieth and twenty-first centuries. This trend is more pronounced when the regional drought atlases are combined into a single global dataset. By definition, outlier patterns at the 10% level are expected to occur once per decade, but from 1950 to 2000 more than 6 years per decade are identified as outliers in the global drought atlas (GDA). Extending the GDA through 2020 using an observational dataset suggests that anomalous global drought conditions are present in 80% of years in the twenty-first century. Our results indicate, without recourse to climate models, that the world is more frequently experiencing drought conditions that are highly unusual in the context of past natural climate variability.

Drought risk↗

Separating Physically Distinct Mechanisms in Complex Infrared Plasmonic Nanostructures via Machine Learning Enhanced Electron Energy Loss Spectroscopy

Electron energy loss spectroscopy (EELS) enables direct exploration of plasmonic phenomena at the nanometer level. To isolate individual plasmon modes, linear unmixing methods can be used to separate different physical mechanisms, but in larger and more complex systems the interpretability of the components becomes uncertain. Here, infrared plasmonic resonances in self-assembled heterogeneous monolayer films of doped-semiconductor nanoparticles are examined beyond linear unmixing techniques, and both supervised and unsupervised machine-learning-based analyses of hyperspectral EELS datasets are demonstrated. Additionally, in the supervised approach, a human operator labels a small number of pixels in the hyperspectral dataset corresponding to features of interest which are then propagated across the entire dataset. In the unsupervised approach, non-linear autoencoders are used to create a highly-reduced latent-space representation of the dataset, within which insight into the relevant physics can be gleaned from straightforward distance metrics that do not depend on operator input and bias. The advantage of these approaches is that the labeling separates physical mechanisms without altering the data, enabling robust analyses of the influence of heterogeneities in mesoscale complex systems.

36 MATERIALS SCIENCE↗

Machine learning to discover mineral trapping signatures due to CO 2 injection

Mineral trapping is pursued as a geological CO 2 sequestration (GCS) mechanism because it permanently stores CO 2 in solid phases or minerals. However, CO 2 mineral-trapping mechanisms are poorly understood due to (1) lack of sufficient field and laboratory data characterizing these complex processes, and (2) challenges to develop site-specific reactive-transport models coupling fluid flow and geochemical reactions occurring at various temporal (from milliseconds to years) and spatial (from pore (millimeters) to field (kilometers)) scales. Reactive transport with additional complexities such as heterogeneity can make the simulation outputs even more difficult to interpret because of complex nonlinearity and multi-scale interdependencies. Furthermore, the values of model outputs such as concentrations can vary by several orders of magnitude, making it harder to correlate and characterize the impact of the variables via traditional data interpretation techniques such as exploratory data analyses. Recently, machine learning (ML) has shown promise in feature discovery and in highlighting hidden mechanisms that cannot be obtained by existing data-analytics and statistical methods. In this study, we applied an unsupervised ML approach, non-negative matrix factorization with custom -means clustering (NMF) to the data generated by reactive-transport simulations of GCS. The reactive-transport data consisted of 19 attributes, including four physio-chemical variables (pH, porosity, aqueous CO 2 , and sequestered CO 2 ), six chemical species (K + , Na + , HCO, Ca 2+ , Mg 2+ , Fe 2+ ), and four carbonate minerals (calcite, dolomite, siderite, and ankerite), a feldspar mineral (albite), and four clay minerals (illite, clinochlore, kaolinite, and smectite) over a period of 200 years of simulation time. Furthermore, the simulation data used was for Morrow B sandstone at the Farnsworth hydrocarbon unit in Texas. Data are sampled at two locations within the model domain: (1) at the injection well and (2) 200 m west of the injection well. The injection was performed for a period of 10 years. Using NMF, we estimated the temporal interdependencies among the 19 attributes over a span of 200 years. We found that NMF was able to identify four reaction stages and their dominant attributes; these cannot be directly discerned through traditional visualization (e.g., line plots, Pareto analysis, Glyph-based visualization methods) or exploratory data analysis tools of the simulation data. The four stages were: reactions in the injection phase followed by short-, mid-, and long-term reactions. The NMF analysis also revealed that 10 among the 19 attributes are dominant. These dominant attributes for mineral trapping include calcite, dolomite at injection well, siderite at 200 m away from the injection well, clinochlore, kaolinite, Na + , K + , Ca 2+ , Mg 2+ , pH, and aqeuous CO 2 . Finally, at late times (65–200 years), our results showed that calcite plays a major role in mineral trapping with insignificant contribution from siderite, ankerite, and clay minerals. These findings make the proposed unsupervised ML-model attractive for reactive-transport sensing towards real-time GCS monitoring.

54 ENVIRONMENTAL SCIENCES↗

PixelLearn

PixelLearn is an integrated user-interface computer program for classifying pixels in scientific images. Heretofore, training a machine-learning algorithm to classify pixels in images has been tedious and difficult. PixelLearn provides a graphical user interface that makes it faster and more intuitive, leading to more interactive exploration of image data sets. PixelLearn also provides image-enhancement controls to make it easier to see subtle details in images. PixelLearn opens images or sets of images in a variety of common scientific file formats and enables the user to interact with several supervised or unsupervised machine-learning pixel-classifying algorithms while the user continues to browse through the images. The machinelearning algorithms in PixelLearn use advanced clustering and classification methods that enable accuracy much higher than is achievable by most other software previously available for this purpose. PixelLearn is written in portable C++ and runs natively on computers running Linux, Windows, or Mac OS X.

Mazzoni, Dominic↗

Automated identification of dominant physical processes

The identification of processes that locally and approximately dominate dynamical system behavior has enabled significant advances in understanding and modeling nonlinear differential dynamical systems. Conventional methods of dominant process identification involve piecemeal and ad hoc (non-rigorous, informal) scaling analyses to identify dominant balances of governing equation terms and to delineate the spatiotemporal boundaries (boundaries in space and/or time) of each dominant balance. For the first time, we present an objective global measure of the fit of dominant balances to observations, which is desirable for automation, and was previously undefined. Furthermore, we propose a formal definition of the dominant balance identification problem in the form of an optimization problem. Here, we show that the optimization can be performed by various machine learning algorithms, enabling the automatic identification of dominant balances. Our method is algorithm agnostic and it eliminates reliance upon expert knowledge to identify dominant balances which are not known beforehand.

42 ENGINEERING↗

Anomaly detection in collider physics via factorized observables

To maximize the discovery potential of high-energy colliders, experimental searches should be sensitive to unforeseen new physics scenarios. This goal has motivated the use of machine learning for unsupervised anomaly detection. In this paper, we introduce a new anomaly detection strategy called : factorized observables for regressing conditional expectations. Our approach is based on the inductive bias of factorization, which is the idea that the physics governing different energy scales can be treated as approximately independent. Assuming factorization holds separately for signal and background processes, the appearance of nontrivial correlations between low- and high-energy observables is a robust indicator of new physics. Under the most restrictive form of factorization, a machine-learned model trained to identify such correlations will in fact converge to the optimal new physics classifier. We test on a benchmark anomaly detection task for the Large Hadron Collider involving collimated sprays of particles called jets. By teasing out correlations between the kinematics and substructure of jets, our method can reliably extract percent-level signal fractions. This strategy for uncovering new physics adds to the growing toolbox of anomaly detection methods for collider physics with a complementary set of assumptions. Published by the American Physical Society 2024

Astronomy & Astrophysics↗

Process Anomaly Detection for Sparsely Labeled Events in Nuclear Power Plants

An essential aspect of online monitoring, subtle anomaly detection increases the detection lead time for equipment failure and enables a nuclear power plant (NPP) to mitigate unexpected partial or full outages, resulting in significant cost saving to the plant. Once an anomaly is detected by plant staff, its cause and severity are investigated. Because the vast majority of anomalies require some level of investigation, including some that require time-consuming examination, before they are passed over to the engineering organization for further analysis, plants are often equipped with tools to assist the staff in performing anomaly detection. Those tools operate as a black box and are often based on statistical methods that establish sensor correlations using preconfigured mathematical models and flag correlation deviations as anomalies. Due to the number of anomalies detected at a given NPP on a daily basis, a significant number of flagged anomalies usually await examination for days or weeks. A primary cause of this backlog is that the methods used by the tools generate many false positives. Though this is usually attributed to oversensitive model settings due to very narrow normal operation bands, it can also be associated with the model development being inadequate for the process being monitored, or with missing model inputs that could have explained misclassified positives. The performance of anomaly detection tools impacts their plant acceptance and utilization, especially when the effort to address false positives generated by the tool depletes the value or cost saved by using that tool. Thus, means to advance anomaly detection performance have been investigated by the Department of Energy’s Light Water Reactor Sustainability program. Previous and ongoing efforts have targeted unsupervised machine-learning (ML) methods, which do not require the labeling of any data fed into the ML model. By contrast, in supervised anomaly detection methods, every data point is labeled as either a normal or abnormal process condition, and the model is trained to replicate the classification process. Supervised methods usually outperform unsupervised methods, due to the added value in differentiating normal from anomalous states of the monitored process. An NPP’s corrective action program requires it to track and document, via a dedicated report, the resolution of any issues that occur within the plant. Once created, each report is reviewed by a plant screening committee, and several classifications and decisions are made. Recently, a collaborating NPP developed an artificial intelligence and ML-based classifier to categorize a condition report (CR) into classes that can serve to label the data as normal or anomalous. Applying CRs as labels represents a semi-supervised use case. Semi-supervised ML assumes that labels exist for some data points (i.e., labeled anomalies, in this case) but not for the rest. In this effort, semi-supervised ML methods were used to fuse data from CRs with anomaly detection methods in order to test the hypothesis that partially labeled anomalies would improve the accuracy of the anomaly detection methods. Specifically, two methods were used. The first is the deep Semi-supervised Anomaly Detection (deep SAD) method, which can handle labels ranging from fully unsupervised to fully supervised cases. The second is a newly designed ML method developed specifically for this effort and referred to as the high-order feature (HOF)-based method. To evaluate these two methods in controlled environments, synthetic data generators were developed and used. The first datasets used a spring-mass-damper (SMD) system simulator commonly found in mechanical engineering references. This was used to create two use cases: a one- and a three-mass system. Anomalies were introduced by changing the spring and damper coefficients while the system was actuated by random forces. The second datasets used the commercial Dymola-Modelica software to build a simplified nuclear reactor model. Anomalies were added in the form of corrupted sensor readings and/or control commands. The deep SAD method was tested using the SMD system, while the HOF method was tested using both datasets. Application of the deep SAD semi-supervised ML method demonstrated that labels can generate increased confidence in detecting true anomalies. This helped increase the number of true positives and decrease the number of false negatives—something that would aid in addressing the backlog of possible anomalies. Application of the HOF method demonstrated that labels can aid in down selecting from a candidate set of features to a more optimal subset in order to better differentiate between normal and anomalous conditions.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Machine-learning based approach to examine ecological processes influencing the diversity of riverine dissolved organic matter composition

Dissolved organic matter (DOM) assemblages in freshwater rivers are formed from mixtures of simple to complex compounds that are highly variable across time and space. These mixtures largely form due to the environmental heterogeneity of river networks and the contribution of diverse allochthonous and autochthonous DOM sources. Most studies are, however, confined to local and regional scales, which precludes an understanding of how these mixtures arise at large, e.g., continental, spatial scales. The processes contributing to these mixtures are also difficult to study because of the complex interactions between various environmental factors and DOM. Here we propose the use of machine learning (ML) approaches to identify ecological processes contributing toward mixtures of DOM at a continental-scale. We related a dataset that characterized the molecular composition of DOM from river water and sediment with Fourier-transform ion cyclotron resonance mass spectrometry to explanatory physicochemical variables such as nutrient concentrations and stable water isotopes ( 2 H and 18 O). Using unsupervised ML, distinctive clusters for sediment and water samples were identified, with unique molecular compositions influenced by environmental factors like terrestrial input and microbial activity. Sediment clusters showed a higher proportion of protein-like and unclassified compounds than water clusters, while water clusters exhibited a more diversified chemical composition. We then applied a supervised ML approach, involving a two-stage use of SHapley Additive exPlanations (SHAP) values. In the first stage, SHAP values were obtained and used to identify key physicochemical variables. These parameters were employed to train models using both the default and subsequently tuned hyperparameters of the Histogram-based Gradient Boosting (HGB) algorithm. The supervised ML approach, using HGB and SHAP values, highlighted complex relationships between environmental factors and DOM diversity, in particular the existence of dams upstream, precipitation events, and other watershed characteristics were important in predicting higher chemical diversity in DOM. Our data-driven approach can now be used more generally to reveal the interplay between physical, chemical, and biological factors in determining the diversity of DOM in other ecosystems.

54 ENVIRONMENTAL SCIENCES↗

Discovering hidden geothermal signatures using non-negative matrix factorization with customized k-means clustering

Discovery of hidden geothermal resources is challenging. It requires the mining of large datasets with diverse data attributes representing subsurface hydrogeological and geothermal conditions. The commonly used play fairway analysis approach typically incorporates subject-matter expertise to analyze regional data to estimate geothermal characteristics and favorability. We demonstrate an alternative approach based on machine learning (ML) to process a geothermal dataset from southwest New Mexico (SWNM). The study region includes low- and medium-temperature hydrothermal systems. Several of these systems are not well characterized because of insufficient existing data and limited past explorative work. This study discovers hidden patterns and relations in the SWNM geothermal dataset to improve our understanding of the regional hydrothermal conditions and energy-production favorability. This understanding is obtained by applying an unsupervised ML algorithm based on non-negative matrix factorization coupled with customized k-means clustering (NMFk). NMFk can automatically identify (1) hidden signatures characterizing analyzed datasets, (2) the optimal number of these signatures, (3) the dominant data attributes associated with each signature, and (4) the spatial distribution of the extracted signatures. Here, in this study, NMFk is applied to analyze 18 geological, geophysical, hydrogeological, and geothermal attributes at 44 locations in SWNM. Using NMFk, we find data patterns and identify the spatial associations of hydrothermal signatures within two physiographic provinces (Colorado Plateau and Basin and Range) and two sub-regions of these provinces (the Mogollon-Datil volcanic field and the Rio Grande rift) in SWNM. The ML algorithm extracted five hydrothermal signatures in the SWNM datasets that differentiate between low (<90°C) and medium (90-150°C)-temperature hydrothermal systems. The algorithm also suggests that the Rio Grande rift and northern Mogollon-Datil volcanic field are the most favorable regions for future geothermal resource discovery. NMFk also identified critical attributes to identify medium-temperature hydrothermal systems in the study area. The resulting NMFk model can be applied to predict geothermal conditions and their uncertainties at new SWNM locations based on limited data from unexplored regions. The code to execute the performed analyses as well as the corresponding data can be found at https://github.com/SmartTensors/GeoThermalCloud.jl.

15 GEOTHERMAL ENERGY↗

Real-time tracking and analysis of gas bubble dynamics in laser powder bed fusion using in-situ X-ray characterization and machine learning

Porosity defects remain a significant challenge in the laser powder bed fusion (LPBF) process, adversely affecting the mechanical properties and reliability of additively manufactured components. Here, this study investigates the real-time formation and trajectory of gas bubbles during LPBF of Al6061 alloy using advanced in-situ X-ray characterization and machine learning. The unsupervised Gaussian mixture model and particle tracking algorithm developed are able to precisely track and quantify the properties of gas bubbles and keyhole pores. Our analysis identified five distinct types of gas bubble formation and movement patterns, emphasizing the diverse origins and behaviors of these defects. It enables precise quantification of trajectories, velocities, and morphological changes of gas bubbles, offering a granular view of the subsurface dynamics within the melt pool. Additionally, we explored keyhole-induced pore dynamics, revealing the critical role of keyhole oscillation and collapse for the formation of both large and small gas pores. It defines four different regions of gas bubble movement within the melt pool, providing a clearer understanding of how local fluid dynamics affect pore behavior. The results underscore the importance of integrating in-situ experimental observation and automated machine learning to develop a more robust predictive model for defect formation in LPBF.

In-situ X-ray imaging↗

Learning to simulate high energy particle collisions from unlabeled data

In many scientific fields which rely on statistical inference, simulations are often used to map from theoretical models to experimental data, allowing scientists to test model predictions against experimental results. Experimental data is often reconstructed from indirect measurements causing the aggregate transformation from theoretical models to experimental data to be poorly-described analytically. Instead, numerical simulations are used at great computational cost. We introduce Optimal-Transport-based Unfolding and Simulation (OTUS), a fast simulator based on unsupervised machine-learning that is capable of predicting experimental data from theoretical models. Without the aid of current simulation information, OTUS trains a probabilistic autoencoder to transform directly between theoretical models and experimental data. Identifying the probabilistic autoencoder’s latent space with the space of theoretical models causes the decoder network to become a fast, predictive simulator with the potential to replace current, computationally-costly simulators. Here, we provide proof-of-principle results on two particle physics examples, Z-boson and top-quark decays, but stress that OTUS can be widely applied to other fields.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Source-agnostic gravitational-wave detection with recurrent autoencoders

Abstract We present an application of anomaly detection techniques based on deep recurrent autoencoders (AEs) to the problem of detecting gravitational wave (GW) signals in laser interferometers. Trained on noise data, this class of algorithms could detect signals using an unsupervised strategy, i.e. without targeting a specific kind of source. We develop a custom architecture to analyze the data from two interferometers. We compare the obtained performance to that obtained with other AE architectures and with a convolutional classifier. The unsupervised nature of the proposed strategy comes with a cost in terms of accuracy, when compared to more traditional supervised techniques. On the other hand, there is a qualitative gain in generalizing the experimental sensitivity beyond the ensemble of pre-computed signal templates. The recurrent AE outperforms other AEs based on different architectures. The class of recurrent AEs presented in this paper could complement the search strategy employed for GW detection and extend the discovery reach of the ongoing detection campaigns.

47 OTHER INSTRUMENTATION↗

Optical Control of Adaptive Nanoscale Domain Networks

Adaptive networks can sense and adjust to dynamic environments to optimize their performance. Understanding their nanoscale responses to external stimuli is essential for applications in nanodevices and neuromorphic computing. However, it is challenging to image such responses on the nanoscale with crystallographic sensitivity. Here, the evolution of nanodomain networks in (PbTiO 3 ) n /(SrTiO 3 ) n superlattices (SLs) is directly visualized in real space as the system adapts to ultrafast repetitive optical excitations that emulate controlled neural inputs. The adaptive response allows the system to explore a wealth of metastable states that are previously inaccessible. Their reconfiguration and competition are quantitatively measured by scanning x-ray nanodiffraction as a function of the number of applied pulses, in which crystallographic characteristics are quantitatively assessed by assorted diffraction patterns using unsupervised machine-learning methods. The corresponding domain boundaries and their connectivity are drastically altered by light, holding promise for light-programable nanocircuits in analogy to neuroplasticity. Phase-field simulations elucidate that the reconfiguration of the domain networks is a result of the interplay between photocarriers and transient lattice temperature. The demonstrated optical control scheme and the uncovered nanoscopic insights open opportunities for the remote control of adaptive nanoscale domain networks.

36 MATERIALS SCIENCE↗

Towards inverse microstructure-centered materials design using generative phase-field modeling and deep variational autoencoders

The field of Integrated Computational Materials Engineering (ICME) combines a broad range of methods to study materials’ responses over a spectrum of length scales. A relatively unexplored aspect of microstructure-sensitive materials design is uncertainty propagation and quantification (UP/UQ) of materials’ microstructure, as well as establishing process-structure–property (PSP) relationships for inverse material design. In this study, an efficient UP technique built on the idea of changing probability measures and a deep generative unsupervised representative machine learning method for microstructure-based design of thermal conductivity of materials is proposed. Probability measures are used to represent microstructure space, and Wasserstein metrics are used to test the efficiency of the UP method. By using deep Variational AutoEncoder (VAE), we identify the correlations between the material/process parameters and the thermal conductivity of heterogeneous dual-phase microstructures. Through high-throughput screening, UP, and the deep-generative VAE method, PSP relationships that are too complex can be revealed by exploiting the materials’ design space with an emphasis on microstructures. As a last point, we demonstrate generative machine learning serves as a useful tool for inverse microstructure-centered materials design, and we demonstrate this by examining the inverse design of thermal conductivity in nano-structured materials. Here, the results reveal the effects of morphology, volume fraction, characteristic length scale, and the individual thermal diffusivity of phases on the thermal conductivity of dual-phase alloys. Our findings emphasize the advantages of high-throughput phase-field modeling and generative deep learning for linking PSP and inverse microstructure-centered materials design.

36 MATERIALS SCIENCE↗