Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Unsupervised machine learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Real-Time Anomaly Detection for Searches Beyond the Standard Model in the ProtoDUNE Horizontal Drift Detector

This paper summarizes work conducted throughout a SULI internship at Fermi National Accelerator Laboratory focused on building an unsupervised machine learning model for real-time anomaly detection in ProtoDUNE Horizontal Drift. Using simulated data, we trained an autoencoder model on a pure cosmic dataset, and evaluated it on both cosmic and neutrino events—making the model an anomaly detector. The goal was to make a model which matches or exceeds the current ADC Simple Window trigger algorithm so that our model can perform at the same rate but provide sensitivity to potential beyond-the-Standard-Model (BSM) signatures. In the end, we were able to construct a model which slightly exceeds the capabilities of the ADC Simple Window while remaining completely unsupervised, achieving 31.9 ± 0.2% (26.6 ± 0.2%) ν efficiency at 5 Hz (2 Hz), a 3.6 (3.2) percentage point increase. Additionally, 17.5 ± 0.3% (18.3 ± 0.3%) of the events that passed the autoencoder at 5 Hz (2 Hz) were missed by the current trigger algorithm. Future work will investigate alternative normalization methods, including quantile transformation, and evaluate the model on ProtoDUNE-HD detector-glitch data if that data becomes available.

Wilson, C. [Cincinnati U., RWC]↗

The Dark Machines Anomaly Score Challenge: Benchmark Data and Model Independent Event Classification for the Large Hadron Collider

We describe the outcome of a data challenge conducted as part of the Dark Machines (https://www.darkmachines.org) initiative and the Les Houches 2019 workshop on Physics at TeV colliders. The challenged aims to detect signals of new physics at the Large Hadron Collider (LHC) using unsupervised machine learning algorithms. First, we propose how an anomaly score could be implemented to define model-independent signal regions in LHC searches. We define and describe a large benchmark dataset, consisting of >1 billion simulated LHC events corresponding to 10\, fb^{-1} 10 f b − 1 of proton-proton collisions at a center-of-mass energy of 13 TeV. We then review a wide range of anomaly detection and density estimation algorithms, developed in the context of the data challenge, and we measure their performance in a set of realistic analysis environments. We draw a number of useful conclusions that will aid the development of unsupervised new physics searches during the third run of the LHC, and provide our benchmark dataset for future studies at https://www.phenoMLdata.org. Code to reproduce the analysis is provided at https://github.com/bostdiek/DarkMachines-UnsupervisedChallenge.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

InClass nets: independent classifier networks for nonparametric estimation of conditional independence mixture models and unsupervised classification

Abstract Conditional independence mixture models (CIMMs) are an important class of statistical models used in many fields of science. We introduce a novel unsupervised machine learning technique called the independent classifier networks (InClass nets) technique for the nonparameteric estimation of CIMMs. InClass nets consist of multiple independent classifier neural networks (NNs), which are trained simultaneously using suitable cost functions. Leveraging the ability of NNs to handle high-dimensional data, the conditionally independent variates of the model are allowed to be individually high-dimensional, which is the main advantage of the proposed technique over existing non-machine-learning-based approaches. Two new theorems on the nonparametric identifiability of bivariate CIMMs are derived in the form of a necessary and a (different) sufficient condition for a bivariate CIMM to be identifiable. We use the InClass nets technique to perform CIMM estimation successfully for several examples. We provide a public implementation as a Python package called RainDancesVI.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Unsupervised Process Anomaly Detection and Identification Using the Leave-One-Variable-Out Approach

Automated anomaly detection and identification can signal equipment issues and pinpoint causes in large-scale industrial systems. For systems with limited failure history, unsupervised machine learning methods can be utilized as they do not require past failures. This study introduces the leave-one-variable-out (LOVO) model, which masks one variable at a time to predict the others, learning underlying process correlations. Detection performance was assessed with synthetic and experimental data, while identification performance used only synthetic data due to its ability to generate labeled anomaly types. For detection using synthetic data, the LOVO model generally outperformed comparative models; while using experimental data, the comparative methods outperformed the LOVO model. However, the comparative methods required selecting a latent size, and these conclusions pertain to using the optimal size. In practice, it would not be feasible to always select the optimal value, and incorrect selections impacted performance. In contrast, the LOVO model does not require a latent space. For identification using synthetic data, the LOVO model was slightly outperformed in interpretability and repeatability but still demonstrated impressive results. These outcomes suggest that the LOVO model is an effective model and may be more easily implemented without the challenging tuning process of selecting a latent size.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Oleaginous Yeast Biology Elucidated With Comparative Transcriptomics

ABSTRACT Extremophilic yeasts have favorable metabolic and tolerance traits for biomanufacturing‐ like lipid biosynthesis, flavinogenesis, and halotolerance – yet the connection between these favorable phenotypes and strain genotype is not well understood. To this end, this study compares the phenotypes and gene expression patterns of biotechnologically relevant yeasts Yarrowia lipolytica , Debaryomyces hansenii , and Debaryomyces subglobosus grown under nitrogen starvation, iron starvation, and salt stress. To analyze the large data set across species and conditions, two approaches were used: a “network‐first” approach where a generalized metabolic network serves as a scaffold for mapping genes and a “cluster‐first” approach where unsupervised machine learning co‐expression analysis clusters genes. Both approaches provide insight into strain behavior. The network‐first approach corroborates that Yarrowia upregulates lipid biosynthesis during nitrogen starvation and provides new evidence that riboflavin overproduction in Debaryomyces yeasts is overflow metabolism that is routed to flavin cofactor production under salt stress. The cluster‐first approach does not rely on annotation; therefore, the coexpression analysis can identify known and novel genes involved in stress responses, mainly transcription factors and transporters. Therefore, this work links the genotype to the phenotype of biotechnologically relevant yeasts and demonstrates the utility of complementary computational approaches to gain insight from transcriptomics data across species and conditions.

Weintraub, Sarah J. [Department of Bioinformatics ↗

Exploring urban typologies using comprehensive analysis of transportation dynamics

Abstract As urban areas continue to expand and develop, categorizing cities into typologies offers a valuable framework for understanding metropolitan dynamics and fostering inter-city collaboration. However, existing typologies related to urban mobility have limitations, failing to consider cities within a single large urban region and often overlooking crucial dimensions such as trip demand and traffic flow. In this paper, we introduce a transportation-focused characterization for cities within a large urban region, specifically the San Francisco Bay Area, California. We incorporate over 40 metrics across five transportation dimensions: trip demand, road network, multi-modal network, traffic flow, and land use. Specifically, for the trip demand dimension, we include metrics capturing residents’ trip characteristics, such as mode share, intra-city trips, and inter-city trips. Additionally, we analyze the purpose of trips entering the city to gain a deeper understanding of incoming trip patterns. In the traffic flow dimension, we examine metrics like vehicle miles traveled, delay, and congestion to assess the traffic conditions on the street network. These, combined with other dimensions, provide a comprehensive view of a city’s transportation dynamics. Using unsupervised machine learning clustering methods, we identified eight distinct typologies for the Bay Area: Live Work Cities; Job and Activity Magnet Cities; Anchor Cities; Multi-modal Cities; Hyper-connected Cities; Low-density Residential Cities; Medium-density Residential Cities; and Mixed-use Residential Cities. Our findings show that many clusters are strongly influenced by trip demand and traffic flow metrics. Finally, we examine the practicality of this typology and its potential to guide collaborative transportation management strategies. The typologies provide a foundation for dialogue among Bay Area cities, focusing on evaluating shared characteristics and leveraging successes or challenges to develop unified strategies for transportation management.

Kuncheria, Anu↗

Mapping structural heterogeneity at the nanoscale with scanning nano-structure electron microscopy (SNEM)

Here, in this work, we explore the use of scanning electron diffraction (also known as 4D-STEM) coupled with electron atomic pair distribution function analysis (ePDF) to understand the local order (structure and chemistry) as a function of position in a complex multicomponent system, a hot rolled, Ni-encapsulated, Zr 65 Cu 17.5 Ni 10 Al 7.5 bulk metallic glass (BMG), with a spatial resolution of 3 nm. We show that it is possible to gain insight into the chemistry and chemical clustering/ordering tendency in different regions of the sample, including in the vicinity of nano-scale crystallites that are identified from virtual dark field images and in heavily deformed regions at the edge of the BMG. In addition to simpler analysis, unsupervised machine learning was used to extract partial PDFs from the material, modeled as a quasi-binary alloy, and map them in space. These maps allowed key insights not only into the local average composition, as validated by EELS, but also a unique insight into chemical short-range ordering tendencies in different regions of the sample during formation. The experiments are straightforward and rapid and, unlike spectroscopic measurements, don’t require energy filters on the instrument. We spatially map different quantities of interest (QoI’s), defined as scalars that can be computed directly from positions and widths of ePDF peaks or parameters refined from fits to the patterns. We developed a flexible and rapid data reduction and analysis software framework that allows experimenters to rapidly explore images of the sample on the basis of different QoI’s. The power and flexibility of this approach are explored and described in detail. Because of the fact that we are getting spatially resolved images of the nanoscale structure obtained from ePDFs we call this approach scanning nano-structure electron microscopy (SNEM), and we believe that it will be powerful and useful extension of current 4D-STEM methods.

36 MATERIALS SCIENCE↗

Anomalously high elastic modulus of a poly(ethylene oxide)-based composite electrolyte

The practical use of lithium metal anodes in solid-state batteries requires a polymer membrane with high lithium-ion conductivity, thermal/electrochemical stability, and mechanical strength. The primary challenge is to effectively decouple the ionic conductivity and mechanical strength of the polymer electrolytes. We report a remarkably facile single step synthetic strategy based on in-situ crosslinking of poly(ethylene oxide) (xPEO) in the presence of a woven glass fiber (GF). Such a simple method yields composite polymer electrolytes (CPE) of anomalously high elastic modulus up to 2.5 GPa over a broad temperature range (20 °C – 245 °C) that has never been previously documented. An unsupervised machine learning algorithm, K-mean clustering analysis, was implemented on the hyperspectral Raman mapping at the xPEO/GF interface. Using such a unique means, we show for the first time that the promoted mechanical strength originates from xPEO and GF interactions through dynamic hydrogen and ionic bonding. High ionic conductivity is achieved by the addition plasticizer (e.g. tetraglyme), where trifluoromethanesulfonate anions are tethered to the xPEO matrix and Li + cations are favorably transported through coordination with the plasticizer. Further, stringent galvanostatic cycling tests indicates the CPE can be stably cycled for >3000 h in a Li-metal symmetric cell at a moderate temperature (nearly 1500 Coulombs/cm 2 Li equivalents), outperforming most of the PEO-based electrolytes. The GF reinforced CPE reported here has multifunctional uses, such as solid electrolytes for all solid-state batteries and membranes for redox-flow batteries. Although the focus of this study is on lithium-based batteries, the results are equally promising for other alkali metal based batteries such as sodium and potassium.

25 ENERGY STORAGE↗

Time-resolved spray characterization via unified optical flow and binarization technique

This work leverages an unsupervised machine learning and advanced image processing techniques to characterize the breakup of fuel sprays in a small-scale combustor under reacting conditions, providing valuable insights into near-nozzle flow phenomenology. The proposed methodology integrates an improved optical flow model on a convolutional neural network to extract flow vectors with a binarization technique to assess droplets’ size and shape across the region of interest. The velocimetry approach demonstrates superior performance compared to a state-of-the-art optical flow model when applied to high-speed X-ray phase contrast spray images, achieving more accurate and reliable flow predictions. Moreover, breakup processes are quantified by breakup length and sphericity in accordance with velocity estimations, allowing a more complete characterization of the flow. This study establishes a robust methodology for analyzing spray morphology and primary breakup in compact combustors, contributing valuable means of understanding and optimizing fuel spray behavior in advanced combustion systems.

42 ENGINEERING↗

Heterogeneous microstructure of yttrium hydride and its relation to mechanical properties

Here, the goal of this study is to investigate the properties of yttrium hydride materials in relation to the microstructure, especially its homogeneity. High-throughput nanoindentation mapping was used to evaluate hardness distribution. Raman spectral imaging demonstrated its sensitivity to the presence of YH2 and impurities. Raman peak position maps were correlated with residual stress in the specimens. Electron backscatter diffraction mapping provided phase distributions with correlation to high-energy X-ray diffraction analysis. The experimental mapping data were combined and analyzed using unsupervised machine learning cluster procedures. The machine learning analysis revealed that yttrium hydride specimens contained a major δ-YH2 – x phase component and minor α-Y and δ-YH2 – x components with significant residual stress. The minor phase fraction decreased with increasing nominal H/Y ratio, which affected the nanoindentation and Vickers hardness. The multimodal mapping procedures described herein affect developing important microstructure–property relationships, as well as correlations in heterogeneity and mechanical properties.

36 MATERIALS SCIENCE↗

Automated phase segmentation and quantification of high-resolution TEM image for alloy design

In the alloy design and development process, a wealth of atomically resolved structural high-resolution transmission electron microscopy (HRTEM) images are produced. Identifying the different nano-precipitate phases and tracking their evolution under various compositions and during manufacturing or post-processing requires hundreds of HRTEM images and thousands of precipitates. The nanoscopic phase information labeling and analysis purely relies on humans are prohibitively costly and time-consuming, sometimes not reliable because of the lack of authoritative knowledge. Here, in this work, we develop a novel unsupervised machine learning approach coupled with adaptive computer vision techniques with features in the Fourier space to automatically determine the number of phases and segment/quantify the phases with nanoscale resolution, allowing for quantitative correlation between nanostructure formation, processing and functional properties. To automate the phase extraction/quantification and ascertain its applicability, we have applied the developed framework to the HRTEM images from several alloy systems, processing conditions, image magnifications, and phase types and morphologies (precipitates, nano-twins, stacking faults, crystalline matrix, and amorphous structures) for verification. This study paves the road for compression, visualization, and translation of raw image structural data into physically relevant information in real-time with minimal human supervision. It shows the promise of enabling high-throughput materials characterization for the acceleration of alloy manufacturing and design.

36 MATERIALS SCIENCE↗

Database development and exploration of process–microstructure relationships using variational autoencoders

The paper demonstrates graphical representation of a large database containing process–microstructure relationships using an unsupervised machine learning algorithm. Correlating microstructural features to processing is an essential first step to answer the difficult problem of process sequence design. Here, a large database of 346,200 orientation distribution functions resulting from a variety of process sequences is constructed, where each sequence comprises up to four stages of tension, compression and rolling along different directions in various permutations. This open-source database is constructed for collaborative development of process design algorithms. The paper demonstrates a novel application of the large database: graphical representation of texture–process relationships. A variational autoencoder is used to reduce the entire database to a two dimensional latent space where variations in processes and properties can be visualized. Using proximity analysis in this latent space, we can quickly unearth multiple process solutions to the problem of texture or property design.

36 MATERIALS SCIENCE↗

Combinatorial Exploration and Mapping of Phase Transformation in a Ni–Ti–Co Thin Film Library

Combinatorial synthesis and high-throughput characterization of a Ni–Ti–Co thin film materials library are reported for exploration of reversible martensitic transformation. The library was prepared by magnetron co-sputtering, annealed in vacuum at 500 °C without atmospheric exposure, and evaluated for shape memory behavior as an indicator of transformation. Composition, structure, and transformation behavior of the 177 pads in the library were characterized using high-throughput wavelength dispersive spectroscopy (WDS), X-ray photoelectron spectroscopy (XPS), X-ray diffraction (XRD), and four-point probe temperature-dependent resistance (R(T)) measurements. A new, expanded composition space having phase transformation with low thermal hysteresis and Co > 10 at. % is found. Unsupervised machine learning methods of hierarchical clustering were employed to streamline data processing of the large XRD and XPS data sets. Through cluster analysis of XRD data, we identified and mapped the constituent structural phases. Finally, composition–structure–property maps for the ternary system are made to correlate the functional properties to the local microstructure and composition of the Ni–Ti–Co thin film library.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Metabolomics and the pig model reveal aberrant cardiac energy metabolism in metabolic syndrome

Abstract Although metabolic syndrome (MS) is a significant risk of cardiovascular disease (CVD), the cardiac response (MR) to MS remains unclear due to traditional MS models’ narrow scope around a limited number of cell-cycle regulation biomarkers and drawbacks of limited human tissue samples. To date, we developed the most comprehensive platform studying MR to MS in a pig model tightly related to human MS criteria. By incorporating comparative metabolomic, transcriptomic, functional analyses, and unsupervised machine learning (UML), we can discover unknown metabolic pathways connections and links on numerous biomarkers across the MS-associated issues in the heart. For the first time, we show severely diminished availability of glycolytic and citric acid cycle (CAC) pathways metabolites, altered expression, GlcNAcylation, and activity of involved enzymes. A notable exception, however, is the excessive succinate accumulation despite reduced succinate dehydrogenase complex iron-sulfur subunit b (SDHB) expression and decreased content of precursor metabolites. Finally, the expression of metabolites and enzymes from the GABA-glutamate, GABA-putrescine, and the glyoxylate pathways significantly increase, suggesting an alternative cardiac means to replenish succinate and malate in MS. Our platform discovers potential therapeutic targets for MS-associated CVD within pathways that were previously unknown to corelate with the disease.

59 BASIC BIOLOGICAL SCIENCES↗

Adaptive hyperparameter updating for training restricted Boltzmann machines on quantum annealers

Restricted Boltzmann Machines (RBMs) have been proposed for developing neural networks for a variety of unsupervised machine learning applications such as image recognition, drug discovery, and materials design. The Boltzmann probability distribution is used as a model to identify network parameters by optimizing the likelihood of predicting an output given hidden states trained on available data. Training such networks often requires sampling over a large probability space that must be approximated during gradient based optimization. Quantum annealing has been proposed as a means to search this space more efficiently which has been experimentally investigated on D-Wave hardware. D-Wave implementation requires selection of an effective inverse temperature or hyperparameter (β) within the Boltzmann distribution which can strongly influence optimization. Here, we show how this parameter can be estimated as a hyperparameter applied to D-Wave hardware during neural network training by maximizing the likelihood or minimizing the Shannon entropy. We find both methods improve training RBMs based upon D-Wave hardware experimental validation on an image recognition problem. Neural network image reconstruction errors are evaluated using Bayesian uncertainty analysis which illustrate more than an order magnitude lower image reconstruction error using the maximum likelihood over manually optimizing the hyperparameter. The maximum likelihood method is also shown to out-perform minimizing the Shannon entropy for image reconstruction.

97 MATHEMATICS AND COMPUTING↗

Dynamic Mode Decomposition of Random Pressure Fields over Bluff Bodies

Fluctuating surface pressures on a bluff body exposed to a boundary layer flow generally are characterized as a spatiotemporally varying random field. In this paper, a dynamic mode decomposition (DMD) was applied to extract dominant features embedded in these random pressure fields. Utilizing an unsupervised machine learning algorithm, spatial modes and their temporal variations were grouped into different clusters at scales, e.g., macro, meso, and micro. A proper orthogonal decomposition (POD) of the experimental data was carried out to observe commonalities and distinctive perspectives each decomposition offers. Here, a comprehensive examination of the DMD/POD for their convergence criteria, data sufficiency, and modal components analysis was conducted. The physical interpretation of the spatiotemporal pressure field based on these decomposition schemes was discussed. At different scales, the DMD modes can capture the evolution of aerodynamic features, e.g., convection of vortices (or vortex tubes) and other structures. The distribution of energy among these three broad scales also reflects an energy cascade in pressure fluctuations akin to turbulence.

97 MATHEMATICS AND COMPUTING↗

X-ray nano-imaging of defects in thin film catalysts via cluster analysis

Functional properties of transition-metal oxides strongly depend on crystallographic defects; crystallographic lattice deviations can affect ionic diffusion and adsorbate binding energies. Scanning x-ray nanodiffraction enables imaging of local structural distortions across an extended spatial region of thin samples. Yet, localized lattice distortions remain challenging to detect and localize using nanodiffraction, due to their weak diffuse scattering. Here, in this study, we apply an unsupervised machine learning clustering algorithm to isolate the low-intensity diffuse scattering in as-grown and alkaline-treated thin epitaxially strained SrIrO 3 films. We pinpoint the defect locations, find additional strain variation in the morphology of electrochemically cycled SrIrO 3 , and interpret the defect type by analyzing the diffraction profile through clustering. Our findings demonstrate the use of a machine learning clustering algorithm for identifying and characterizing hard-to-find crystallographic defects in thin films of electrocatalysts and highlight the potential to study electrochemical reactions at defect sites in operando experiments.

42 ENGINEERING↗