Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “unsupervised machine learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Real-space visualization of a defect-mediated charge density wave transition

Here, we study the coupled charge density wave (CDW) and insulator-to-metal transitions in the 2D quantum material 1T-TaS 2 . By applying in situ cryogenic 4D scanning transmission electron microscopy with in situ electrical resistance measurements, we directly visualize the CDW transition and establish that the transition is mediated by basal dislocations (stacking solitons). We find that dislocations can both nucleate and pin the transition and locally alter the transition temperature T c by nearly ~75 K. This finding was enabled by the application of unsupervised machine learning to cluster five-dimensional, terabyte scale datasets, which demonstrate a one-to-one correlation between resistance—a global property—and local CDW domain-dislocation dynamics, thereby linking the material microstructure to device properties. This work represents a major step toward defect-engineering of quantum materials, which will become increasingly important as we aim to utilize such materials in real devices.

4D-STEM↗

Secondary structure determines electron transport in peptides

Proteins play a key role in biological electron transport, but the structure–function relationships governing the electronic properties of peptides are not fully understood. Despite recent progress, understanding the link between peptide conformational flexibility, hierarchical structures, and electron transport pathways has been challenging. Here, we use single-molecule experiments, molecular dynamics (MD) simulations, nonequilibrium Green’s function-density functional theory (NEGF-DFT), and unsupervised machine learning to understand the role of secondary structure on electron transport in peptides. Our results reveal a two-state molecular conductance behavior for peptides across several different amino acid sequences. MD simulations and Gaussian mixture modeling are used to show that this two-state molecular conductance behavior arises due to the conformational flexibility of peptide backbones, with a high-conductance state arising due to a more defined secondary structure (beta turn or 3 10 helices) and a low-conductance state occurring for extended peptide structures. These results highlight the importance of helical conformations on electron transport in peptides. Conformer selection for the peptide structures is rationalized using principal component analysis of intramolecular hydrogen bonding distances along peptide backbones. Molecular conformations from MD simulations are used to model charge transport in NEGF-DFT calculations, and the results are in reasonable qualitative agreement with experiments. Projected density of states calculations and molecular orbital visualizations are further used to understand the role of amino acid side chains on transport. Overall, our results show that secondary structure plays a key role in electron transport in peptides, which provides broad avenues for understanding the electronic properties of proteins.

Science & Technology - Other Topics↗

Optimizing the shape of photometric redshift distributions with clustering cross-correlations

We present an optimization method for the assignment of photometric galaxies to a chosen set of redshift bins. This is achieved by combining simulated annealing, an optimization algorithm inspired by solid-state physics, with an unsupervised machine learning method, a self-organizing map (SOM) of the observed colours of galaxies. Starting with a sample of galaxies that is divided into redshift bins based on a photometric redshift point estimate, the simulated annealing algorithm repeatedly reassigns SOM-selected subsamples of galaxies, which are close in colour, to alternative redshift bins. We optimize the clustering cross-correlation signal between photometric galaxies and a reference sample of galaxies with well-calibrated redshifts. Depending on the effect on the clustering signal, the reassignment is either accepted or rejected. By dynamically increasing the resolution of the SOM, the algorithm eventually converges to a solution that minimizes the number of mismatched galaxies in each tomographic redshift bin and thus improves the compactness of their corresponding redshift distribution. This method is demonstrated on the synthetic Legacy Survey of Space and Time cosmoDC2 catalogue. We find a significant decrease in the fraction of catastrophic outliers in the redshift distribution in all tomographic bins, most notably in the highest redshift bin with a decrease in the outlier fraction from 57 percent to 16 percent.

79 ASTRONOMY AND ASTROPHYSICS↗

The mass profiles of dwarf galaxies from Dark Energy Survey lensing

We present a novel approach to extracting dwarf galaxies from photometric data to measure their average halo mass profile with weak lensing. We characterize their stellar mass and redshift distributions with a spectroscopic calibration sample. By combining the ${\sim} 5000\,\mathrm{deg}^2$ multiband photometry from the Dark Energy Survey and redshifts from the Satellites Around Galactic Analogs Survey with an unsupervised machine learning method, we select a low-mass galaxy sample spanning redshifts $z\lt 0.3$ and divide it into three mass bins. From low to high median mass, the bins contain [146 420, 330 146, 275 028] galaxies and have median stellar masses of $\log _{10}(M_*/\text{M}_\odot)=\left[8.52\substack{+0.57 -0.76},\, 9.02\substack{+0.50 -0.64},\, 9.49\substack{+0.50 -0.58}\right]$ . We measure the stacked excess surface mass density profiles, $\Delta \Sigma (R)$, of these galaxies using galaxy–galaxy lensing with a signal-to-noise ratio of [14, 23, 28]. Through a simulation-based forward-modelling approach, we fit the measurements to constrain the stellar-to-halo mass relation and find the median halo mass of these samples to be $\log _{10}(M_{\rm halo}/\text{M}_\odot)$ = [$10.67\substack{+0.2 -0.4}$, $11.01\substack{+0.14 -0.27}$, $11.40\substack{+0.08 -0.15}$]. The cold dark matter profiles are consistent with NFW (Navarro, Frenk, and White) profiles over scales ${\lesssim} 0.15 \, {h}^{-1}$ Mpc. We find that ${\sim} 20$ per cent of the dwarf galaxy sample are satellites. This is the first measurement of the halo profiles and masses of such a comprehensive, low-mass galaxy sample. The techniques presented here pave the way for extracting and analysing even lower mass dwarf galaxies and for more finely splitting galaxies by their properties with future photometric and spectroscopic survey data.

dark matter↗

Topological and magnetic properties of the interacting Bernevig-Hughes-Zhang model

We investigate the effects of electronic correlations on the Bernevig-Hughes-Zhang model using the real-space density matrix renormalization group (DMRG) algorithm. We introduce a method to probe topological phase transitions in systems with strong correlations using DMRG, substantiated by an unsupervised machine learning methodology that analyzes the orbital structure of the real-space edges. Including the full multi-orbital Hubbard interaction term, we construct a phase diagram as a function of a gap parameter (m) and the Hubbard interaction strength (U) via exact DMRG simulations on N×4 cylinders. Our analysis confirms that the topological phase persists in the presence of interactions, consistent with previous studies, but it also reveals an intriguing phase transition from a paramagnetic to a stripey antiferromagnetic topological insulator. The combination of the magnetic structure factor, strength of magnetic moments, and the orbitally resolved density, provides real-space information on both topology and magnetism in a strongly correlated system.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

The AXEAP2 program for K β X-ray emission spectra analysis using artificial intelligence

The processing and analysis of synchrotron data can be a complex task, requiring specialized expertise and knowledge. Our previous work addressed the challenge of X-ray emission spectrum (XES) data processing by developing a standalone application using unsupervised machine learning. However, the task of analyzing the processed spectra remains another challenge. Although the non-resonant K β XES of 3 d transition metals are known to provide electronic structure information such as oxidation and spin state, finding appropriate parameters to match experimental data is a time-consuming and labor-intensive process. Here, a new XES data analysis method based on the genetic algorithm is demonstrated, applying it to Mn, Co and Ni oxides. This approach is also implemented as a standalone application, Argonne X-ray Emission Analysis 2 ( AXEAP2 ), which finds a set of parameters that result in a high-quality fit of the experimental spectrum with minimal intervention. AXEAP2 is able to find a set of parameters that reproduce the experimental spectrum, and provide insights into the 3 d electron spin state, 3 d –3 p electron exchange force and K β emission core-hole lifetime.

36 MATERIALS SCIENCE↗

Identifying Outliers in AI-based Image Compression

Image compression using artificial intelligence (AI) is becoming increasingly prevalent across various fields, including scientific research. Scientific instruments can generate hundreds of images per second, and effectively compressing these images with high compression ratios is crucial for facilitating scientific discoveries. However, automatically detecting outlier cases, where compression may not have succeeded or where interesting scientific phenomena are present, poses a significant challenge. To address this, we have developed a methodology based on unsupervised machine learning techniques for detecting outlier compressed images. This methodology utilizes metrics such as peak signal-to-noise ratio (PSNR), structural similarity index measure (SSIM), structural texture similarity index measure (STSIM), and deep image and structural texture similarity index (DISTS). We have evaluated our methodology on several unlabeled datasets, including microscopy and x-ray images, and have successfully identified multiple outlier images using our proposed approach. Furthermore, our approach has enabled us to identify image semantics that are valuable for post-experiment analysis by scientists.

Data Analysis↗

Efficient Clustering of Software Vulnerabilities using Self Organizing Map (SOM)

The common vulnerabilities and exposures (CVE) database was created with a mission to ``identify, define, and catalog publicly disclosed cybersecurity vulnerabilities''. This rich body of information can be used to enable rapid and efficient response to secure and defend cyber operations and protect critical cyber infrastructure. The main goal of this paper is to develop a visual analytics tool to enable deep analysis of CVEs using unsupervised clustering techniques. We enhance our analysis by first mapping CVEs to hierarchical-classes in Common Weakness Enumeration (CWE) using information in the National Vulnerability Database (NVD). Both the mapping and the numerical representation of CVEs are enabled by V2W-BERT, which uses natural language processing of the extensive information in NVD to generate a large tabular database of 137,226 CVE entries from 1999 to 2020, where each CVE is represented by a vector of 768 numerical features. The vectorized data is processed by Self-Organizing Maps (SOM), which is an unsupervised machine learning technique for dimensionality reduction, visual representation and clustering. Using a Torus map of 6417 units, we achieve ~10-fold data compression of ~140k CVEs using SOM. The trained map is further clustered using standard K-means clustering into 138 clusters of CVEs. We conducted a brief investigation of the rich mapping of CVEs to best-matching-units to K-means clusters, as well as CVEs to CWEs. For example, this novel mapping provided insight into the role of CWE-59 and CWE-264 in several CVEs that is otherwise hard to explore in the original data. We conclude that our this novel approach will not only enable deep analysis of the complex relationships between CVEs and CWEs, but also a mechanism to quickly respond to and design mitigation actions for rapidly evolving vulnerabilities that have not been mapped to existing CWEs.

Panchal, Khyati↗

Anomaly detection for MPC forecast in Fleet of Water Heaters

Among residential devices, water heaters consume 20% of home energy use in the United States. Water heaters possess the capability to store energy within their reservoirs, enabling the ability to decouple energy use from hot water use. This capability can be used to reduce energy usage and costs while also supporting grid services. This requires accurate forecasting of the parameters of the water heater such as upper and lower temperatures. In this study, we analyzed the performance and behavior of a water heater model used in the real-world to predict a control mechanism that is implemented in a smart residential neighborhood. The model forecasts are accurate in most cases but not all. In such scenarios, error correction of the model is necessary to further improve model predictive control accuracy. Anomaly detection is the first step of error correction. This study complements existing research by grouping time series data into two clusters one with anomalies and another without anomalies. To achieve this task, we explored and compared multiple unsupervised machine learning algorithms to perform clustering. Among these algorithms, Ward clustering has the lowest running time and identified the highest number of anomalies for the upper temperature limit. The proposed approach is tested based on the data collected in a neighborhood with 46 townhomes located in Atlanta, GA.

Lebakula, Viswadeep↗

Event-Based Analysis of Solar Power Distribution Feeder Using Micro-PMU Measurements

Solar distribution feeders are commonly used in solar farms that are integrated into distribution substations. In this paper, we focus on a real-world solar distribution feeder and conduct an event-based analysis by using micro-PMU measurements. The solar distribution feeder of interest is a behind-the-meter solar farm with a generation capacity of over 4 MW that has about 200 low-voltage distributed photovoltaic (PV) inverters. The event-based analysis in this study seeks to address the following practical matters. First, we conduct event detection by using an unsupervised machine learning approach. For each event, we determine the event’s source region by an impedancebased analysis, coupled with a descriptive analytic method. We segregate the events that are caused by the solar farm, i.e., locallyinduced events, versus the events that are initiated in the grid, i.e., grid-induced events, which caused a response by the solar farm. Second, for the locally-induced events, we examine the impact of solar production level and other significant parameters to make statistical conclusions. Third, for the grid-induced events, we characterize the response of the solar farm; and make comparisons with the response of an auxiliary neighboring feeder to the same events. Fourth, we scrutinize multiple specific events; such as by revealing the dynamics to the control system of the solar distribution feeder. The results and discoveries in this study are informative to utilities and solar power industry.

14 SOLAR ENERGY↗

Autonomous Anomaly Detection for MPC Forecasts of HVAC Systems in Residential Communities

The use of residential heating, ventilation, and air conditioning (HVAC) to shift peak demand or provide ancillary services is a potential solution in the presence of older grids and distributed renewables. However, to ensure the efficient use of devices, utilities need to accurately forecast the load and adopt error correction schemes when necessary. While significant theoretical research exists in the area of predictive control of HVAC, little experimental evidence exists. The lack of experimental data in turn causes researchers to be unprepared for unsystematic errors which emerge due to the higher complexity of the data generating process. This study offers an anomaly detection methodology that uses unsupervised machine learning algorithms to detect and isolate these errors with different forecast error ranges. The results of anomaly detection procedure can then be used for error correction and would eventually help develop better predictive controllers. The methodology is tested using real world data from a smart neighborhood that currently operates in Atlanta. GA.

Lebakula, Viswadeep↗

Lowering of Tc in Van Der Waals Layered Materials Under In-Plane Strain

The dependence of electromechanical behavior on strain in ferroelectric materials can be leveraged as parameter to tune ferroelectric properties such as the Curie temperature. For van der Waals materials, a unique opportunity arises because of wrinkling, bubbling, and Moiré phenomena accessible due to structural properties inherent to the van der Waals gap. Here, we use piezoresponse force microscopy and unsupervised machine learning methods to gain insight into the ferroelectric properties of layered CuInP2S6 where local areas are strained in-plane due to a partial delamination, resulting in a topographic bubble feature. We observe significant differences between strained and unstrained areas in piezoresponse images as well as voltage spectroscopy, during which strained areas show a sigmoid-shaped response usually associated with the response measured around the Curie temperature, indicating a lowering of the Curie temperature under tensile strain. These results suggest that strain engineering might be used to further increase the functionality of CuInP2S6 through locally modifying ferroelectric properties on the micro- and nanoscale.

36 MATERIALS SCIENCE↗

REC protein family expansion by the emergence of a new signaling pathway

This report presents multi-genome evidence that REC protein family expansion occurs when the emergence of new pathways gives rise to functional discordance. Specificity between residues in REC domain containing response regulators with paired histidine kinases is under negative purifying selection, constrained by the presence of other bacterial two-component systems signaling cascades that share sequence and structural identity. Presuming that the two-component systems can evolve by neutral amino acid changes (neutral drift) when purifying evolutionary constraints are relaxed, how might the REC protein family expand by amino acid changes when these constraints remain intact? Using an unsupervised machine learning approach to observe the sequence landscape of REC domains across long phylogenetic distances, we find that within-gene recombination, a subcategory of gene conversion, switched the effector domain and, consequently, the regulatory context of a duplicated response regulator from transcriptional regulation by σ54 to that by σ70. We determined that the recombined response regulator diverged from its parent by episodic diversifying selection and neutral drift. Functional experiments of the parent of recombined response regulators in a model Pseudomonas putida KT2440 model system revealed that the parent and recombined response regulators sense and respond to different carboxylic acids. Finally, a residue-switching experiment using structural predictions and functional characterization suggests that the new residues in the recombined regulator could form a new interaction interface and mediate condition-specific phosphotransfer. Overall, our study finds that genetic perturbations can create conditions of functional discordance, whereby the REC protein family can evolve by episodic diversifying selection.

59 BASIC BIOLOGICAL SCIENCES↗

Search for Beyond the Standard Model physics with anomaly detection in multilepton final states in pp collisions at s=13TeV with the ATLAS detector

A model-agnostic search for Beyond the Standard Model physics is presented, targeting final states with at least four light leptons (electrons or muons). The search regions are separated by event topology and unsupervised machine learning is used to identify anomalous events in the full 140 fb-1$$^{-1}$$ of proton–proton collision data collected with the ATLAS detector during Run 2. No significant excess above the Standard Model background expectation is observed. Model-agnostic limits are presented in each topology, along with limits on several benchmark models including vector-like leptons, wino-like charginos and neutralinos, or smuons. Limits are set on the flavourful vector-like lepton model for the first time.

Aad, G↗

Enhancing the hunt for new phenomena in dijet final states using anomaly detection filters at the high-luminosity large Hadron Collider

In the realm of dijet searches in high-energy physics, a significant challenge has emerged: with experiments producing more and more data, the traditional methods of using analytic functions to describe dijet mass spectra start to fail. Here, to address this, we suggest the application of an anomaly detection approach to eliminate less interesting background events based on event final states. This method not only bypasses the limitations of conventional background models but also significantly enhances our ability to detect potential signals of new physics. Through simulations that mimic the conditions of the upcoming high-luminosity large Hadron collider, we demonstrate the strength and efficiency of this approach in dealing with large data volumes. The integration of unsupervised machine learning into our experimental framework paves the way for a promising avenue to unveil hidden physics discoveries within the overwhelming influx of data.

47 OTHER INSTRUMENTATION↗

Data-Driven Clustering and Classification of Outage Patterns with Insights into their Links to Extreme Events

At a global level extreme events have increased in both scale and impact. These events have the potential to affect the electrical grid infrastructure and cause a wide range of outages, which can lead to a disruption in daily patterns, cost millions of dollars and also the loss of life. Currently, to track these outage events there have been various approaches developed ranging from regional to national level quantifications for what defines an outage. However, this variation in methods can potentially lead to subjective decision-making and a lack of proper management in relation to the event. While previous work has made strides in determining spatio-temporal patterns, minimal attention has been given to the type and number of outages an area may be exposed to. The differences in incurred cost and the overall severity of an event between a transformer box malfunction and a hurricane are drastic, and by finding historical signals, we can allow for more efficient management, potentially saving lives and millions of dollars. Here, we leverage unsupervised machine learning techniques to delineate outage patterns among 22 counties within the United States and find that there are clear, segregated clusters (0.93 silhouette) of data which are related by event behavior and underlying cause. This finding will allow for energy stakeholders, policy makers, and researchers to gain a deeper understanding of the extent and severity of historic events and to better prepare for electrical grid infrastructure planning and management.

Koob, Benjamin [ORNL]↗

ENSIGN

ENSIGN is a data analytics software package offering a modern unsupervised machine learning solution for scalable discovery in Big Data. The analytics in ENSIGN are based on an advanced mathematical tool called tensor decomposition and they are optimized to run efficiently on a range of computing platforms (from small multicore Desktop platforms to large Supercomputing clusters and novel high-end memory-driven computing platforms such as HPE Superdome Flex). ENSIGN enables the user to extract deep insights from the entirety of massive-scale (100s of Gigabytes or Terabytes scale) multidimensional data. ENSIGN uncovers latent patterns in data without the user having to specify or describe what the patterns are; the user, in the first place, may not even know such patterns existed and that they have to look for such patterns. The insights gained from ENSIGN could be trailheads that can be used as starting points for deeper forensic investigation.

Baskaran, Muthu↗

Dynamic Ride-Matching for Large-Scale Transportation Systems

Efficient dynamic ride-matching (DRM) in large-scale transportation systems is a key driver in transport simulations to yield answers to challenging problems. Although the DRM problem is simple to solve, it quickly becomes a computationally challenging problem in large-scale transportation system simulations. Therefore, this study thoroughly examines the DRM problem dynamics and proposes an optimization-based solution framework to solve the problem efficiently. To benefit from parallel computing and reduce computational times, the problem’s network is divided into clusters utilizing a commonly used unsupervised machine learning algorithm along with a linear programming model. Then, these sub-problems are solved using another linear program to finalize the ride-matching. At the clustering level, the framework allows users adjusting cluster sizes to balance the trade-off between the computational time savings and the solution quality deviation. A case study in the Chicago Metropolitan Area, U.S., illustrates that the framework can reduce the average computational time by 58% at the cost of increasing the average pick up time by 26% compared with a system optimum, that is, non-clustered, approach. Another case study in a relatively small city, Bloomington, Illinois, U.S., shows that the framework provides quite similar results to the system-optimum approach in approximately 62% less computational time.

33 ADVANCED PROPULSION SYSTEMS↗