Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “pattern clustering”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Identifying Climate Patterns Using Clustering Autoencoder Techniques

Abstract The complexity of growing spatiotemporal resolution of climate simulations produces a variety of climate patterns under different projection scenarios. This paper proposes a new data-driven climate classification workflow via an unsupervised deep learning technique that can dimensionally reduce the vast volume of spatiotemporal numerical climate projection data into a compact representation. We aim to identify distinct zones that capture multiple climate variables as well as their future changes under different climate change scenarios. Our approach leverages convolutional autoencoders combined with k -means clustering (standard autoencoder) and online clustering based on the Sinkhorn–Knopp algorithm (clustering autoencoder) across the conterminous United States (CONUS) to capture unique climate patterns in a data-driven fashion from the Geophysical Fluid Dynamics Laboratory Earth System Model with GOLD component (GFDL-ESM2G). The developed approach compresses 70 years of GFDL-ESM2G simulation at 0.125° spatial resolution across the CONUS under multiple warming scenarios to a lower-dimensional space by a factor of 660 000 and then tested on 150 years of GFDL-ESM2G simulation data. The results show that five climate clusters capture physically reasonable and spatially stable climatological patterns matched to known climate classes defined by human experts. Results also show that using a clustering autoencoder can reduce the computational time for clustering by up to 9.2 times when compared to using a standard autoencoder. Our five unique climate patterns resulting from the deep learning–based clustering of the lower-dimensional space thereby enable us to provide insights on hydrometeorology and its spatial heterogeneity across the conterminous United States immediately without downloading large climate datasets. Significance Statement This paper presents a data-driven climate classification approach using unsupervised deep learning to dimensionally reduce climate model outputs and to identify distinct climate regions for their future changes. Our approach compresses climate information for 70 years of Geophysical Fluid Dynamics Laboratory Earth System Model data across the conterminous United States (CONUS) at 0.125° spatial resolution. The results reveal that five climate clusters capture reasonable and stable climatological patterns matched to known climate patterns. The embedded clustering process in deep learning provides ×9.2 times faster execution than the k -means clustering technique. These results give us insight about climate spatial patterns and heterogeneity of hydrological patterns across the conterminous United States without downloading large climate datasets.

Kurihana, Takuya↗

A Taxonomic Classification Approach for Global Spatio-temporal Data

The World Bank, World Health Organization, and other major vendors collectively provide thousands of global time series datasets that focus on issues of the environment, public health, economics, violence, education, and national security. Sorting these data into meaningful information requires the use of data mining techniques to cluster trends into an orderly and manageable number of cases. The World SpatioTemporal Analytics and Mapping (WSTAMP) project database (wstamp.ornl.gov) was developed to spatiotemporally harmonize global vendor data (23,300+ attributes, 200+ locations, 50+ years). Within the WSTAMP analytical environment, Dynamic Time Warping (DTW) has been a highly effective data-driven approach for clustering and mapping these time series into national spatiotemporal behavior maps. Two significant properties have surfaced from this work. First, several recognizable cluster patterns have emerged and persist across a range of locations, attributes, and time frames (e.g., increasing, decreasing, rebounding, peak, oscillating). Secondly, practitioners engaging WSTAMP have noted the explanatory and anticipatory value of these patterns and articulated particular interest in detecting them within the spatiotemporal cube. This need was addressed by shifting DTW-based clustering from an open ended, data-driven implementation to a taxonomic pattern matching approach. This paper presents the method including implementation strategies for visualization and human computer interaction and applies the approach to a sample data set and concludes with next steps.

Stewart, Robert↗

The Tianlai-WIYN North Celestial Cap Redshift Survey

We present the results of a small, low redshift spectroscopic survey of galaxies within 3 degrees of the North Celestial Pole (NCP) selected using V-band photometry obtained from the North Celestial Cap Survey (NCCS) (Gorbikov & Brosch 2014). The purpose of the current survey is to create a redshift space template for 21 cm emission from neutral hydrogen with which to correlate radio line intensity observations by the Tianlai dish and cylinder interferometers. A total of 898 redshifts were obtained from the 2102 extended objects in the NCCS with m_V < 19 in the survey area. After accounting for extinction, the survey geometry and selection effects, the number density and clustering pattern of galaxies in the redshift catalog are consistent with other low redshift surveys. We were also able to identify 11 galaxy cluster candidates from this redshift catalog.

Ansari, Reza [AIM, Saclay]↗

Graph Analytics on Jellyfish topology

Because large unstructured datasets is important for many science domains, distributed graph analytics is critical to many scientists. Unfortunately, obtaining scaling and performance for irregular communication is challenging because contemporary network interconnects are primarily designed to maximize bandwidths of fixed-neighborhoods large-message exchanges (e.g., stencils). Although there is no consensus on the “best” network topologies for irregular communication, unstructured graph-based interconnects can be more suitable. We analyze three popular graph workloads – clustering, pattern enumeration, and traversal — on comparable networks (in terms of resources and costs) constructed from Jellyfish Random Regular, Dragonfly and Fat tree topologies, varying the routing algorithms. Using packet-level simulations, we demonstrate up to 60% improvement in communication time with Jellyfish due to diversity of the short paths between arbitrary endpoints, which can reduce overall network stalls and congestion.

Graph Analytics, network topology, interconnect, H↗

Q-BEEP: Quantum Bayesian Error Mitigation Employing Poisson Modeling over the Hamming Spectrum

Quantum computing technology has grown rapidly in recent years, with new technologies being explored, error rates being reduced, and quantum processor’s qubit capacity growing. However, near-term quantum algorithms are still unable to be induced without compounding consequential levels of noise, leading to non-trivial erroneous results. Quantum Error Correction (in-situ error mitigation) and Quantum Error Mitigation (post-induction error mitigation) are promising fields of research within the quantum algorithm scene, aiming to alleviate quantum errors, increasing the overall fidelity and hence the overall quality of circuit induction. Earlier this year, a pioneering work, namely HAMMER, published in ASPLOS-22 demonstrated the existence of a latent structure regarding post-circuit induction errors when mapping to the Hamming spectrum. However, they intuitively assumed that errors occur in local clusters, and that at higher average Hamming distances this structure falls away. In this work, we show that such a correlation structure is not only local but extends certain non-local clustering patterns which can be precisely described by a Poisson distribution model taking the input circuit, the device run time status (i.e., calibration statistics) and qubit topology into consideration. Using this quantum error characterizing model, we developed an iterative algorithm over the generated Bayesian network state-graph for post-induction error mitigation. Thanks to more precise modeling of the error distribution latent structure and the new iterative method, our Q-Beep approach provides state of the art performance and can boost circuit execution fidelity by up to 234.6% on Bernstein-Vazirani circuits and on average 71.0% on QAOA solution quality, using 16 practical IBMQ quantum processors. For other benchmarks such as those in QASMBench, the fidelity improvement is up to 17.8%. Q-Beep is a light-weight post-processing technique that can be performed offline and remotely, making it a useful tool for quantum vendors to integrate and provide more reliable circuit induction results.

Stein, Samuel A.↗

Streamlining heterologous expression of top carbonic anhydrases in Escherichia coli : bioinformatic and experimental approaches

Carbonic anhydrase (CA) enzymes facilitate the reversible hydration of CO 2 to bicarbonate ions and protons. Identifying efficient and robust CAs and expressing them in model host cells, such as Escherichia coli, enables more efficient engineering of these enzymes for industrial CO 2 capture. However, expression of CAs in E. coli is challenging due to the possible formation of insoluble protein aggregates, or inclusion bodies. This makes the production of soluble and active CA protein a prerequisite for downstream applications. In this study, we streamlined the process of CA expression by selecting seven top CA candidates and used two bioinformatic tools to predict their solubility for expression in E. coli. The prediction results place these enzymes in two categories: low and high solubility. Our expression of high solubility score CAs (namely CA5-SspCA, CA6-SazCAtrunc, CA7-PabCA and CA8-PhoCA) led to significantly higher protein yields (5 to 75 mg purified protein per liter) in flask cultures, indicating a strong correlation between the solubility prediction score and protein expression yields. Furthermore, phylogenetic tree analysis demonstrated CA class-specific clustering patterns for protein solubility and production yields. Unexpectedly, we also found that the unique N-terminal, 11-amino acid segment found after the signal sequence (not present in its homologs), was essential for CA6-SazCA activity. Overall, this work demonstrated that protein solubility prediction, phylogenetic tree analysis, and experimental validation are potent tools for identifying top CA candidates and then producing soluble, active forms of these enzymes in E. coli. The comprehensive approaches we report here should be extendable to the expression of other heterogeneous proteins in E. coli.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Test Vector Development for Verification and Validation of Heavy-Duty Autonomous Vehicle Operations

The current focus in the ongoing development of autonomous driving systems (ADS) for heavy duty vehicles is that of vehicle operational safety. To this end, developers and researchers alike are working towards a complete understanding of the operating environments and conditions that autonomous vehicles are subject to during their mission. This understanding is critical to the testing and validation phases of the development of autonomous vehicles and allows for the identification of both the nominal and edge case scenarios encountered by these systems. Previous work by the authors saw the development of a comprehensive scenario generation framework to identify an operating domain specification (ODS), or external and internal conditions an autonomous driving system can expect to encounter on its mission to form critical scenario groups for autonomous vehicle testing and validating using statistical patterns, clustering, and correlation. Continuing this prior work, this paper focuses on the generation of test cases based on the critical scenarios identified that can be used to prioritize either the most common nominal driving scenarios or the least common severe driving scenarios. These test cases can then be used validate, through simulation or real-world testing, the operating design domain (ODD) for a generalized driving mission and built upon to identify spatial and temporal impacts on the driving mission of an autonomous vehicle.

Siekmann, Adam↗

SCULPT (Supervised Clustering and Uncovering Latent Patterns with Training) v1

SCULPT (Supervised Clustering and Uncovering Latent Patterns with Training) is a comprehensive data visualization and analysis application focused on working with COLTRIMS (COLd Target Recoil Ion Momentum Spectroscopy) data, which is used in atomic and molecular physics experiments. The application offers several powerful features: - Data uploading and processing capabilities for COLTRIMS files - Multiple visualization methods using UMAP (Uniform Manifold Approximation and Projection) for dimensionality reduction - Interactive selection of data points across multiple views - Feature engineering through various methods: - Manual feature selection from calculated physics parameters - Deep autoencoder for dimension reduction - Genetic programming for discovering meaningful features - Mutual information-based feature selection - Multiple clustering approaches (DBSCAN, KMeans, Agglomerative) - Quality metrics for evaluating clustering results - Export capabilities for selections and generated features

Daoud, Hazem [Lawrence Berkeley National Laborator↗

Field-driven cluster formation in two-dimensional colloidal binary mixtures

Here, we study size- and charge-asymmetric oppositely charged colloids driven by an external electric field. The large particles are connected by harmonic springs, forming a hexagonal-lattice network, while the small particles are free of bonds and exhibit fluidlike motion. We show that this model exhibits a cluster formation pattern when the external driving force exceeds a critical value. The clustering is accompanied with stable wave packets in vibrational motions of the large particles.

42 ENGINEERING↗

A multi-level load shape clustering and disaggregation approach to characterize patterns of energy consumption behavior

This study presents representative electrical load shapes, disaggregated to the end-use level, for over 5000 customer clusters across California’s residential, commercial, industrial and agricultural sectors. We developed a novel, multi-level load shape clustering approach for residential and commercial sectors leveraging interval meter data for over 350,000 California utility customers collected as a part of the Phase 4 California Demand Response (DR) Potential Study. The clustering approach allowed us to identify typical consumption patterns and categorize customers based on their daily load shape displayed throughout the year. For example, we were able to identify customers with particular energy technologies such as electric vehicles and rooftop solar, as well as building occupancy types such as restaurants, grocery stores and even unoccupied buildings, based solely on whole-building interval data. We then combined the load shape-based clusters with other customer information including building type, climate, geographical area, total consumption and low-income status, to create a set of customer clusters based on both demographics and usage patterns. Total cluster electricity demand was then disaggregated into a wide variety of end-uses using weather normalization and other publicly available end-use load shape datasets. The resulting disaggregated cluster load shapes will be released in anonymized form as part of the Phase 4 DR Potential Study. They will have wide-ranging applications in energy research and policy analysis, including estimation of energy efficiency (EE) and DR potential on the end-use level, time-dependent valuation of EE savings, building stock modeling, and developing customer targeting strategies for EE and DR programs.

Murthy, Samanvitha↗

Data-Driven Clustering and Classification of Outage Patterns with Insights into their Links to Extreme Events

At a global level extreme events have increased in both scale and impact. These events have the potential to affect the electrical grid infrastructure and cause a wide range of outages, which can lead to a disruption in daily patterns, cost millions of dollars and also the loss of life. Currently, to track these outage events there have been various approaches developed ranging from regional to national level quantifications for what defines an outage. However, this variation in methods can potentially lead to subjective decision-making and a lack of proper management in relation to the event. While previous work has made strides in determining spatio-temporal patterns, minimal attention has been given to the type and number of outages an area may be exposed to. The differences in incurred cost and the overall severity of an event between a transformer box malfunction and a hurricane are drastic, and by finding historical signals, we can allow for more efficient management, potentially saving lives and millions of dollars. Here, we leverage unsupervised machine learning techniques to delineate outage patterns among 22 counties within the United States and find that there are clear, segregated clusters (0.93 silhouette) of data which are related by event behavior and underlying cause. This finding will allow for energy stakeholders, policy makers, and researchers to gain a deeper understanding of the extent and severity of historic events and to better prepare for electrical grid infrastructure planning and management.

Koob, Benjamin [ORNL]↗

Ion Clusters Reveal the Sources, Impacts, and Drivers of Freshwater Salinization

Population growth, land use change, climate change, and natural resource extraction are driving the salinization of freshwater resources worldwide. Reversing these trends will require data-centric approaches that identify salt sources, environmental drivers, and ecosystem responses. In this study, we applied principal component analysis and hierarchical clustering to identify ion covariance patterns, or “ion clusters,” in Broad Run, an urban stream in the Mid-Atlantic United States. These clusters correspond to distinct hydrologic regimes and reveal specific salinization risks: (1) phosphorus pollution mobilized during summer storms (Cluster 1); (2) elevated concentrations of sulfate and bicarbonate during baseflow (Cluster 2), likely reflecting groundwater discharge; and (3) elevated specific conductance and sodium, chloride, and potassium ion concentrations during snowmelt and rain-on-snow events (Cluster 3), driven by deicer and anti-icer wash-off. These ion fingerprints offer a transferable framework for diagnosing salt sources, assessing ecological risk, and identifying management targets. Our findings underscore the need for next-generation stormwater infrastructure and smart growth policies to protect aquatic life in rapidly urbanizing watersheds.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Exploring the S-process History in the Galactic Disk: Cerium Abundances and Gradients in Open Clusters from the OCCAM/APOGEE Sample

The APOGEE Open Cluster Chemical Abundances and Mapping survey is used to probe the chemical evolution of the s-process element cerium in the Galactic disk. Cerium abundances were derived from measurements of Ce ii lines in the APOGEE spectra using the Brussels Automatic Code for Characterizing High Accuracy Spectra in 218 stars belonging to 42 open clusters. Our results indicate that, in general, for ages < 4 Gyr, younger open clusters have higher [Ce/Fe] and [Ce/α-element] ratios than older clusters. In addition, metallicity segregates open clusters in the [Ce/X]–age plane (where X can be H, Fe, or the α-elements O, Mg, Si, or Ca). These metallicity-dependent relations result in [Ce/Fe] and [Ce/α] ratios with ages that are not universal clocks. Radial gradients of [Ce/H] and [Ce/Fe] ratios in open clusters, binned by age, were derived for the first time, with d[Ce/H]/dR GC being negative, while d[Ce/Fe]/dR GC is positive. [Ce/H] and [Ce/Fe] gradients are approximately constant over time, with the [Ce/Fe] gradient becoming slightly steeper, changing by ~+0.009 dex kpc -1 Gyr -1 . Both the [Ce/H] and [Ce/Fe] gradients are shifted to lower values of [Ce/H] and [Ce/Fe] for older open clusters. The chemical pattern of Ce in open clusters across the Galactic disk is discussed within the context of s-process yields from asymptotic giant branch (AGB) stars, gigayear time delays in Ce enrichment of the interstellar medium, and the strong dependence of Ce nucleosynthesis on the metallicity of its AGB stellar sources.

79 ASTRONOMY AND ASTROPHYSICS↗

The Role of Ligand–Ligand Interactions in Multimodal Ligand Conformational Equilibria and Surface Pattern Formation

Multimodal chromatography uses multiple modes of interaction such as charge, hydrophobic, or hydrogen bonding to separate proteins. Recently, we used molecular dynamics (MD) simulations to show that ligands immobilized on surfaces can interact and associate with neighboring ligands to form hydrophobic and charge patches, which may have important implications for the nature of protein–surface interactions. Here, we study interfacial systems of increasing complexity—from a single immobilized multimodal ligand to high density surfaces—to better understand how ligand behavior is affected by the presence of a surface and the presence of other ligands in the vicinity, and how this behavior scales to larger systems. Furthermore, we find that tethering a ligand to a surface restricts its conformations to a subset of those observed in free solution, yet the ligand maintains flexibility in the plane of the surface and can form contacts with neighboring ligands. We find that although the formation of a contact between two neighboring ligands is slightly unfavorable, three neighboring ligands exhibit a preference for the formation of a fully connected cluster. To explore how these trends in ligand association extend to a larger surface with high density of ligands, we performed coarse-grained Monte Carlo (MC) simulations of a 132-ligand surface using ligand interactions parametrized based on free energies obtained from the three-ligand MD simulations. Despite their simplicity, the coarse-grained simulations qualitatively capture the cluster size distribution of ligands observed in detailed MD simulations. Quantitative differences between the two suggest opportunities for improvements in the coarse-grained energy function for efficient predictions of cluster and pattern formations. Our approach presents a promising route to the engineering of multimodal patterns for future chromatographic resin design.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Collective effects and pattern formation for directional locking of disks moving through obstacle arrays

We examine directional locking effects in an assembly of disks driven through a square array of obstacles as the angle of drive rotates from 0° to 90°. For increasing disk densities, the system exhibits a series of different dynamic patterns along certain locking directions, including one-dimensional or multiple-row chain phases and density-modulated phases. For nonlocking driving directions, the disks form disordered patterns or clusters. Here, when the obstacles are small or far apart, a large number of locking phases appear; however, as the number of disks increases, the number of possible locking phases drops due to the increasing frequency of collisions between the disks and obstacles. For dense arrays or large obstacles, we find an increased clogging effect in which immobile and moving disks coexist.

97 MATHEMATICS AND COMPUTING↗

Phase identification using co‐association matrix ensemble clustering

Calibrating distribution system models to aid in the accuracy of simulations such as hosting capacity analysis is increasingly important in the pursuit of the goal of integrating more distributed energy resources. The recent availability of smart meter data is enabling the use of machine learning tools to automatically achieve model calibration tasks. This research focuses on applying machine learning to the phase identification task, using a co‐association matrix‐based, ensemble spectral clustering approach. The proposed method leverages voltage time series from smart meters and does not require existing or accurate phase labels. This work demonstrates the success of the proposed method on both synthetic and real data, surpassing the accuracy of other phase identification research.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Distinct cellular expression and subcellular localization of Kv2 voltage‐gated K + channel subtypes in dorsal root ganglion neurons conserved between mice and humans

Abstract The distinct organization of Kv2 voltage‐gated potassium channels on and near the cell body of brain neurons enables their regulation of action potentials and specialized membrane contact sites. Somatosensory neurons have a pseudounipolar morphology and transmit action potentials from peripheral nerve endings through axons that bifurcate to the spinal cord and the cell body within ganglia including the dorsal root ganglia (DRG). Kv2 channels regulate action potentials in somatosensory neurons, yet little is known about where Kv2 channels are located. Here, we define the cellular and subcellular localization of the Kv2 paralogs, Kv2.1 and Kv2.2, in DRG somatosensory neurons with a panel of antibodies, cell markers, and genetically modified mice. We find that relative to spinal cord neurons, DRG neurons have similar levels of detectable Kv2.1 and higher levels of Kv2.2. In older mice, detectable Kv2.2 remains similar, while detectable Kv2.1 decreases. Both Kv2 subtypes adopt clustered subcellular patterns that are distinct from central neurons. Most DRG neurons co‐express Kv2.1 and Kv2.2, although neuron subpopulations show preferential expression of Kv2.1 or Kv2.2. We find that Kv2 protein expression and subcellular localization are similar between mouse and human DRG neurons. We conclude that the organization of both Kv2 channels is consistent with physiological roles in the somata and stem axons of DRG neurons. The general prevalence of Kv2.2 in DRG as compared to central neurons and the enrichment of Kv2.2 relative to detectable Kv2.1 in older mice, proprioceptors, and axons suggest more widespread roles for Kv2.2 in DRG neurons.

Neurosciences & Neurology↗