Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “distributed clustering methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Exploring the Landscape of Distributed Graph Clustering on Leadership Supercomputers

The rapid growth of large-scale datasets in fields like biology and social networks has driven the need for advanced graph analytics techniques. Community detection, a fundamental task in graph analytics, identifies closely connected groups of nodes within a network, providing valuable insights across various disciplines. This study focuses on two classic community detection methods, the Louvain algorithm and Markov Clustering (MCL), and evaluates the performance of two prominent distributed community detection algorithms: HiPDPL-GPU, our prior implementation, and HipMCL. We conduct experiments on GPU-accelerated heterogeneous HPC systems, Summit and Frontier, to assess their performance under varying conditions. Our objective is to identify the strengths and weaknesses of these algorithms in terms of scalability, and quality of solutions. We evaluate these algorithms on a diverse set of 70+ networks spanning 13 domains, with sizes ranging up to 4.2 billion edges. Our results demonstrate that HiPDPL-GPU consistently outperforms HipMCL, especially for large-scale networks. HiPDPL-GPU achieves significantly faster runtimes (47x to 1439x), higher modularity scores, and improved scalability. These findings highlight HiPDPL-GPU as a promising solution for efficient and effective large-scale graph analytics in diverse application domains, and provide insights into the feasibility of using MCL-based approaches for certain application domains.

Community detection, graph algorithms↗

LBNL CRADA (FP00009949) with the American Public Power Association: Electricity Reliability Metrics, Analysis, and Planning (Final Technical Report)

LBNL and APPA (the team) jointly examined the extent to which differences in distribution feeder characteristics are correlated with differences in their reliability performance when exposed to three different types of natural hazards (wildlife, weather, and vegetation). The team employed data-driven approaches to quantify the relationships between various measures of feeder reliability and a suite of feeder characteristics individually and jointly via a statistically-based clustering method.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Substructure in the stellar halo near the Sun: II. Characterisation of independent structures

In an accompanying paper, we present a data-driven method for clustering in ‘integrals of motion’ space and apply it to a large sample of nearby halo stars with 6D phase-space information. The algorithm identified a large number of clusters, many of which could tentatively be merged into larger groups. The goal here is to establish the reality of the clusters and groups through a combined study of their stellar populations (average age, metallicity, and chemical and dynamical properties) to gain more insights into the accretion history of the Milky Way. To this end, we developed a procedure that quantifies the similarity of clusters based on the Kolmogorov–Smirnov test using their metallicity distribution functions, and an isochrone fitting method to determine their average age, which is also used to compare the distribution of stars in the colour–absolute magnitude diagram. Also taking into consideration how the clusters are distributed in integrals of motion space allows us to group clusters into substructures and to compare substructures with one another. We find that the 67 clusters identified by our algorithm can be merged into 12 extended substructures and 8 small clusters that remain as such. The large substructures include the previously known Gaia-Enceladus, Helmi streams, Sequoia, and Thamnos 1 and 2. We identify a few over-densities that can be associated with the hot thick disc and host a small metal-poor population. Especially notable is the largest (by number of member stars) substructure in our sample which, although peaking at the metallicity characteristic of the thick disc, has a very well populated metal-poor component, and dynamics intermediate between the hot thick disc and the halo. We also identify additional debris in the region occupied by Sequoia with clearly distinct kinematics, likely remnants of three different accretion events with progenitors of similar masses. Although only a small subset of the stars in our sample have chemical abundance information, we are able to identify different trends of [Mg/Fe] versus [Fe/H] for the various substructures, confirming our dissection of the nearby halo. We find that at least 20% of the halo near the Sun is associated to substructures. When comparing their global properties, we note that those substructures on retrograde orbits are not only more metal-poor on average but are also older. We provide a table summarising the properties of the substructures, as well as a membership list that can be used for follow-up chemical abundance studies for example.

79 ASTRONOMY AND ASTROPHYSICS↗

Precise relative magnitude measurement improves fracture characterization during hydraulic fracturing

SUMMARY Microseismic monitoring is an important technique to obtain detailed knowledge of in-situ fracture size and orientation during stimulation to maximize fluid flow throughout the rock volume and optimize production. Furthermore, considering that the frequency of earthquake magnitudes empirically follows a power law (i.e. Gutenberg–Richter), the accuracy of microseismic event magnitude distributions is potentially crucial for seismic risk management. In this study, we analyse microseismicity observed during four hydraulic fracture treatments of the legacy Cotton Valley experiment in 1997 at the Carthage gas field of East Texas, where fractures were activated at the base of the sand-shale Upper Cotton Valley formation. We perform waveform cross-correlation to detect similar event clusters, measure relative amplitude from aligned waveform pairs with a principal component analysis, then measure precise relative magnitudes. The new magnitudes significantly reduce the deviations between magnitude differences and relative amplitudes of event pairs. This subsequently reduces the magnitude differences between clusters located at different depths. Reduction in magnitude differences between clusters suggests that some attenuation-related biases could be effectively mitigated with relative magnitude measurements. The maximum likelihood method is applied to understand the magnitude frequency distributions and quantify the seismogenic index of the clusters. Statistical analyses with new magnitudes suggest that fractures that are more favourably oriented for shear failure have lower b-value and higher seismogenic index, suggesting higher potential for relatively larger earthquakes, rather than fractures subparallel to maximum horizontal principal stress orientation.

58 GEOSCIENCES↗

Constraining Galaxy-Halo connection using machine learning

We investigate the potential of machine learning (ML) methods to model small-scale galaxy clustering for constraining Halo Occupation Distribution (HOD) parameters. Our analysis reveals that while many ML algorithms report good statistical fits, they often yield likelihood contours that are significantly biased in both mean values and variances relative to the true model parameters. This highlights the importance of careful data processing and algorithm selection in ML applications for galaxy clustering, as even seemingly robust methods can lead to biased results if not applied correctly. ML tools offer a promising approach to exploring the HOD parameter space with significantly reduced computational costs compared to traditional brute-force methods if their robustness is established. Using our ANN-based pipeline, we successfully recreate some standard results from recent literature. Properly restricting the HOD parameter space, transforming the training data, and carefully selecting ML algorithms are essential for achieving unbiased and robust predictions. Among the methods tested, artificial neural networks (ANNs) outperform random forests (RF) and ridge regression in predicting clustering statistics, when the HOD prior space is appropriately restricted. We demonstrate these findings using the projected two-point correlation function (w p (r p )), angular multipoles of the correlation function (ξ ℓ (r)), and the void probability function (VPF) of Luminous Red Galaxies from Dark Energy Spectroscopic Instrument mocks. Our results show that while combining w p (r p ) and VPF improves parameter constraints, adding the multipoles ξ 0 , ξ 2 , and ξ 4 to w p (r p ) does not significantly improve the constraints.

cosmology↗

Scalable edge clustering of dynamic graphs via weighted line graphs

Timestamped relational datasets consisting of records (or connections) between pairs of entities are ubiquitous in network science. For applications like peer-to-peer communication, email, various social network interactions, and computer network security, it is useful to organize these records into groups based on how and when they are occurring. Weighted line graphs offer a natural way to model how records are related in such datasets but for large real-world graph topologies, building and utilizing the line graph is prohibitively expensive. Here, we present the framework to cluster the edges of a dynamic graph via the associated line graph that contains two major contributions. The first is a method to work with the line graph implicitly and the second is a distributed scale implementation of an agglomerative hierarchical graph clustering algorithm. We outline a novel hierarchical dynamic graph edge clustering approach that efficiently breaks massive relational datasets into small sets of edges containing events at various timescales. This is in stark contrast to traditional graph clustering algorithms that prioritize highly connected (clique-like) community structures. Our approach relies on constructing a sufficient subgraph of a weighted line graph and applying a hierarchical agglomerative clustering. This approach is related to scalable techniques from spatial clustering, nonlinear-dimension reduction, topological data analysis, and draws particular inspiration from HDBSCAN. As an edge clustering, this method yields an overlapping node clustering. Our algorithm is parallelizable and we demonstrate efficient clustering of a billion-scale, real-world dynamic graph into small edge sets that correlate in topology and time. The entire clustering process for a graph with tens of billions of edges takes just a few minutes of run time on 256 nodes of a distributed compute environment. We argue how the output of the edge clustering is useful for a multitude of data visualization and powerful machine learning tasks, both involving the original massive dynamic graph data and metadata associated with the nodes and edges. Finally, we describe how this approach can be extended to dynamic hypergraphs and dynamic graphs/hypergraphs with unstructured data living on vertices and edges.

Data Analysis↗

A Two-Step Time-Series Data Clustering Method for Building-Level Load Profile: Preprint

Residential and commercial buildings have huge potential to contribute value to improve grid resilience by participating grid services. To reveal the significant value, it is critical to estimate the grid service capability from these buildings. Unlike the large-scale distributed energy resources such as wind and solar farms, those buildings need to participate grid services in aggregation, not by individual. Therefore, it is important to appropriately group buildings for aggregation. In this paper, we develop a load profile clustering method to classify the building-level load profiles for grid service capability estimation. In our two-step clustering approach, we first calculate the total load consumption for each building, clustering the load profiles based on energy consumption level. Then, we further cluster the load profiles in each energy cluster based on the load shape. The parameter selection for each clustering step is discussed. The proposed method is applied on actual building-level load profiles, and the results have proved the effectiveness of this method.

advanced metering infrastructure (AMI)↗

From Femtoseconds to Gigaseconds: The SolDeg Platform for the Performance Degradation Analysis of Silicon Heterojunction Solar Cells

Heterojunction Si solar cells exhibit notable performance degradation. Here, we modeled this degradation by electronic defects getting generated by thermal activation across energy barriers over time. To analyze the physics of this degradation, we developed the SolDeg platform to simulate the dynamics of electronic defect generation. First, femtosecond molecular dynamics simulations were performed to create a-Si/c-Si stacks, using the machine learning-based Gaussian approximation potential. Second, we created shocked clusters by a cluster blaster method. Third, the shocked clusters were analyzed to identify which of them supported electronic defects. Fourth, the distribution of energy barriers that control the generation of these electronic defects was determined. Fifth, an accelerated Monte Carlo method was developed to simulate the thermally activated time-dependent defect generation across the barriers. Our main conclusions are as follows. (1) The degradation of a-Si/c-Si heterojunction solar cells via defect generation is controlled by a broad distribution of energy barriers. (2) We developed the SolDeg platform to track the microscopic dynamics of defect generation across this wide barrier distribution and determined the time-dependent defect density N(t) from femtoseconds to gigaseconds, over 24 orders of magnitude in time. (3) We have shown that a stretched exponential analytical form can successfully describe the defect generation N(t) over at least 10 orders of magnitude in time. (4) We found that in relative terms, Voc degrades at a rate of 0.2%/year over the first year, slowing with advancing time. (5) We developed the time correspondence curve to calibrate and validate the accelerated testing of solar cells. We found a compellingly simple scaling relationship between accelerated and normal times t normal ∝ t accel T(accel)/T(normal) . (6) We also carried out experimental studies of defect generation in a-Si:H/c-Si stacks. We found a relatively high degradation rate at early times that slowed considerably at longer time scales.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Using Signal Clustering Similarity for Detecting CAN Masquerade Attacks

The computer code assumes that time series representing the physical signals of the vehicle have been extracted from the CAN bus. The main input of the computer code is the multivariate time series representation of the signals in the CAN bus. The computer code cluster these time series using agglomerative hierarchical clustering from benign and attack datasets. Based on this, it generates probability distributions from the similarity of the obtained clusters based in each scenario---benign and attack---using the CluSim method (https://github.com/Hoosier-Clusters/clusim). Finally, it compares how a new data collection compares with the previous distribution to provide and probability score for an intrusion.

Moriano, Pablo↗

Deep Generative Models that Solve PDEs: Distributed Computing for Training Large Data-Free Models

Recent progress in scientific machine learning (SciML) has opened up the possibility of training novel neural network architectures that solve complex partial differential equations (PDEs). Several (nearly data free) approaches have been recently reported that successfully solve PDEs, with examples including deep feed forward networks, generative networks, and deep encoder-decoder networks. However, practical adoption of these approaches is limited by the difficulty in training these models, especially to make predictions at large output resolutions (≥1024×1024). Here we report on a software framework for data parallel distributed deep learning that resolves the twin challenges of training these large SciML models - training in reasonable time as well as distributing the storage requirements. Our framework provides several out of the box functionality including (a) loss integrity independent of number of processes, (b) synchronized batch normalization, and (c) distributed higher-order optimization methods. We show excellent scalability of this framework on both cloud as well as HPC clusters, and report on the interplay between bandwidth, network topology and bare metal vs cloud. We deploy this approach to train generative models of sizes hitherto not possible, showing that neural PDE solvers can be viably trained for practical applications. We also demonstrate that distributed higher-order optimization methods are 2-3× faster than stochastic gradient-based methods and provide minimal convergence drift with higher batch-size.

PDEs↗

A Two-Step Time-Series Data Clustering Method for Building-Level Load Profile

Residential and commercial buildings have huge potential to contribute value to improve grid resilience by participating grid services. To reveal the significant value, it is critical to estimate the grid service capability from these buildings. Unlike the large-scale distributed energy resources such as wind and solar farms, those buildings need to participate grid services in aggregation, not by individual. Therefore, it is important to appropriately group buildings for aggregation. The load profiles in the same group will have similar characteristics at the same time step, so grid operators can send the grid service signal to the customer group with a higher chance to respond at that time step. In this paper, we develop a load profile clustering method to classify the building-level load profiles for grid service capability estimation and the results have proved the effectiveness of this method.

AMI↗

The Initial–Final Mass Relation for Hydrogen-deficient White Dwarfs

The initial–final mass relation represents the total mass lost by a star during the entirety of its evolution from the zero age main sequence to the white-dwarf cooling track. The semiempirical initial–final mass relation (IFMR) is largely based on observations of DA white dwarfs, the most common spectral type of white dwarf and the simplest atmosphere to model. We present a first derivation of the semiempirical IFMR for hydrogen-deficient (non-DA) white dwarfs in open star clusters. We identify a possible discrepancy between the DA and non-DA IFMRs, with non-DA white dwarfs ≈0.07 M {sub ⊙} less massive at a given initial mass. Such a discrepancy is unexpected based on theoretical models of non-DA formation and observations of field white dwarf mass distributions. If real, the discrepancy is likely due to enhanced mass loss during the final thermal pulse and renewed post-AGB evolution of the star. However, we are dubious that the mass discrepancy is physical and instead is due to the small sample size, to systematic issues in model atmospheres of non-DAs, and to the uncertain evolutionary history of Procyon B (spectral type DQZ). A significantly larger sample size is needed to test these assertions. In addition, we also present Monte Carlo models of the correlated errors for DA and non-DA white dwarfs in the initial–final mass plane. We find the uncertainties in initial–final mass determinations for individual white dwarfs can be significantly asymmetric, but the recovered functional form of the IFMR is grossly unaffected by the correlated errors.

79 ASTRONOMY AND ASTROPHYSICS↗

A Two-Step Time-Series Data Clustering Method for Building-Level Load Profile

Residential and commercial buildings have huge potential to contribute value to improve grid resilience by participating grid services. To reveal the significant value, it is critical to estimate the grid service capability from these buildings. Unlike the large-scale distributed energy resources such as wind and solar farms, those buildings need to participate grid services in aggregation, not by individual. Therefore, it is important to appropriately group buildings for aggregation. The load profiles in the same group will have similar characteristics at the same time step, so grid operators can send the grid service signal to the customer group with a higher chance to respond at that time step. In this paper, we develop a load profile clustering method to classify the building-level load profiles for grid service capability estimation. In our two-step clustering approach, we first calculate the total load consumption for each building, clustering the load profiles based on energy consumption level. Then, we further cluster the load profiles in each energy cluster based on the load shape. The parameter selection for each clustering step is discussed. The proposed method is applied on actual building-level load profiles, and the results have proved the effectiveness of this method.

advanced metering infrastructure (AMI)↗

MLEC-Sim: A Simulator for Evaluating Multi-Level Erasure Coding

We present MLEC-Sim, a sophisticated simulator for Multi-Level Erasure Coding (MLEC), developed in approximately 13 KLOC. The simulator is engineered to analyze the impact of various system configurations and erasure coding policies on system durability and network overhead. It supports a comprehensive range of parameters including disk capacity, disk I/O bandwidth, failure rates, network bandwidth, and system scale, accommodating various erasure coding approaches such as Single-Level Erasure Coding (SLEC), Multi-Level Erasure Coding (MLEC), and Local Reconstruction Codes (LRC). MLEC-Sim provides support for multiple chunk placement policies, including clustered parity and declustered parity, and encompasses a variety of repair methods like Repair-ALL, Repair-FCO, Repair-HYB, and Repair-MIN. It is capable of simulating disk failures through a variety of means, including distribution-based or trace-based mechanisms, and can handle complex multi-level (de)clustered placements and repair processes. A key feature of MLEC-Sim is its adoption of the splitting simulation method for evaluating system durabilities at extremely high levels, which are challenging to assess with traditional simulation approaches. This feature allows for a detailed evaluation of system resilience under a range of conditions, aiding in the selection of appropriate erasure coding solutions for enhancing system durability. MLEC-Sim contributes to the field of data storage and reliability by providing a tool for the detailed evaluation of the durability and efficiency of erasure coding configurations, intended for use by researchers and practitioners in the design and optimization of storage systems.

Wang, Meng↗

Generating realistic building electrical load profiles through the Generative Adversarial Network (GAN)

Building electrical load profiles can improve understanding of building energy efficiency, demand flexibility, and building-grid interactions. Current approaches to generating load profiles are time-consuming and not capable of reflecting the dynamic and stochastic behaviors of real buildings; some approaches also trigger data privacy concerns. In this study, we proposed a novel approach for generating realistic electrical load profiles of buildings through the Generative Adversarial Network (GAN), a machine learning technique that is capable of revealing an unknown probability distribution purely from data. The proposed approach has three main steps: (1) normalizing the daily 24-hour load profiles, (2) clustering the daily load profiles with the k-means algorithm, and (3) using GAN to generate daily load profiles for each cluster. The approach was tested with an open-source database – the Building Data Genome Project. We validated the proposed method by comparing the mean, standard deviation, and distribution of key parameters of the generated load profiles with those of the real ones. The KL divergence of the generated and real load profiles are within 0.3 for majority of parameters and clusters. Additionally, results showed the load profiles generated by GAN can capture not only the general trend but also the random variations of the actual electrical loads in buildings. We report the proposed GAN approach can be used to generate building electrical load profiles, verify other load profile generation models, detect changes to load profiles, and more importantly, anonymize smart meter data for sharing, to support research and applications of grid-interactive efficient buildings.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Joint Analysis of Small-scale Galaxy Clustering and Galaxy–Galaxy Lensing from BOSS Galaxies

We present a joint analysis of galaxy clustering and galaxy–galaxy lensing measurements from BOSS galaxies using a simulation-based emulation method combined with a halo occupation distribution model. Our emulators are constructed with the Aemulus ν simulations, a suite of wνCDM N-body simulations with massive neutrinos as independent particle species. We combine small-scale analysis of clustering from 0.1 to 60.2 h −1 Mpc and lensing from 1.7 to 60.2 h −1 Mpc to perform cosmological constraints. We split the BOSS galaxies into three redshift bins to measure their clustering and employ galaxies from Dark Energy Camera Legacy Survey and Hyper Suprime-Cam as source galaxies to measure lensing separately. We find that the addition of lensing significantly improves the constraining power on $S_8 = σ_8(Ω_m/0.3)^{0.5}$, with a weak improvement for fσ 8 . Our results of fσ 8 indicate tensions of around 1σ−4σ below the results of the cosmic microwave background observations of Planck. For S 8 , our results are also lower than Planck, and the tension can be mitigated when considering possible systematics in lensing measurement. As a by-product, our analysis prefers a nonzero neutrino mass but without strong significance, with the constraining power dominated by the clustering. Given the accuracy and precision of our model and the observational data, it is anticipated that larger and higher-quality spectroscopic data sets will improve the constraints on this fundamental property in the near future.

Gao, Wenhao [Shanghai Jiao Tong University (China)↗

Sub-10 nm Probing of Ferroelectricity in Heterogeneous Materials by Machine Learning Enabled Contact Kelvin Probe Force Microscopy

Reducing the dimensions of ferroelectric materials down to the nanoscale has strong implications on the ferroelectric polarization pattern and on the ability to switch the polarization. As the size of ferroelectric domains shrinks to the nanometer scale, the heterogeneity of the polarization pattern becomes increasingly pronounced, enabling a large variety of possible polar textures in nanocrystalline and nanocomposite materials. Critical to the understanding of fundamental physics of such materials and hence their applications in electronic nanodevices is the ability to investigate their ferroelectric polarization at the nanoscale in a nondestructive way. We show that contact Kelvin probe force microscopy (cKPFM) combined with a k-means response clustering algorithm enables to measure the ferroelectric response at a mapping resolution of 8 nm. In a BaTiO 3 thin film on silicon composed of tetragonal and hexagonal nanocrystals, we determine a nanoscale lateral distribution of discrete ferroelectric response clusters, fully consistent with the nanostructure determined by transmission electron microscopy. Moreover, we apply this data clustering method to the cKPFM responses measured at different temperatures, which allows us to follow the corresponding change in the polarization pattern as the Curie temperature is approached and across the phase transition. This work opens up perspectives for mapping complex ferroelectric polarization textures such as curled/swirled polar textures that can be stabilized in epitaxial heterostructures and more generally for mapping the polar domain distribution of any spatially highly heterogeneous ferroelectric materials.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

ADDGALS: Simulated Sky Catalogs for Wide Field Galaxy Surveys

Abstract We present a method for creating simulated galaxy catalogs with realistic galaxy luminosities, broadband colors, and projected clustering over large cosmic volumes. The technique, denoted Addgals (Adding Density Dependent GAlaxies to Lightcone Simulations), uses an empirical approach to place galaxies within lightcone outputs of cosmological simulations. It can be applied to significantly lower-resolution simulations than those required for commonly used methods such as halo occupation distributions, subhalo abundance matching, and semi-analytic models, while still accurately reproducing projected galaxy clustering statistics down to scales of r ∼ 100 h −1 kpc . We show that Addgals catalogs reproduce several statistical properties of the galaxy distribution as measured by the Sloan Digital Sky Survey (SDSS) main galaxy sample, including galaxy number densities, observed magnitude and color distributions, as well as luminosity- and color-dependent clustering. We also compare to cluster–galaxy cross correlations, where we find significant discrepancies with measurements from SDSS that are likely linked to artificial subhalo disruption in the simulations. Applications of this model to simulations of deep wide-area photometric surveys, including modeling weak-lensing statistics, photometric redshifts, and galaxy cluster finding, are presented in DeRose et al., and an application to a full cosmology analysis of Dark Energy Survey (DES) Year 3 like data is presented in DeRose et al. We plan to publicly release a 10,313 square degree catalog constructed using Addgals with magnitudes appropriate for several existing and planned surveys, including SDSS, DES, VISTA, Wide-field Infrared Survey Explorer, and Rubin Observatory’s Legacy Survey of Space and Time.

79 ASTRONOMY AND ASTROPHYSICS↗