Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “distributed clustering methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Analytic marginalization of N(z) uncertainties in tomographic galaxy surveys

In this paper, we present a new method to marginalize over uncertainties in redshift distributions, N(z), within tomographic cosmological analyses applicable to current and upcoming photometric galaxy surveys. We allow for arbitrary deviations from the best-guess N(z) governed by a general covariance matrix describing the uncertainty in our knowledge of redshift distributions. In principle, this is marginalization over hundreds or thousands of new parameters describing potential deviations as a function of redshift and tomographic bin. However, by linearly expanding the theory predictions around a fiducial model, this marginalization can be performed analytically, resulting in a modified data covariance matrix that effectively downweights the modes of the data vector that are more sensitive to redshift distribution variations. We showcase this method by applying it to the galaxy clustering measurements from the Hyper Suprime-Cam first data release. We illustrate how to marginalize over sample variance of the calibration sample and a large general systematic uncertainty in photometric estimation methods, and explore the impact of priors imposing smoothness in the redshift distributions.

79 ASTRONOMY AND ASTROPHYSICS↗

Structured background grids for generation of unstructured grids by advancing front method

A new method of background grid construction is introduced for generation of unstructured tetrahedral grids using the advancing-front technique. Unlike the conventional triangular/tetrahedral background grids which are difficult to construct and usually inadequate in performance, the new method exploits the simplicity of uniform Cartesian meshes and provides grids of better quality. The approach is analogous to solving a steady-state heat conduction problem with discrete heat sources. The spacing parameters of grid points are distributed over the nodes of a Cartesian background grid by interpolating from a few prescribed sources and solving a Poisson equation. To increase the control over the grid point distribution, a directional clustering approach is used. The new method is convenient to use and provides better grid quality and flexibility. Sample results are presented to demonstrate the power of the method.

Pirzadeh, Shahyar↗

Measurement of the photometric baryon acoustic oscillations with self-calibrated redshift distribution

ABSTRACT We use a galaxy sample derived from the Dark Energy Camera Legacy Survey Data Release 9 to measure the baryonic acoustic oscillations (BAO). The magnitude-limited sample consists of 10.6 million galaxies in an area of 4974 deg2 over the redshift range of [0.6, 1]. A key novelty of this work is that the true redshift distribution of the photo-z sample is derived from the self-calibration method, which determines the true redshift distribution using the clustering information of the photometric data alone. Through the angular correlation function in four tomographic bins, we constrain the BAO scale dilation parameter α to be 1.025 ± 0.033, consistent with the fiducial Planck cosmology. Alternatively, the ratio between the comoving angular diameter distance and the sound horizon, DM/rs, is constrained to be 18.94 ± 0.61 at the effective redshift of 0.749. We corroborate our results with the true redshift distribution obtained from a weighted spectroscopic sample, finding very good agreement. We have conducted a series of tests to demonstrate the robustness of the measurement. Our work demonstrates that the self-calibration method can effectively constrain the true redshift distribution in cosmological applications, especially in the context of photometric BAO measurement.

Astronomy & Astrophysics↗

Core Mass Estimates in Strong Lensing Galaxy Clusters: A Comparison between Masses Obtained from Detailed Lens Models, Single-halo Lens Models, and Einstein Radii

The core mass of galaxy clusters is both an important anchor of the radial mass distribution profile and probe of structure formation. With thousands of strong lensing galaxy clusters being discovered by current and upcoming surveys, timely, efficient, and accurate core mass estimates are needed. Here, we assess the results of two efficient methods to estimate the core mass of strong lensing clusters: the mass enclosed by the Einstein radius (M(<θ E ) where θ E is approximated from arc positions; Remolina González et al. 2020), and single-halo lens model (M SHM ; Remolina González et al. 2021), against measurements from publicly available detailed lens models (M DLM ) of the same clusters. We use data from the Sloan Giant Arc Survey, the Reionization Lensing Cluster Survey, the Hubble Frontier Fields, and the Cluster Lensing and Supernova Survey with Hubble. We find a scatter of 18.2% (8.2%) with a bias of -7.1% (1.0%) between M corr (<θ arcs ) (M SHM ) and M DLM . Last, we compare the statistical uncertainties measured in this work to those from simulations. This work demonstrates the successful application of these methods to observational data. As the effort to efficiently model the mass distribution of strong lensing galaxy clusters continues, we need fast, reliable methods to advance the field.

79 ASTRONOMY AND ASTROPHYSICS↗

High-accuracy method for modeling nucleation and growth of particles

State-of-the-art numerical models describing the kinetics of aerosol particle nucleation and growth from a cooling vapor primarily use a nodal method, in which particles that are smaller than the critical size are omitted from consideration because they are thermodynamically unfavorable. This omission is based on the assumption that most newly formed particles are above the critical size, so that subcritical-size particles are not important to take into account. Due to the nature of the nodal method, it suffers from numerical diffusion, which can cause an artificial broadening of the cluster size distribution leading to a significant overestimation of the number of large-size particles. To address these issues, we propose a more accurate numerical method that explicitly models particles of all sizes, and uses a special numerical scheme that substantially reduces the numerical diffusion and provides high solution accuracy and numerical stability. We extensively compare this novel method to the commonly used nodal solver of the general dynamic equation (GDE) for particle growth and demonstrate that it offers GDE solutions with higher accuracy with low numerical diffusion. Incorporating small subcritical clusters into the solution is crucial for: 1) more precise determination of the entire particle size distribution function and 2) wider applicability of the model to experimental studies with non-monotonic temperature variations leading to particle evaporation. The computational code implementing this numerical method in Python is available upon request.

42 ENGINEERING↗

Exploring the Landscape of Distributed Graph Clustering on Leadership Supercomputers

The rapid growth of large-scale datasets in fields like biology and social networks has driven the need for advanced graph analytics techniques. Community detection, a fundamental task in graph analytics, identifies closely connected groups of nodes within a network, providing valuable insights across various disciplines. This study focuses on two classic community detection methods, the Louvain algorithm and Markov Clustering (MCL), and evaluates the performance of two prominent distributed community detection algorithms: HiPDPL-GPU, our prior implementation, and HipMCL. We conduct experiments on GPU-accelerated heterogeneous HPC systems, Summit and Frontier, to assess their performance under varying conditions. Our objective is to identify the strengths and weaknesses of these algorithms in terms of scalability, and quality of solutions. We evaluate these algorithms on a diverse set of 70+ networks spanning 13 domains, with sizes ranging up to 4.2 billion edges. Our results demonstrate that HiPDPL-GPU consistently outperforms HipMCL, especially for large-scale networks. HiPDPL-GPU achieves significantly faster runtimes (47x to 1439x), higher modularity scores, and improved scalability. These findings highlight HiPDPL-GPU as a promising solution for efficient and effective large-scale graph analytics in diverse application domains, and provide insights into the feasibility of using MCL-based approaches for certain application domains.

Community detection, graph algorithms↗

LBNL CRADA (FP00009949) with the American Public Power Association: Electricity Reliability Metrics, Analysis, and Planning (Final Technical Report)

LBNL and APPA (the team) jointly examined the extent to which differences in distribution feeder characteristics are correlated with differences in their reliability performance when exposed to three different types of natural hazards (wildlife, weather, and vegetation). The team employed data-driven approaches to quantify the relationships between various measures of feeder reliability and a suite of feeder characteristics individually and jointly via a statistically-based clustering method.

24 POWER TRANSMISSION AND DISTRIBUTION↗

The use of unsupervised clustering as a classifier for LACIE MSS data

The author has identified the following significant results. This classification method appears to give accurate field center results and to give practical, statistically consistent and accurate estimates of crop proportions. The accuracy of this method is attributable to certain qualities of the particular clustering algorithm. These qualities are freedom from assumptions about Gaussian data, and the continual updating of distribution estimates, including updating the number of modes. This method is relatively tolerant of errors in the determination of crop type, as crop identity is used only for identifying clusters, and not for computing signatures.

Pentland, A. P.↗

Performance and Application of Parallel OVERFLOW Codes on Distributed and Shared Memory Platforms

The presentation discusses recent studies on the performance of the two parallel versions of the aerodynamics CFD code, OVERFLOW_MPI and _MLP. Developed at NASA Ames, the serial version, OVERFLOW, is a multidimensional Navier-Stokes flow solver based on overset (Chimera) grid technology. The code has recently been parallelized in two ways. One is based on the explicit message-passing interface (MPI) across processors and uses the _MPI communication package. This approach is primarily suited for distributed memory systems and workstation clusters. The second, termed the multi-level parallel (MLP) method, is simple and uses shared memory for all communications. The _MLP code is suitable on distributed-shared memory systems. For both methods, the message passing takes place across the processors or processes at the advancement of each time step. This procedure is, in effect, the Chimera boundary conditions update, which is done in an explicit "Jacobi" style. In contrast, the update in the serial code is done in more of the "Gauss-Sidel" fashion. The programming efforts for the _MPI code is more complicated than for the _MLP code; the former requires modification of the outer and some inner shells of the serial code, whereas the latter focuses only on the outer shell of the code. The _MPI version offers a great deal of flexibility in distributing grid zones across a specified number of processors in order to achieve load balancing. The approach is capable of partitioning zones across multiple processors or sending each zone and/or cluster of several zones into a single processor. The message passing across the processors consists of Chimera boundary and/or an overlap of "halo" boundary points for each partitioned zone. The MLP version is a new coarse-grain parallel concept at the zonal and intra-zonal levels. A grouping strategy is used to distribute zones into several groups forming sub-processes which will run in parallel. The total volume of grid points in each group are approximately balanced. A proper number of threads are initially allocated to each group, and in subsequent iterations during the run-time, the number of threads are adjusted to achieve load balancing across the processes. Each process exploits the multitasking directives already established in Overflow.

Djomehri, M. Jahed↗

Substructure in the stellar halo near the Sun: II. Characterisation of independent structures

In an accompanying paper, we present a data-driven method for clustering in ‘integrals of motion’ space and apply it to a large sample of nearby halo stars with 6D phase-space information. The algorithm identified a large number of clusters, many of which could tentatively be merged into larger groups. The goal here is to establish the reality of the clusters and groups through a combined study of their stellar populations (average age, metallicity, and chemical and dynamical properties) to gain more insights into the accretion history of the Milky Way. To this end, we developed a procedure that quantifies the similarity of clusters based on the Kolmogorov–Smirnov test using their metallicity distribution functions, and an isochrone fitting method to determine their average age, which is also used to compare the distribution of stars in the colour–absolute magnitude diagram. Also taking into consideration how the clusters are distributed in integrals of motion space allows us to group clusters into substructures and to compare substructures with one another. We find that the 67 clusters identified by our algorithm can be merged into 12 extended substructures and 8 small clusters that remain as such. The large substructures include the previously known Gaia-Enceladus, Helmi streams, Sequoia, and Thamnos 1 and 2. We identify a few over-densities that can be associated with the hot thick disc and host a small metal-poor population. Especially notable is the largest (by number of member stars) substructure in our sample which, although peaking at the metallicity characteristic of the thick disc, has a very well populated metal-poor component, and dynamics intermediate between the hot thick disc and the halo. We also identify additional debris in the region occupied by Sequoia with clearly distinct kinematics, likely remnants of three different accretion events with progenitors of similar masses. Although only a small subset of the stars in our sample have chemical abundance information, we are able to identify different trends of [Mg/Fe] versus [Fe/H] for the various substructures, confirming our dissection of the nearby halo. We find that at least 20% of the halo near the Sun is associated to substructures. When comparing their global properties, we note that those substructures on retrograde orbits are not only more metal-poor on average but are also older. We provide a table summarising the properties of the substructures, as well as a membership list that can be used for follow-up chemical abundance studies for example.

79 ASTRONOMY AND ASTROPHYSICS↗

Precise relative magnitude measurement improves fracture characterization during hydraulic fracturing

SUMMARY Microseismic monitoring is an important technique to obtain detailed knowledge of in-situ fracture size and orientation during stimulation to maximize fluid flow throughout the rock volume and optimize production. Furthermore, considering that the frequency of earthquake magnitudes empirically follows a power law (i.e. Gutenberg–Richter), the accuracy of microseismic event magnitude distributions is potentially crucial for seismic risk management. In this study, we analyse microseismicity observed during four hydraulic fracture treatments of the legacy Cotton Valley experiment in 1997 at the Carthage gas field of East Texas, where fractures were activated at the base of the sand-shale Upper Cotton Valley formation. We perform waveform cross-correlation to detect similar event clusters, measure relative amplitude from aligned waveform pairs with a principal component analysis, then measure precise relative magnitudes. The new magnitudes significantly reduce the deviations between magnitude differences and relative amplitudes of event pairs. This subsequently reduces the magnitude differences between clusters located at different depths. Reduction in magnitude differences between clusters suggests that some attenuation-related biases could be effectively mitigated with relative magnitude measurements. The maximum likelihood method is applied to understand the magnitude frequency distributions and quantify the seismogenic index of the clusters. Statistical analyses with new magnitudes suggest that fractures that are more favourably oriented for shear failure have lower b-value and higher seismogenic index, suggesting higher potential for relatively larger earthquakes, rather than fractures subparallel to maximum horizontal principal stress orientation.

58 GEOSCIENCES↗

Constraining Galaxy-Halo connection using machine learning

We investigate the potential of machine learning (ML) methods to model small-scale galaxy clustering for constraining Halo Occupation Distribution (HOD) parameters. Our analysis reveals that while many ML algorithms report good statistical fits, they often yield likelihood contours that are significantly biased in both mean values and variances relative to the true model parameters. This highlights the importance of careful data processing and algorithm selection in ML applications for galaxy clustering, as even seemingly robust methods can lead to biased results if not applied correctly. ML tools offer a promising approach to exploring the HOD parameter space with significantly reduced computational costs compared to traditional brute-force methods if their robustness is established. Using our ANN-based pipeline, we successfully recreate some standard results from recent literature. Properly restricting the HOD parameter space, transforming the training data, and carefully selecting ML algorithms are essential for achieving unbiased and robust predictions. Among the methods tested, artificial neural networks (ANNs) outperform random forests (RF) and ridge regression in predicting clustering statistics, when the HOD prior space is appropriately restricted. We demonstrate these findings using the projected two-point correlation function (w p (r p )), angular multipoles of the correlation function (ξ ℓ (r)), and the void probability function (VPF) of Luminous Red Galaxies from Dark Energy Spectroscopic Instrument mocks. Our results show that while combining w p (r p ) and VPF improves parameter constraints, adding the multipoles ξ 0 , ξ 2 , and ξ 4 to w p (r p ) does not significantly improve the constraints.

cosmology↗

Scalable edge clustering of dynamic graphs via weighted line graphs

Timestamped relational datasets consisting of records (or connections) between pairs of entities are ubiquitous in network science. For applications like peer-to-peer communication, email, various social network interactions, and computer network security, it is useful to organize these records into groups based on how and when they are occurring. Weighted line graphs offer a natural way to model how records are related in such datasets but for large real-world graph topologies, building and utilizing the line graph is prohibitively expensive. Here, we present the framework to cluster the edges of a dynamic graph via the associated line graph that contains two major contributions. The first is a method to work with the line graph implicitly and the second is a distributed scale implementation of an agglomerative hierarchical graph clustering algorithm. We outline a novel hierarchical dynamic graph edge clustering approach that efficiently breaks massive relational datasets into small sets of edges containing events at various timescales. This is in stark contrast to traditional graph clustering algorithms that prioritize highly connected (clique-like) community structures. Our approach relies on constructing a sufficient subgraph of a weighted line graph and applying a hierarchical agglomerative clustering. This approach is related to scalable techniques from spatial clustering, nonlinear-dimension reduction, topological data analysis, and draws particular inspiration from HDBSCAN. As an edge clustering, this method yields an overlapping node clustering. Our algorithm is parallelizable and we demonstrate efficient clustering of a billion-scale, real-world dynamic graph into small edge sets that correlate in topology and time. The entire clustering process for a graph with tens of billions of edges takes just a few minutes of run time on 256 nodes of a distributed compute environment. We argue how the output of the edge clustering is useful for a multitude of data visualization and powerful machine learning tasks, both involving the original massive dynamic graph data and metadata associated with the nodes and edges. Finally, we describe how this approach can be extended to dynamic hypergraphs and dynamic graphs/hypergraphs with unstructured data living on vertices and edges.

Data Analysis↗

Quantitative characterization of spatial distribution of particles in materials: Application to materials processing

Most engineering materials contain second phase particles or fibers which serve to reinforce the matrix phase. The effect of reinforcements on material properties is usually analyzed in terms of the average volume fraction and spacing of reinforcements, quantities which are global microstructural characteristics. However, material properties can also depend on local microstructural characteristics; for example, on how uniformly the reinforcing phase is distributed in the material. The analysis method will then be applied to a materials processing problem to discover how processing parameters can be selected to maximize redistribution of the reinforcing phase during processing. Several mathematical analysis methods could be adapted to the problem of characterizing the distribution of particles in materials. A tessellation-based method was selected. In the first phase of the investigation, a software package was written to automate the analysis. Typical results are shown. The analysis technique allows the degree to which particles are clustered together, the size and spacing of particle clusters, and the particle density in clusters to be found. The analysis methods were applied to computer-generated distributions and to a few real particle-containing materials. Methods for analyzing a nonuniform particle distribution in a material can be applied to two broad classes of materials science problems: understanding how the resulting particle distribution affects properties. The analysis method described is applied to a materials processing problem: how to select extrusion conditions to maximize the redistribution of reinforcing particles that are initially nonuniformly distributed. In addition, the tessellation-based method to analyze star distributions in spiral galaxies was adapted, illustrating the diverse types of problems to which the analysis method can be applied.

Parse, J. B.↗

A Two-Step Time-Series Data Clustering Method for Building-Level Load Profile: Preprint

Residential and commercial buildings have huge potential to contribute value to improve grid resilience by participating grid services. To reveal the significant value, it is critical to estimate the grid service capability from these buildings. Unlike the large-scale distributed energy resources such as wind and solar farms, those buildings need to participate grid services in aggregation, not by individual. Therefore, it is important to appropriately group buildings for aggregation. In this paper, we develop a load profile clustering method to classify the building-level load profiles for grid service capability estimation. In our two-step clustering approach, we first calculate the total load consumption for each building, clustering the load profiles based on energy consumption level. Then, we further cluster the load profiles in each energy cluster based on the load shape. The parameter selection for each clustering step is discussed. The proposed method is applied on actual building-level load profiles, and the results have proved the effectiveness of this method.

advanced metering infrastructure (AMI)↗

From Femtoseconds to Gigaseconds: The SolDeg Platform for the Performance Degradation Analysis of Silicon Heterojunction Solar Cells

Heterojunction Si solar cells exhibit notable performance degradation. Here, we modeled this degradation by electronic defects getting generated by thermal activation across energy barriers over time. To analyze the physics of this degradation, we developed the SolDeg platform to simulate the dynamics of electronic defect generation. First, femtosecond molecular dynamics simulations were performed to create a-Si/c-Si stacks, using the machine learning-based Gaussian approximation potential. Second, we created shocked clusters by a cluster blaster method. Third, the shocked clusters were analyzed to identify which of them supported electronic defects. Fourth, the distribution of energy barriers that control the generation of these electronic defects was determined. Fifth, an accelerated Monte Carlo method was developed to simulate the thermally activated time-dependent defect generation across the barriers. Our main conclusions are as follows. (1) The degradation of a-Si/c-Si heterojunction solar cells via defect generation is controlled by a broad distribution of energy barriers. (2) We developed the SolDeg platform to track the microscopic dynamics of defect generation across this wide barrier distribution and determined the time-dependent defect density N(t) from femtoseconds to gigaseconds, over 24 orders of magnitude in time. (3) We have shown that a stretched exponential analytical form can successfully describe the defect generation N(t) over at least 10 orders of magnitude in time. (4) We found that in relative terms, Voc degrades at a rate of 0.2%/year over the first year, slowing with advancing time. (5) We developed the time correspondence curve to calibrate and validate the accelerated testing of solar cells. We found a compellingly simple scaling relationship between accelerated and normal times t normal ∝ t accel T(accel)/T(normal) . (6) We also carried out experimental studies of defect generation in a-Si:H/c-Si stacks. We found a relatively high degradation rate at early times that slowed considerably at longer time scales.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Using Signal Clustering Similarity for Detecting CAN Masquerade Attacks

The computer code assumes that time series representing the physical signals of the vehicle have been extracted from the CAN bus. The main input of the computer code is the multivariate time series representation of the signals in the CAN bus. The computer code cluster these time series using agglomerative hierarchical clustering from benign and attack datasets. Based on this, it generates probability distributions from the similarity of the obtained clusters based in each scenario---benign and attack---using the CluSim method (https://github.com/Hoosier-Clusters/clusim). Finally, it compares how a new data collection compares with the previous distribution to provide and probability score for an intrusion.

Moriano, Pablo↗

Convergence rate enhancement of navier-stokes codes on clustered grids

Our Sensitivity-Based Minimal Residual (SBMR) method which is based on our earlier Distributed Minimal Residual (DMR) method allows each component of the solution vector in a system of equations to have its own convergence speed. Our global SBMR method was found to consistently outperform the DMR method while requiring considerably less computer memory. Recently, we have developed and tested a new Line SBMR or LSBMR method and a Time-Step-Scaling (TSS) method that are even more robust and computationally efficient than our global SBMR method, especially on highly clustered computational grids in laminar and turbulent flow computations.

Choi, Kwang-Yoon↗