Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “local clustering algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Thermodynamic non-ideality and disorder heterogeneity in actinide silicate solid solutions

Non-ideal thermodynamics of solid solutions can greatly impact materials degradation behavior. We have investigated an actinide silicate solid solution system (USiO 4 –ThSiO 4 ), demonstrating that thermodynamic non-ideality follows a distinctive, atomic-scale disordering process, which is usually considered as a random distribution. Neutron total scattering implemented by pair distribution function analysis confirmed a random distribution model for U and Th in first three coordination shells; however, a machine-learning algorithm suggested heterogeneous U and Th clusters at nanoscale (~2 nm). The local disorder and nanosized heterogeneous is an example of the non-ideality of mixing that has an electronic origin. Partial covalency from the U/Th 5f–O 2p hybridization promotes electron transfer during mixing and leads to local polyhedral distortions. The electronic origin accounts for the strong non-ideality in thermodynamic parameters that extends the stability field of the actinide silicates in nature and under typical nuclear waste repository conditions.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Benchmarking image processing techniques for porosity measurement in polymer additive manufacturing: Review and experimental analysis

An image processing workflow is proposed for porosity measurement in polymer additive manufacturing. Various techniques, including global and local thresholding, region growing, and K-means clustering, were applied to microscopic images of carbon fiber reinforced acrylonitrile butadiene styrene (CF-ABS) and benchmarked for their ability to accurately measure porosity. Global methods included Otsu, minimum error, iterative, and entropy-based thresholding, while local methods included Niblack, Bernsen, Sauvola, and Bradley-Roth algorithms. Artificial uneven illumination was introduced to test local adaptive thresholds. Results showed significant differences in porosity values across methods. Otsu, region growing, and K-means clustering excelled under uniform illumination, while Sauvola and Bradley-Roth performed better with uneven illumination. Comparison with X-ray computed tomography (XCT) revealed slightly lower porosity values (2.55 %) than optimized methods (2.73–2.79 %) due to XCT's lower resolution excluding smaller pores. While XCT offers finer pore detection, it limits sample volume and underestimates porosity due to spatial variation. Validation using artificial grayscale images with 5 % porosity confirmed that Otsu, Bradley-Roth, region growing, and Sauvola algorithms produced accurate results. Although tested on a single material system, these methods can be adapted to others with optimization. In conclusion, given XCT's high computational and time costs, this study highlights suitable image processing techniques as cost-effective alternatives for porosity analysis in polymer composites.

Additive manufacturing↗

Manifold Sampling for Optimizing Nonsmooth Nonconvex Compositions

Here we propose a manifold sampling algorithm for minimizing a nonsmooth composition $f= h\circ F$, where we assume $h$ is nonsmooth and may be inexpensively computed in closed form and $F$ is smooth but its Jacobian may not be available. We additionally assume that the composition $h\circ F$ defines a continuous selection. Manifold sampling algorithms can be classified as model-based derivative-free methods, in that models of $F$ are combined with particularly sampled information about $h$ to yield local models for use within a trust-region framework. We demonstrate that cluster points of the sequence of iterates generated by the manifold sampling algorithm are Clarke stationary. We consider the tractability of three particular subproblems generated by the manifold sampling algorithm and the extent to which inexact solutions to these subproblems may be tolerated. Numerical results demonstrate that manifold sampling as a derivative-free algorithm is competitive with state-of-the-art algorithms for nonsmooth optimization that utilize first-order information about $f$.

97 MATHEMATICS AND COMPUTING↗

Robust clustering of the local Milky Way stellar kinematic substructures with Gaia eDR3

Understanding local stellar kinematic substructures in the solar neighbourhood helps build a complete picture of the formation of the Milky Way, as well as an empirical phase space distribution of dark matter that would inform detection experiments. We apply the clustering algorithm HDBSCAN on the Gaia early third data release to identify a list of stable clusters in velocity space and action-angle space by taking into account the measurement uncertainties and studying the stability of the clustering results. We find 1405 (497) stars in 23 (6) robust clusters in velocity space (action-angle space) that are consistently not associated with noise. We discuss the kinematic properties of these structures and study whether many of the small clusters belong to a similar larger cluster based on their chemical abundances. They are attributed to the known structures: the Gaia Sausage-Enceladus, the Helmi Stream, and globular cluster NGC 3201 are found in both spaces, while NGC 104 and the thick disc (Sequoia) are identified in velocity space (action-angle space). Although we do not identify any new structures, we find that the HDBSCAN member selection of already known structures is unstable to input kinematics of the stars when resampled within their uncertainties. We therefore present the stable subset of local kinematic structures, which are consistently identified by the clustering algorithm, and emphasize the need to take into account error propagation during both the manual and automated identification of stellar structures, both for existing ones as well as future discoveries.

79 ASTRONOMY AND ASTROPHYSICS↗

VoroClust

SAND2025-11465O VoroClust, also known as Voronoi Clustering, is a fast, density-based unsupervised clustering algorithm applicable to high-resolution and high-dimensional data. It operates as quickly as distance-based clustering methods while effectively capturing complex regional geometries, matching the performance of current density-based methods. VoroClust employs a data-centered sphere cover to reduce computational demands while preserving data topology. It propagates clusters outward from local density peaks. Although supervised machine learning is powerful for applications like image classification and segmentation, it requires comprehensive, consistent datasets, which many applications lack. Unsupervised clustering algorithms analyze the structure of each dataset rather than relying on similarities with other examples, making them well-suited for practical applications with insufficient or inappropriate data for supervised learning. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Ebeida, Mohamed [Sandia National Lab. (SNL-CA), Li↗

X-ray nano-imaging of defects in thin film catalysts via cluster analysis

Functional properties of transition-metal oxides strongly depend on crystallographic defects; crystallographic lattice deviations can affect ionic diffusion and adsorbate binding energies. Scanning x-ray nanodiffraction enables imaging of local structural distortions across an extended spatial region of thin samples. Yet, localized lattice distortions remain challenging to detect and localize using nanodiffraction, due to their weak diffuse scattering. Here, in this study, we apply an unsupervised machine learning clustering algorithm to isolate the low-intensity diffuse scattering in as-grown and alkaline-treated thin epitaxially strained SrIrO 3 films. We pinpoint the defect locations, find additional strain variation in the morphology of electrochemically cycled SrIrO 3 , and interpret the defect type by analyzing the diffraction profile through clustering. Our findings demonstrate the use of a machine learning clustering algorithm for identifying and characterizing hard-to-find crystallographic defects in thin films of electrocatalysts and highlight the potential to study electrochemical reactions at defect sites in operando experiments.

42 ENGINEERING↗

I-GCN: A Graph Convolutional Network Accelerator with Runtime Locality Enhancement through Islandization

In this paper, we propose a novel hardware accelerator for GCN inference called I-GCN that significantly improves data locality and reduces unnecessary computation through a new online graph restructuring algorithm we refer to as islandization. The proposed algorithm finds clusters of nodes with strong internal but weak external connections. The islandization process yields two major benefits. First, by processing islands rather than individual nodes, there is better on-chip data reuse and fewer off-chip memory accesses. Second, there is less redundant computation as aggregation for common/shared neighbors in an island can be reused. The parallel search, identification, and leverage of graph islands are all handled purely in hardware at runtime working in an incremental pipelined manner. This is done without any preprocessing of the graph data or adjustment of the GCN model structure.

Geng, Tong↗

Local Pair Natural Orbital-Based Coupled-Cluster Theory through Full Quadruples (DLPNO–CCSDTQ)

In this work, we implement a local pair natural orbitalbased coupled-cluster method through the full treatment of quadruple excitations (CCSDTQ). The domain-based local pair natural orbital (DLPNO) approach, which has successfully been applied to lower levels of coupled-cluster theory, is utilized in our algorithm, and thus our algorithm is called DLPNO-CCSDTQ. For simplicity in the working equations and in the implementation, we t 1 -dress the twoelectron integrals as well as Fock matrix elements. Our method can recover CCSDTQ-CCSDT and CCSDTQ-CCSDT(Q) energy differences on the order of 0.01−0.05 kcal mol −1 , even at a loose quadruples natural orbital (QNO) occupation number cutoff of 3.33 × 10 −6 . To highlight the capabilities of our code and its potential future applications, we showcase computations that would be intractable with canonical CCSDTQ, such as the benzene dimer, (H 2 O) 17 , and adamantane. With sufficient computing resources, computations up to 15 heavy atoms (40 atoms overall) may be feasible for fully bonded 3D systems.

Cluster chemistry↗

Linear-scaling quadruple excitations in local pair natural orbital coupled-cluster theory

Here, we present a fast, asymptotically linear-scaling implementation of the perturbative quadruples energy correction in coupled-cluster theory using local natural orbitals. Our work follows the domain-based local pair natural orbital (DLPNO) approach previously applied to lower levels of excitations in coupled-cluster theory. Our DLPNO-CCSDT(Q) algorithm uses converged doubles and triples amplitudes from a preceding DLPNO-CCSDT computation to compute the quadruples amplitude and energy in the quadruples natural orbital (QNO) basis. We demonstrate the compactness of the QNO space, showing that more than 95% of the (Q) correction can be recovered using relatively loose natural orbital cutoffs, compared to the tighter cutoffs used in pair and triples natural orbitals at lower levels of coupled-cluster theory. We also highlight the accuracy of our algorithm in the computation of relative energies, which yields deviations of sub-kJ mol −1 in relative energy compared to the canonical CCSDT(Q). Timings are conducted on a series of growing linear alkanes (up to 10 carbons and 608 basis functions) and water clusters (up to 49 water molecules and 2842 basis functions) to establish the asymptotic linear-scaling of our DLPNO-(Q) algorithm.

Auxiliary functions↗

Q-BEEP: Quantum Bayesian Error Mitigation Employing Poisson Modeling over the Hamming Spectrum

Quantum computing technology has grown rapidly in recent years, with new technologies being explored, error rates being reduced, and quantum processor’s qubit capacity growing. However, near-term quantum algorithms are still unable to be induced without compounding consequential levels of noise, leading to non-trivial erroneous results. Quantum Error Correction (in-situ error mitigation) and Quantum Error Mitigation (post-induction error mitigation) are promising fields of research within the quantum algorithm scene, aiming to alleviate quantum errors, increasing the overall fidelity and hence the overall quality of circuit induction. Earlier this year, a pioneering work, namely HAMMER, published in ASPLOS-22 demonstrated the existence of a latent structure regarding post-circuit induction errors when mapping to the Hamming spectrum. However, they intuitively assumed that errors occur in local clusters, and that at higher average Hamming distances this structure falls away. In this work, we show that such a correlation structure is not only local but extends certain non-local clustering patterns which can be precisely described by a Poisson distribution model taking the input circuit, the device run time status (i.e., calibration statistics) and qubit topology into consideration. Using this quantum error characterizing model, we developed an iterative algorithm over the generated Bayesian network state-graph for post-induction error mitigation. Thanks to more precise modeling of the error distribution latent structure and the new iterative method, our Q-Beep approach provides state of the art performance and can boost circuit execution fidelity by up to 234.6% on Bernstein-Vazirani circuits and on average 71.0% on QAOA solution quality, using 16 practical IBMQ quantum processors. For other benchmarks such as those in QASMBench, the fidelity improvement is up to 17.8%. Q-Beep is a light-weight post-processing technique that can be performed offline and remotely, making it a useful tool for quantum vendors to integrate and provide more reliable circuit induction results.

Stein, Samuel A.↗

TEMImageNet training library and AtomSegNet deep-learning models for high-precision atom segmentation, localization, denoising, and deblurring of atomic-resolution images

Abstract Atom segmentation and localization, noise reduction and deblurring of atomic-resolution scanning transmission electron microscopy (STEM) images with high precision and robustness is a challenging task. Although several conventional algorithms, such has thresholding, edge detection and clustering, can achieve reasonable performance in some predefined sceneries, they tend to fail when interferences from the background are strong and unpredictable. Particularly, for atomic-resolution STEM images, so far there is no well-established algorithm that is robust enough to segment or detect all atomic columns when there is large thickness variation in a recorded image. Herein, we report the development of a training library and a deep learning method that can perform robust and precise atom segmentation, localization, denoising, and super-resolution processing of experimental images. Despite using simulated images as training datasets, the deep-learning model can self-adapt to experimental STEM images and shows outstanding performance in atom detection and localization in challenging contrast conditions and the precision consistently outperforms the state-of-the-art two-dimensional Gaussian fit method. Taking a step further, we have deployed our deep-learning models to a desktop app with a graphical user interface and the app is free and open-source. We have also built a TEM ImageNet project website for easy browsing and downloading of the training data.

25 ENERGY STORAGE↗

CONUS-wide Projected Flood Frequency and Uncertainty Estimates, Version 1.0

This dataset presents a large-ensemble of CONUS-wide projected flood frequency and uncertainty estimates across ~2.7 million NHDPlusV2 river reaches over the CONUS. The framework producing this dataset leverages a multi-model, uncertainty-aware modeling framework that allows evaluating shifts in flood frequences at the stream reach level across the CONUS. CONUS-wide ensemble streamflow projections generated from hydrologic simulations driven by downscaled and bias-corrected Coupled Model Intercomparison Project Phase 6 (CMIP6) outputs are used to derive these flood frequency and uncertainty estimates over the period 1980 - 2099. A spatially consistent regional L-moment algorithm is applied across clusters defined by the US Hydrologic Unit Code Subregions (HUC4s and HUC8s) and NHDPlusV2 stream orders to estimate flood frequencies. The dataset also includes at-site based flood estimates that allow for the comparison between local and regional approach-based estimates, assess projected changes, and characterize their uncertainties. For more reliable estimation of rare flood frequencies such as 500 and 1000-year return periods, super-ensemble based estimates are also included in the dataset. This dataset is derived to support the "Impact-Informed Dam Safety Risk Assessment for Securing Hydropower Assests" project for the US Department of Energy (DOE) Hydropower and Hydrokinetic Office (H2O). For further details, refer to Kao et al. (2022), Ghimire et al. (2023), Ghimire et al. (2025), and Hosking and Wallis (1997).

Ghimire, Ganesh [ORNL] (ORCID:0000000242843941)↗

S-PLUS DR1 galaxy clusters and groups catalogue using PzWav

ABSTRACT We present a catalogue of 4499 groups and clusters of galaxies from the first data release of the multi-filter (5 broad, 7 narrow) Southern Photometric Local Universe Survey (S-PLUS). These groups and clusters are distributed over 273 deg2 in the Stripe 82 region. They are found using the PzWav algorithm, which identifies peaks in galaxy density maps that have been smoothed by a cluster scale difference-of-Gaussians kernel to isolate clusters and groups. Using a simulation-based mock catalogue, we estimate the purity and completeness of cluster detections: at S/N > 3.3, we define a catalogue that is 80 per cent pure and complete in the redshift range 0.1 < z < 0.4, for clusters with M200 > 1014 M⊙. We also assessed the accuracy of the catalogue in terms of central positions and redshifts, finding scatter of σR = 12 kpc and σz = 8.8 × 10−3, respectively. Moreover, less than 1 per cent of the sample suffers from fragmentation or overmerging. The S-PLUS cluster catalogue recovers ∼80 per cent of all known X-ray and Sunyaev-Zel’dovich selected clusters in this field. This fraction is very close to the estimated completeness, thus validating the mock data analysis and paving an efficient way to find new groups and clusters of galaxies using data from the ongoing S-PLUS project. When complete, S-PLUS will have surveyed 9300 deg2 of the sky, representing the widest uninterrupted areas with narrow-through-broad multi-band photometry for cluster follow-up studies.

Astronomy & Astrophysics↗

Adaptive Hierarchical Cyber Attack Detection and Localization in Active Distribution Systems

Development of a cyber security strategy for the active distribution systems is challenging due to the inclusion of distributed renewable energy generations. Here this paper proposes an adaptive hierarchical cyber attack detection and localization framework for distributed active distribution systems via analyzing electrical waveforms. Cyber attack detection is based on a sequential deep learning model, via which even minor cyber attacks can be identified. The two-stage cyber attack localization algorithm first estimates the cyber attack sub-region, and then localize the specified cyber attack within the estimated subregion. We propose a modified spectral clustering-based network partitioning method for the hierarchical cyber attack ‘coarse’ localization. Next, to further narrow down the cyber attack location, a normalized impact score based on waveform statistical metrics is proposed to obtain a ‘fine’ cyber attack location by characterizing different waveform properties. Finally, compared with classical and state-of-art methods, a comprehensive quantitative evaluation with two case studies shows promising estimation results of the proposed framework.

42 ENGINEERING↗

Global Optimization of Chemical Cluster Structures: Methods, Applications, and Challenges

Chemical clusters are relevant to many applications in catalysis, separations, materials, and energy sciences. Experimentally, the structure of clusters is difficult to determine, but it is very important in understanding their chemistry and properties. Computational methods can be used to examine cluster structure, however finding the most stable structure is not simple, particularly as the cluster size increases. Global optimization techniques have long been used to tackle the problem of the most stable structure, but such approaches would have to look for a global minimum, while sampling local minima over the whole potential energy surface as well. In this review, the state-of-the-art theory of global optimization theory is summarized. First, the definition, significance, relation to experiments, and a brief history of global optimization is presented. We then discuss, in more detail, three versatile global optimization methods: the basin hopping, the artificial bee colony algorithm, and the genetic algorithm. We close with some representative application examples of global optimization of clusters since 2016 and the challenges, open questions and opportunities in this field.

Global optimization, Chemical clusters, Artificial↗

Theory of Trotter Error with Commutator Scaling

The Lie-Trotter formula, together with its higher-order generalizations, provides a simple approach to decomposing the exponential of a sum of operators. Despite significant effort, the error scaling of such product formulas remains poorly understood. We develop a theory of Trotter error that overcomes the limitations of truncating the Baker-Campbell-Hausdorff expansion. Our analysis directly exploits the commutativity of operator summands, producing tighter error bounds for both real- and imaginary-time evolutions. Whereas previous work achieves similar goals for systems with geometric locality or Lie-algebraic structure, our approach holds in general. We give a host of improved algorithms for digital quantum simulation and quantum Monte Carlo methods, nearly matching or even outperforming the best previous results. Our applications include: (i) a simulation of second-quantized plane-wave electronic structure, nearly matching the interaction-picture algorithm of Low and Wiebe; (ii) a simulation of $k$-local Hamiltonians almost with induced one-norm scaling, faster than the qubitization algorithm of Low and Chuang; (iii) a simulation of rapidly decaying power-law interactions, outperforming the Lieb-Robinson-based approach of Tran et al.; (iv) a hybrid simulation of clustered Hamiltonians, dramatically improving the result of Peng, Harrow, Ozols, and Wu; and (v) quantum Monte Carlo simulations of the transverse field Ising model and quantum ferromagnets, tightening previous analyses of Bravyi and Gosset. We obtain further speedups using the fact that product formulas can preserve the locality of the simulated system. Specifically, we show that local observables can be simulated with complexity independent of the system size for power-law interacting systems, which implies a Lieb-Robinson bound nearly matching a recent result of Tran et al. Our analysis reproduces known tight bounds for first- and second-order formulas. We further investigate the tightness of our bounds for higher-order formulas. For quantum simulation of a one-dimensional Heisenberg model with an even-odd ordering of terms, our result overestimates the complexity by only a factor of $5$. Our bound is also close to tight for power-law interactions and other orderings of terms. This suggests that our theory can accurately characterize Trotter error in terms of both the asymptotic scaling and the constant prefactor.

quantum computing, numerical analysis↗

A group finder algorithm optimised for the study of local galaxy environments

Context. The majority of galaxy group catalogues available in the literature use the popular friends-of-friends algorithm which links galaxies using a linking length. One potential drawback to this approach is that clusters of points can be linked with thin bridges which may not be desirable. In order to study galaxy groups, it is important to obtain realistic group structures. Aim. Here, in this study, we present a new simple group finder algorithm, TD-ENCLOSER, that finds the group that encloses a target galaxy of interest. Methods. TD-ENCLOSER is based on the kernel density estimation method which treats each galaxy, represented by a zero-dimensional particle, as a two-dimensional circular Gaussian. The algorithm assigns galaxies to peaks in the density field in order of density in descending order (‘top down’) so that galaxy groups ‘grow’ around the density peaks. Outliers in under-dense regions are prevented from joining groups by a specified hard threshold, while outliers at the group edges are clipped below a soft (blurred) interior density level. Results. The group assignments are largely insensitive to all free parameter variations apart from the hard density threshold and the kernel standard deviation, although this is a known feature of density-based group finder algorithms and it operates with a computing speed that increases linearly with the size of the input sample. In preparation for a companion paper, we also present a simple algorithm to select unique representative groups when duplicates occur. Conclusions. TD-ENCLOSER is tested on a mock galaxy catalogue using a smoothing scale of 0.3 Mpc and is found to be able to recover the input group distribution with sufficient accuracy to be applied to observed galaxy distributions.

79 ASTRONOMY AND ASTROPHYSICS↗

StarHorse results for spectroscopic surveys and Gaia DR3: Chrono-chemical populations in the solar vicinity, the genuine thick disk, and young alpha-rich stars

The Gaia mission has provided an invaluable wealth of astrometric data for more than a billion stars in our Galaxy. The synergy between Gaia astrometry, photometry, and spectroscopic surveys gives us comprehensive information about the Milky Way. Using the Bayesian isochrone-fitting code StarHorse, we derive distances and extinctions for more than 10 million unique stars listed in both Gaia Data Release 3 and public spectroscopic surveys: 557 559 in GALAH+ DR3, 4 531 028 in LAMOST DR7 LRS, 347 535 in LAMOST DR7 MRS, 562 424 in APOGEE DR17, 471 490 in RAVE DR6, 249 991 in SDSS DR12 (optical spectra from BOSS and SEGUE), 67 562 in the Gaia-ESO DR5 survey, and 4 211 087 in the Gaia RVS part of the Gaia DR3 release. StarHorse can increase the precision of distance and extinction measurements where Gaia parallaxes alone would be uncertain. We used StarHorse for the first time to derive stellar ages for main-sequence turnoff and subgiant branch stars, around 2.5 million stars, with age uncertainties typically around 30%; the uncertainties drop to 15% for subgiant-branch-only stars, depending on the resolution of the survey. With the derived ages in hand, we investigated the chemical-age relations. In particular, the α and neutron-capture element ratios versus age in the solar neighbourhood show trends similar to previous works, validating our ages. We used the chemical abundances from local subgiant samples of GALAH DR3, APOGEE DR17, and LAMOST MRS DR7 to map groups with similar chemical compositions and StarHorse ages, using the dimensionality reduction technique t-SNE and the clustering algorithm HDBSCAN. We identify three distinct groups in all three samples, confirmed by their kinematic properties: the genuine chemical thick disk, the thin disk, and a considerable number of young alpha-rich stars (427) that are also a part of the delivered catalogues. We confirm that the genuine thick disk’s kinematics and age properties are radically different from those of the thin disk and compatible with high-redshift (z ≈ 2) star-forming disks with high dispersion velocities. We also find a few extra chemical populations in GALAH DR3 thanks to the availability of neutron-capture element information.

79 ASTRONOMY AND ASTROPHYSICS↗