Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “local clustering algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Materials Learning Algorithms (MALA): Scalable machine learning for electronic structure calculations in large-scale atomistic simulations

We present the Materials Learning Algorithms (MALA) package, a scalable machine learning framework designed to accelerate density functional theory (DFT) calculations suitable for large-scale atomistic simulations. Using local descriptors of the atomic environment, MALA models efficiently predict key electronic observables, including local density of states, electronic density, density of states, and total energy. The package integrates data sampling, model training and scalable inference into a unified library, while ensuring compatibility with standard DFT and molecular dynamics codes. We demonstrate MALA's capabilities with examples including boron clusters, aluminum across its solid-liquid phase boundary, and predicting the electronic structure of a stacking fault in a large beryllium slab. Scaling analyses reveal MALA's computational efficiency and identify bottlenecks for future optimization. With its ability to model electronic structures at scales far beyond standard DFT, MALA is well suited for modeling complex material systems, making it a versatile tool for advanced materials research.

Density functional theory↗

Nearby stellar substructures in the Galactic halo from DESI Milky Way Survey Year 1 Data Release

We report five nearby ($d_{\mathrm{helio}} < 5$ kpc) stellar substructures in the Galactic halo from a subset of 138 661 stars in the Dark Energy Spectroscopic Instrument (DESI) Milky Way Survey Year 1 Data Release. With an unsupervised clustering algorithm, HDBSCAN*, these substructures are independently identified in Integrals of Motion ($E_{\rm tot}$, $L_{\rm z}$, $\log {J_r}$, $\log {J_z}$) space and Galactocentric cylindrical velocity space ($V_{R}$, $V_{\phi }$, $V_{z}$). We associate all identified clusters with known nearby substructures (Helmi streams, M18-Cand10/MMH-1, Sequoia, Antaeus, and ED-2) previously reported in various studies. With metallicities precisely measured by DESI, we confirm that the Helmi streams, M18-Cand10, and ED-2 are chemically distinct from local halo stars. We have characterized the chemodynamic properties of each dynamic group, including their metallicity dispersions, to associate them with their progenitor types (globular cluster or dwarf galaxy). Our approach for searching substructures with HDBSCAN* reliably detects real substructures in the Galactic halo, suggesting that applying the same method can lead to the discovery of new substructures in future DESI data. With more stars from future DESI data releases and improved astrometry from the upcoming Gaia Data Release 4, we will have a more detailed blueprint of the Galactic halo, offering a significant improvement in our understanding of the formation and evolutionary history of the Milky Way Galaxy.

dynamics↗

An X-ray image of the violent interstellar medium in 30 Doradus

A detailed analysis of the X-ray emission from the largest H II region complex in the Local Group, 30 Dor, is presented. Applying a new maximum entropy deconvolution algorithm to the Einstein Observatory data, reveals striking correlations among the X-ray, radio, and optical morphologies of the region, with X-ray-emitting bubbles filling cavities surrounded by H-alpha shells and coextensive diffuse X-ray and radio continuum emission from throughout the region. The total X-ray luminosity in the 0.16-3.5 keV band from an area within 160 pc of the central cluster R136 is about 2 x 10 to the 37th ergs/sec.

Wang, Q.↗

Towards Superior Software Portability with SHAD and HPX C++ Libraries

As hardware architectures and software stacks complexity grows, development productivity, performance and software portability, quickly evolve from desirable features to actual needs. SHAD, the Scalable High-performance Algorithms and Data-structures C++ library is designed to mitigate these issues: it provides general purpose building blocks as well as high-level custom utilities, and offers a shared-memory programming abstraction which facilitates the programming of complex systems, scaling up to High Performance Computing clusters. SHAD’s portability is achieved through an abstract runtime interface, which decouples the upper layers of the library and hides the low level details of the underlying architecture. This layer enables SHAD to interface with different runtime/threading systems, e.g. Intel TBB and Global Memory and Threading (GMT). However, current backends targeting distributed systems, rely on a centralized controller which may possibly limit scalability up to hundreds of nodes and creates a network hot spot due to all to one communication for synchronization, and possibly resulting in degraded performance at high process counts. In this research, we explore HPX, the C++ standard library for parallelism and concurrency, as an additional backend in support of the SHAD library, and present the methodologies in support of local and remote task executions in SHAD with respect to HPX. Finally, we evaluate the proposed system by comparing against existing backends of SHAD and analyzing their performance on C++ Standard Template Library algorithms.

Wu, Nanmiao↗

Geometric Interpretation of the Cluster Location Problem Part I: Theory

We present a new framing of the seismic location problem using principles drawn from differential geometry. Our interpretation relies upon the common assumption that travel times observed across a network are continuous, differentiable functions of source location. In consequence, travel‐time functions constitute a differentiable map between the source region and a Riemannian manifold. The manifold is said to be the image of the source region embedded in a generally high‐dimension travel‐time vector space. A cluster of events in the source region has an image of discrete points on the manifold, that, except in the simplest cases, cannot be viewed directly. However, it is possible to project the image of a cluster into a tangent space of the manifold for direct visualization. The projection operator can be computed directly from the data without a velocity model, but produces a distorted rendering of the cluster geometry. With a model we can predict the distortions and correct them to estimate cluster geometry. We develop these points with the simplest possible example, one for which direct visualization of the manifold is possible, using the example as an introduction to the relevant concepts from differential geometry in a familiar setting. The tangent space, a local linearization of the manifold, plays a key role. We develop a metric to estimate the limits of linearization, that is, to determine when the curvature of the manifold invalidates the linear assumption. We also examine the interplay of model error, inadequate network geometry, and pick error. We then generalize our results from the simple case to the general case of 3D source regions observed by general networks. Although we do suggest a new “project and correct” method for location, we do not develop it into a practical algorithm. In conclusion, our intention rather is to highlight new analytical methods grounded in differential geometry.

East Pacific Ocean Islands↗

Deterministic High-Fidelity Neutronics Simulation of Pebble Bed Reactors Using Pebble Tracking Transport

The pebble tracking transport (PTT) algorithm offers a high-fidelity deterministic approach for neutron transport for pebble bed reactors (PBRs). This approach requires the mesh for the active-core region to consist exclusively of tetrahedral elements, where each node in the pebble-packing region represents a pebble centroid. This paper investigates the application of PTT for full-scale PBRs, considering both the isothermal and the temperature-dependent core conditions. Macroscopic cross sections are generated using Serpent 2 full-core eigenvalue simulations where pebbles are grouped into disjoint subsets using machine learning. To minimize the need for individual cross-section sets for each pebble in the core, K-means clustering is used to group pebbles by temperature and neutronic environment parameters. Here, we compare the multiplication factor and power rate distributions between PTT simulations using the Griffin reactor physics software and reference solutions from Serpent 2. Our analysis shows that a full-core, high-fidelity PTT calculation produces accurate results with minimal local (pebblewise) errors. Additionally, timing results indicate that PTT simulations converge rapidly on modern supercomputing platforms.

Griffin↗

Hail Size Distribution Mapping

A 3-D weather radar visualization software program was developed and implemented as part of an experimental Launch Pad 39 Hail Monitor System. 3DRadPlot, a radar plotting program, is one of several software modules that form building blocks of the hail data processing and analysis system (the complete software processing system under development). The spatial and temporal mapping algorithms were originally developed through research at the University of Central Florida, funded by NASA s Tropical Rainfall Measurement Mission (TRMM), where the goal was to merge National Weather Service (NWS) Next-Generation Weather Radar (NEXRAD) volume reflectivity data with drop size distribution data acquired from a cluster of raindrop disdrometers. In this current work, we adapted these algorithms to process data from a cluster of hail disdrometers positioned around Launch Pads 39A or 39B, along with the corresponding NWS radar data. Radar data from all NWS NEXRAD sites is archived at the National Climatic Data Center (NCDC). That data can be readily accessed at . 3DRadPlot plots Level III reflectivity data at four scan elevations (this software is available at Open Channel Software, ). By using spatial and temporal interpolation/extrapolation based on hydrometeor fall dynamics, we can merge the hail disdrometer array data coupled with local Weather Surveillance Radar-1988, Doppler (WSR-88D) radial velocity and reflectivity data into a 4-D (3-D space and time) picture of hail size distributions. Hail flux maps can then be generated and used for damage prediction and assessment over specific surfaces corresponding to structures within the disdrometer array volume. Immediately following a hail storm, specific damage areas and degree of damage can be identified for inspection crews.

Source record↗

A generalizable machine learning-assisted fast Fourier transform algorithm to simulate the large strain phenomena in polycrystalline materials

Machine learning methods have shown initial promise in constitutive modeling for single crystals or homogenized polycrystals, delivering notable computational efficiency. However, existing machine learning-based constitutive models often lack generalizability, limiting their application across diverse boundary value problems. This study introduces a thermodynamics-informed artificial neural network model to accelerate rate-tangent crystal plasticity fast Fourier transform simulations for cross-scale deformation behaviors of polycrystals under complex loading. Our model integrates microstructural variability and local interactions effectively. To address local effects in each grain, we employ K-means clustering to group Gauss points within the microstructure into clusters assumed to be in similar mechanical states. This approach, based on self-clustering analysis, extends model scope from macroscopic stress response to the granular level, capturing mechanical responses and orientation evolution across grains. This reduces the number of nonlinear problems to solve, with cluster responses propagated throughout each group. The thermodynamics-based artificial neural network-extracted features are further processed using local material state clusters to account for history-dependent deformation and evolving microstructures. Additionally, representative volume element simulations with rate-tangent crystal plasticity fast Fourier transform provide reliable datasets for model training. The proposed model demonstrates high efficiency, accuracy, self-consistency, and enhanced generalizability in predicting strain–stress responses and orientation evolution at both individual grain and aggregate scales under complex loading conditions, such as biaxial tension and arbitrary loading scenarios.

36 MATERIALS SCIENCE↗

Extracting sea ice data from satellite SAR imagery

With the prospect operational satellite SAR's by the end of the decade, there is a clear need to develop automated algorithms for the extraction of geophysical data about sea ice from high resolution radar imagery. To this end, techniques were developed for distinguishing ice from open water and for resolving the details of deformation within areas 100 km square imaged by SEASAT SAR. The classification of ice and open water is based on the creation of two bands of lower resolution image data: local average brightness, and the local variance of brightness. In the space of these two variables, ice and open water are separated into two distinct clusters. The deformation is found on a 3.4 km mesh by local cross-correlations of the brightness. Comparison with manually determined deformation shows room for improvement in regions of high deformation by using smaller areas for cross-correlation. Th concentration and deformation data are used together to determine localized regions of the scene where open water is produced or lost.

Fily, M.↗

Primordial non-Gaussianity from the completed SDSS-IV extended Baryon Oscillation Spectroscopic Survey – I: Catalogue preparation and systematic mitigation

ABSTRACT We investigate the large-scale clustering of the final spectroscopic sample of quasars from the recently completed extended Baryon Oscillation Spectroscopic Survey (eBOSS). The sample contains 343 708 objects in the redshift range 0.8 < z < 2.2 and 72 667 objects with redshifts 2.2 < z < 3.5, covering an effective area of $4699\, {\rm deg}^{2}$. We develop a neural network-based approach to mitigate spurious fluctuations in the density field caused by spatial variations in the quality of the imaging data used to select targets for follow-up spectroscopy. Simulations are used with the same angular and radial distributions as the real data to estimate covariance matrices, perform error analyses, and assess residual systematic uncertainties. We measure the mean density contrast and cross-correlations of the eBOSS quasars against maps of potential sources of imaging systematics to address algorithm effectiveness, finding that the neural network-based approach outperforms standard linear regression. Stellar density is one of the most important sources of spurious fluctuations, and a new template constructed using data from the Gaia spacecraft provides the best match to the observed quasar clustering. The end-product from this work is a new value-added quasar catalogue with the improved weights to correct for non-linear imaging systematic effects, which will be made public. Our quasar catalogue is used to measure the local-type primordial non-Gaussianity in a companion paper.

79 ASTRONOMY AND ASTROPHYSICS↗

Augmenting Graph Convolution with Distance Preserving Embedding for Improved Learning

Graph convolution incorporates topological information of a graph into learning. Message passing corresponds to traversal of a local neighborhood in classical graph algorithms. We show that incorporating additional global structures, such as shortest paths, through distance preserving embedding can improve performance. Our approach, Gavotte, significantly improves the performance of a range of popular graph neu-ral networks such as GCN, GA T,Graph SAGE, and GCNII for transductive learning. Gavotte also improves the performance of graph neural networks for full-supervised tasks, albeit to a smaller degree. As high-quality embeddings are generated by Gavotte as a by-product, we leverage clustering algorithms on these embed dings to augment the training set and introduce Gavotte+. Our results of Gavotte+ on datasets with very few labels demonstrate the advantage of augmenting graph convolution with distance preserving embedding.

Cong, Guojing↗

A Massively Parallel Implementation of the CCSD(T) Method Using the Resolution-of-the-Identity Approximation and a Hybrid Distributed/Shared Memory Parallelization Model

In this work, a parallel algorithm is described for the coupled-cluster singles and doubles method augmented with a perturbative correction for triple excitations [CCSD(T)] using the resolution-of-the-identity (RI) approximation for two-electron repulsion integrals (ERIs). The algorithm bypasses the storage of four-center ERIs by adopting an integral-direct strategy. The CCSD amplitude equations are given in a compact quasi-linear form by factorizing them in terms of amplitude-dressed three-center intermediates. A hybrid MPI/OpenMP parallelization scheme is employed, which uses the OpenMP-based shared memory model for intranode parallelization and the MPI-based distributed memory model for internode parallelization. Parallel efficiency has been optimized for all terms in the CCSD amplitude equations. Two different algorithms have been implemented for the rate-limiting terms in the CCSD amplitude equations that entail and -scaling computational costs, where N O and N V denote the number of correlated occupied and virtual orbitals, respectively. One of the algorithms assembles the four-center ERIs requiring N V 4 and N O 2 N V 2 -scaling memory costs in a distributed manner on a number of MPI ranks, while the other algorithm completely bypasses the assembling of quartic memory-scaling ERIs and thus largely reduces the memory demand. It is demonstrated that the former memory-expensive algorithm is faster on a few hundred cores, while the latter memory-economic algorithm shows a better strong scaling in the limit of a few thousand cores. The program is shown to exhibit a near-linear scaling, in particular for the compute-intensive triples correction step, on up to 8000 cores. The performance of the program is demonstrated via calculations involving molecules with 24–51 atoms and up to 1624 atomic basis functions. As the first application, the complete basis set (CBS) limit for the interaction energy of the π-stacked uracil dimer from the S66 data set has been investigated. This work reports the first calculation of the interaction energy at the CCSD(T)/aug-cc-pVQZ level without local orbital approximation. The CBS limit for the CCSD correlation contribution to the interaction energy was found to be -8.01 kcal/mol, which agrees very well with the value -7.99 kcal/mol reported by Schmitz, Hättig, and Tew [ Phys. Chem. Chem. Phys. 2014 , 16 , 22167-22178]. The CBS limit for the total interaction energy was estimated to be -9.64 kcal/mol.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

GPU-acceleration of the ELPA2 distributed eigensolver for dense symmetric and hermitian eigenproblems

The solution of eigenproblems is often a key computational bottleneck that limits the tractable system size of numerical algorithms, among them electronic structure theory in chemistry and in condensed matter physics. Large eigenproblems can easily exceed the capacity of a single compute node, thus must be solved on distributed-memory parallel computers. We here present GPU-oriented optimizations of the ELPA two-stage tridiagonalization eigensolver (ELPA2). On top of cuBLAS-based GPU offloading, we add a CUDA kernel to speed up the back-transformation of eigenvectors, which can be the computationally most expensive part of the two-stage tridiagonalization algorithm. Furthermore, we benchmark the performance of this GPU-accelerated eigensolver on two hybrid CPU–GPU architectures, namely a compute cluster based on Intel Xeon Gold CPUs and NVIDIA Volta GPUs, and the Summit supercomputer based on IBM POWER9 CPUs and NVIDIA Volta GPUs. Consistent with previous benchmarks on CPU-only architectures, the GPU-accelerated two-stage solver exhibits a parallel performance superior to the one-stage counterpart. Finally, we demonstrate the performance of the GPU-accelerated eigensolver developed in this work for routine semi-local KS-DFT calculations comprising thousands of atoms.

97 MATHEMATICS AND COMPUTING↗

Particle Filter Based Inference Testing

The primary intent of PAR-FIT (Particle Filter based Inference Testing) is to provide hard inductive evidence that a machine learning model is capable and proven for an individual test input. By examining training data used to form the underlying model functional correlation, an estimate of the reliability that a model will make the correct prediction can be made. The Sequential Probability Ratio Test is used to derive a qualitative evaluation for reliability based on hypothesis testing. The PAR-FIT framework achieves this by implementing a particle filter and the sequential probability ratio test algorithms on the machine learning model training data to determine relevancy of new individual test samples to the training dataset. The kernel function evaluates the local proximity and density of training data used to derive a prediction outcome. Particles are used to probabilistically determine which training data to evaluate for proximity. For test samples that are within a close proximity to and surrounded by multiple training data points, the evaluated reliability of the prediction is high. For test samples that are anomalies not represented by the training dataset, in low density data clusters, or are far from existing data points, the evaluated reliability is low as insufficient training evidence exists to suggest the model is capable of making the correct prediction. Sequential Probability Ratio Test is further used to determine when a hypothesis on whether a signal can be rejected or accepted for use. The ratio test collects sequence information from the particle filter to test whether the signal is anomalous or normal via hypothesis testing of the underlying distributions.

Chen, Edward [Idaho National Laboratory (INL), Ida↗

Discovery and Spectroscopic Characterization of a Distant, Compact Milky Way Satellite in Gemini

We present the discovery of a compact Milky Way satellite in the constellation of Gemini. This system was discovered by cross-matching detections from two independent search algorithms applied to Blanco/DECam data from the third data release of the DECam Local Volume Exploration survey (DELVE DR3), and confirmed with deeper imaging from Gemini/GMOS-N. Based on these data, we determine that the system is an ultra-faint ($M_V = -2.1^{+0.4}_{-0.6}$), compact ($r_{1/2} = 8.6^{+1.4}_{-1.2}$ pc) system located at a heliocentric distance of $120^{+7}_{-6}$ kpc. These physical properties place the system in the regime of ambiguous, ultra-faint compact Milky Way halo satellites that cannot be confidently classified as dwarf galaxies or star clusters from morphology alone; we therefore name the system DELVE 8/Gemini I. From medium-resolution Keck/DEIMOS spectroscopy, we securely identify four members including two blue horizontal branch stars, confirming the system as a bound satellite moving at a mean radial velocity of $v_{\rm hel} = -82.7^{+3.7}_{-3.9} {\rm km\,s}^{-1}$. We also use these spectra to place an upper limit of $\rm [Fe/H] \lesssim -2.5$ on the metallicity of DELVE 8/Gemini I's brightest star, supporting the classification of the system as either an ancient star cluster or ultra-faint dwarf galaxy. The discovery of faint, distant systems similar to DELVE 8/Gemini I is expected to become more common with upcoming surveys.

Overdeck, K. [Chicago U., Astron. Astrophys. Ctr.;↗

Ultra-faint Milky Way Satellites Discovered in Carina, Phoenix, and Telescopium with DELVE Data Release 3

We report the discovery of three Milky Way satellite candidates: Carina IV, Phoenix III, and DELVE 7, in the third data release of the DECam Local Volume Exploration survey (DELVE). The candidate systems were identified by cross-matching results from two independent search algorithms. All three are extremely faint systems composed of old, metal-poor stellar populations (τ ≳ 10 Gyr, [Fe/H] ≲−1.4). Carina IV (M V = −2.8; r 1/2 = 40 pc) and Phoenix III (M V = −1.2; r 1/2 = 19 pc) have half-light radii that are consistent with the known population of dwarf galaxies, while DELVE 7 (M V = 1.2; r 1/2 = 2 pc) is very compact and seems more likely to be a star cluster, though its nature remains ambiguous without spectroscopic follow-up. The Gaia proper motions of stars in Carina IV ($M_{\star} = 2250^{+1180}_{-830} M_⊙$) indicate that it is unlikely to be associated with the LMC, while DECam CaHK photometry confirms that its member stars are metal poor. Phoenix III ($M_{\star} = 520^{+660}_{-290} M_⊙$) is the faintest known satellite in the extreme outer stellar halo (D GC > 100 kpc), while DELVE 7 ($M_{\star} = 60^{+120}_{-40} M_⊙$) is the faintest known satellite with D GC > 20 kpc.

Tan, Chin Yi [Univ. of Chicago, IL (United States)↗

Message Passing vs. Shared Address Space on a Cluster of SMPs

The convergence of scalable computer architectures using clusters of PCs (or PC-SMPs) with commodity networking has become an attractive platform for high end scientific computing. Currently, message-passing and shared address space (SAS) are the two leading programming paradigms for these systems. Message-passing has been standardized with MPI, and is the most common and mature programming approach. However message-passing code development can be extremely difficult, especially for irregular structured computations. SAS offers substantial ease of programming, but may suffer from performance limitations due to poor spatial locality, and high protocol overhead. In this paper, we compare the performance of and programming effort, required for six applications under both programming models on a 32 CPU PC-SMP cluster. Our application suite consists of codes that typically do not exhibit high efficiency under shared memory programming. due to their high communication to computation ratios and complex communication patterns. Results indicate that SAS can achieve about half the parallel efficiency of MPI for most of our applications: however, on certain classes of problems SAS performance is competitive with MPI. We also present new algorithms for improving the PC cluster performance of MPI collective operations.

Shan, Hongzhang↗

PDBspheres: a method for finding 3D similarities in local regions in proteins

Abstract We present a structure-based method for finding and evaluating structural similarities in protein regions relevant to ligand binding. PDBspheres comprises an exhaustive library of protein structure regions (‘spheres’) adjacent to complexed ligands derived from the Protein Data Bank (PDB), along with methods to find and evaluate structural matches between a protein of interest and spheres in the library. PDBspheres uses the LGA (Local–Global Alignment) structure alignment algorithm as the main engine for detecting structural similarities between the protein of interest and template spheres from the library, which currently contains >2 million spheres. To assess confidence in structural matches, an all-atom-based similarity metric takes side chain placement into account. Here, we describe the PDBspheres method, demonstrate its ability to detect and characterize binding sites in protein structures, show how PDBspheres—a strictly structure-based method—performs on a curated dataset of 2528 ligand-bound and ligand-free crystal structures, and use PDBspheres to cluster pockets and assess structural similarities among protein binding sites of 4876 structures in the ‘refined set’ of the PDBbind 2019 dataset.

59 BASIC BIOLOGICAL SCIENCES↗