Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “clustering algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Fast shared-memory streaming multilevel graph partitioning

In this report we show that a fast parallel graph partitioner can benefit many applications by reducing data transfers. The online methods for partitioning graphs have to be fast and they often rely on simple one-pass streaming algorithms, while the offline methods for partitioning graphs contain more involved algorithms and the most successful methods in this category belong to the multilevel approaches. In this work, we assess the feasibility of using streaming graph partitioning algorithms within the multilevel framework. Our end goal is to come up with a fast parallel offline multilevel partitioner that can produce competitive cutsize quality. We rely on a simple but fast and flexible streaming algorithm throughout the entire multilevel framework. This streaming algorithm serves multiple purposes in the partitioning process: a clustering algorithm in the coarsening, an effective algorithm for the initial partitioning, and a fast refinement algorithm in the uncoarsening. Its simple nature also lends itself easily for parallelization. The experiments on various graphs show that our approach is on the average up to 5.1x faster than the multi-threaded MeTiS, which comes at the expense of only 2x worse cutsize.

97 MATHEMATICS AND COMPUTING↗

Deep Learning Based Frequency Stability Assessment in Power Grid with High Renewables

Frequency stability assessment is one critical aspect of power system security assessment. Traditional N-1 screening method is based on the simulations of a few typical daily and seasonal operation scenarios. However, the increasing integration of inverter-based renewables and the retirement of conventional synchronous generators result in decreasing system inertia and growing complexity of system operating conditions. Selecting a few typical operation scenarios cannot cover all operating conditions, and the time-domain simulation of all operation conditions requires tremendous time. This paper proposes a more efficient frequency stability assessment method based on deep learning. The affinity propagation clustering algorithm is used to divide the dataset into different clusters, so the selected dataset for training can cover the diversified operating conditions as much as possible. Also, feature normalization is applied to both the training dataset and testing dataset in order to remove any unnecessary bias. Especially, trained model based on full dataset normalization has bounded error in the prediction. The case study on the reduced 240-bus WECC system demonstrates that the proposed method can predict accurate frequency nadir with limited training dataset. The deep learning model using the revised feature normalization can predict more accurate frequency nadir than that using the traditional feature normalization and has very small maximum prediction error.

affinity propagation↗

ArborX

ArborX library tackles a problem of efficiently finding geometric objects that are close in space. Variations of this problem, such as finding the nearest neighbors of a point, or finding all objects within a certain distance, are inherent components of applications in many fields. The data may be large so that solving the problem efficiently may require significant computational resources, such as multiple processors or accelerators such as general purpose GPUs. ArborX' main advantage in its ability to solve large problems efficiently utilizing a combination of distributed and on-node parallelism. ArborX can be run efficiently on a wide variety of hardware, including GPUs from different vendors, which distinguishes it from other available libraries which typically choose only few of these. The other advantage is that it supports both types of user problems: spatial problems (useful for intersections and finding objects within certain distance), and nearest neighbor problems. ArborX also supports flexible interface in its interaction with a user. Particularly, it allows a user to call user's own function on a positive match, a functionality not rarely available in other libraries. ArborX implements construction and traversal algorithms using efficient tree structures, such as bounding volume hierarchy (BVH). At its core, it uses linear BVH for its low construction cost and sufficient quality. ArborX is written using C++, and is parallelized using the message passing interface (MPI) for the distributed communication, and the Kokkos library for on-node parallelism. This approach allows ArborX to be run on a wide variety of hardware, from common laptops and desktops to supercomputers while using the same codebase. ArborX also implements several advanced algorithms using geometric search, such as density-based clustering algorithm DBSCAN.

ECP↗

A bi-level spatiotemporal clustering approach and its application to drought extraction

We present a novel flexible bi-level spatiotemporal clustering algorithm to extract events based on their intensity and spatiotemporal structures. Our algorithm consists of using (i) a novel space-time k-means clustering to obtain spatiotemporally coherent intensity clusters, and (ii) a density-based spatial clustering of applications with noise (DBSCAN) to spatiotemporally section the intensity clusters into individual events. We discuss the development of the algorithm, the selection, tuning and meaning of the parameters within each step, as well as its validation. Finally, we apply the algorithm to a spatiotemporal drought index, standardized vapor pressure deficit drought index (SVDI), over the continental United States (US) from 1980–2021 and show that it captures historical drought events over the continental United States and their spatiotemporal extents.

17 WIND ENERGY↗

Serial crystallography with multi-stage merging of thousands of images

KAMO and BLEND provide particularly effective tools to automatically manage the merging of large numbers of data sets from serial crystallography. The requirement for manual intervention in the process can be reduced by extending BLEND to support additional clustering options such as the use of more accurate cell distance metrics and the use of reflection-intensity correlation coefficients to infer `distances' among sets of reflections. This increases the sensitivity to differences in unit-cell parameters and allows clustering to assemble nearly complete data sets on the basis of intensity or amplitude differences. If the data sets are already sufficiently complete to permit it, one applies KAMO once and clusters the data using intensities only. When starting from incomplete data sets, one applies KAMO twice, first using unit-cell parameters. In this step, either the simple cell vector distance of the original BLEND or the more sensitive NCDist is used. This step tends to find clusters of sufficient size such that, when merged, each cluster is sufficiently complete to allow reflection intensities or amplitudes to be compared. One then uses KAMO again using the correlation between reflections with a common hkl to merge clusters in a way that is sensitive to structural differences that may not have perturbed the unit-cell parameters sufficiently to make meaningful clusters. Many groups have developed effective clustering algorithms that use a measurable physical parameter from each diffraction still or wedge to cluster the data into categories which then can be merged, one hopes, to yield the electron density from a single protein form. Since these physical parameters are often largely independent of one another, it should be possible to greatly improve the efficacy of data-clustering software by using a multi-stage partitioning strategy. Here, one possible approach to multi-stage data clustering is demonstrated. The strategy is to use unit-cell clustering until the merged data are sufficiently complete and then to use intensity-based clustering. Using this strategy, it is demonstrated that it is possible to accurately cluster data sets from crystals that have subtle differences.

36 MATERIALS SCIENCE↗

Mitigate: An Adaptive Network Data Anonymization Tool Using Condensation-Based Differential Privacy

Modern network devices collect a large amount of data that can be analyzed to identify bottlenecks, anomalies, cyber-attacks, etc. Therefore, there is often a need to analyze such collections of network data quite often by an external expert or by the research community. However, these collections of data contain sensitive, proprietary information. In order for the network data to be shared, it must first be anonymized. The overall objective of this project is to develop an innovative privacy management tool to anonymize network data and achieve sufficient privacy, acceptable data utility, and efficient data analysis at the same time. No existing anonymization methods can achieve all of these at the same time. The core of this technology is a differential private clustering algorithm that provides strong privacy protection, preserves data properties important for subsequent analysis, and allows the party receiving the anonymized data to conduct analysis directly on anonymized data without the need of decryption or any extra processing. The research carried out was to design, implement and verify a solution to this problem by completing the following tasks: 1) developing the core technology; 2) developing a context based method that automatically recommends fields that must be anonymized; 3) conducted experiments showing superior results using our approach compared to existing tools, and 4) developed an intuitive but basic user interface. The research that was conducted generated novel algorithmic techniques that utilize state-of-the-art methods such as condensation, differential privacy preservation, clustering, automated tuning based on contextual awareness, and recommendation techniques to specify columns to users for anonymization leading to optimal privacy that allows research analysis on the dataset. Experiments were conducted to evaluate the efficacy of these novel algorithmic techniques by performing analysis on original non-anonymized datasets, then conducting analysis on the same yet anonymized datasets and comparing the results of the analyses. Overall, the anonymized analysis results were within 1% of the original results, verifying that the generated technology not only guarantees a high level of privacy but also enables research analysis as if it were conducted on the original dataset. Potential applications of this technology include anonymization of any type of structured network datasets that contain sensitive identifiers, such as IP addresses, that can be used in multiple applications. For example, to create an AI or machine learning model for cyber security, e.g., to detect attacks, or for performance analysis, e.g., identify bottlenecks or predict performance. In addition, a market analysis that was conducted for potential applications of this technology identified a broader range of applications of our anonymization technology beyond the network sector that includes healthcare, banking, insurance, securities, finance (FISB), data brokering, cloud services, ad sales, and government.

97 MATHEMATICS AND COMPUTING↗

Comparing Synoptic Pattern Evolution for Flash‐Flood‐Producing and Non‐Flash‐Flood‐Producing Mesoscale Convective Systems in the United States

Understanding how the short-term evolution of synoptic weather patterns influence Mesoscale Convective Systems (MCSs) is essential, as these systems are responsible for over half of central U.S. flash floods, leading to substantial socioeconomic and water resource management impacts. This study analyzes long-term MCS data, flash flood reports, and atmospheric reanalyses from 2007 to 2017 using a machine learning clustering algorithm to examine how the synoptic weather patterns evolve prior to MCS initiation. While the clusters reflect seasonal and regional differences in MCS occurrence, they do not consistently distinguish between MCSs that do and do not produce flash floods. Systems in the southern Great Plains are more flood-prone when a synoptic-scale forcing, located near the system, drives strong water vapor transport from the nearby moisture source. More generally under different synoptic weather patterns, a broader precipitating area is the most dominant factor governing MCS flash flood potential.

atmospheric dynamics↗

Intelligent Experiments through Real-Time AI: Fast Data Processing and Autonomous Detector Control for High-Energy Nuclear Experiments

The aim of this project is to develop software and hardware for fast real-time data processing and autonomous detector control and calibration for the sPHENIX and the future EIC experiments. Below summarizes Georgia Tech team efforts in the past year: 1. We developed a real-time clustering algorithm and FPGA-based pipeline architecture for processing fired pixel data from ALPIDE sensors in sPHENIX experiments. Our Columnar Clustering Co-Design introduces a hardware-aware, stream-friendly approach that segments pixel data by column pairs using a Column Pair Clustering (CPC) strategy, followed by Cluster Stitching to merge adjacent subclusters. Implemented in Vitis HLS, the pipeline comprises five stages—read-in, subclustering, stitching, analysis, and write-out—connected by tagged HLS streams with custom end-of-event signaling for robust synchronization. We designed a pipelined dataflow model optimized for throughput, low latency, and minimal buffering, enabling scalable clustering across events of arbitrary size. Our system maintains spatial precision via center-of-mass and shape key extraction and efficiently handles edge cases such as fragmented or nested clusters. Compared against DBSCAN in both software and hardware, our approach demonstrates competitive performance under FPGA constraints. 2. We also conducted a comprehensive algorithm-to-hardware co-design of connected component analysis tailored for sPHENIX experiments, focusing on real-time, low-latency processing using FPGAs and High-Level Synthesis (HLS). Starting from a Python-based particle tracking pipeline, the team translated the core logic—graph traversal via DFS and Union-Find—into an HLS-compatible C++ model, replacing dynamic memory and recursion with static arrays and pipelined control flow. The final design includes a fully streamed and dataflow-compatible Union-Find kernel optimized across five iterations, incorporating loop pipelining, array partitioning, AXI/FIFO interface tuning, and function flattening. Experimental results show up to 14.8× speedup over the CPU baseline, reducing per-graph latency to 1.58 μs and demonstrating strong resource efficiency with only ~7k LUTs and zero BRAM usage. The design maintains functional correctness against the Python reference using a Python-based C-simulation framework and Mean Squared Error metrics. This work validates the potential of HLS-driven FPGA designs for edge-level HEP data acquisition, laying a scalable foundation for future integration with real-time detector pipelines and multi-graph processing systems.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Unsupervised learning for identifying events in active target experiments

This article presents novel applications of unsupervised machine learning methods to the problem of event separation in an active target detector, the Active-Target Time Projection Chamber (AT-TPC). The overarching goal is to group similar events in the early stages of the data analysis, thereby improving efficiency by limiting the computationally expensive processing of unnecessary events. The application of unsupervised clustering algorithms to the analysis of two-dimensional projections of particle tracks from a resonant proton scattering experiment on 46 Ar is introduced. We explore the performance of autoencoder neural networks and a pre-trained VGG16 Simonyan and Zisserman (2015) convolutional neural network. We study clustering performance on both data from a simulated 46 Ar experiment, and real events from the AT-TPC detector. We find that a -means algorithm applied to simulated data in the VGG16 latent space forms almost perfect clusters. Additionally, the VGG16+-means approach finds high purity clusters of proton events for real experimental data. Here, we also explore the application of clustering the latent space of autoencoder neural networks for event separation. While these networks show strong performance, they suffer from high variability in their results.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

A Decentralized Approach for Modeling Organized Convection Based on Thermal Populations on Microgrids

Abstract In this study, a spectral model for convective transport is coupled to a thermal population model on a two‐dimensional horizontal “microgrid,” covering the typical gridbox size of general circulation models. The goal is to explore new ways of representing impacts of spatial organization in cumulus cloud fields. The thermals are considered the smallest building block of convection, with thermal life cycle and movement represented through binomial functions. Thermals interact through two simple rules, reflecting pulsating growth and environmental deformation. Long‐lived thermal clusters thus form on the microgrid, exhibiting scale growth and spacing that represent simple forms of spatial organization and memory. Size distributions of cluster number are diagnosed from the microgrid through an online clustering algorithm, and provided as input to a spectral multiplume eddy‐diffusivity mass flux scheme. This yields a decentralized transport system, in that the thermal clusters acting as independent but interacting nodes that carry information about spatial structure. The main objectives of this study are (a) to seek proof of concept of this approach, and (b) to gain insight into impacts of spatial organization on convective transport. Single‐column model experiments demonstrate satisfactory skill in reproducing two observed cases of continental shallow convection. Metrics expressing self‐organization and spatial organization match well with large‐eddy simulation results. We find that in this coupled system, spatial organization impacts convective transport primarily through the scale break in the size distribution of cluster number. The rooting of saturated plumes in the subcloud mixed layer plays a key role in this process.

54 ENVIRONMENTAL SCIENCES↗

Bi-cross validation of spectral clustering hyperparameters

One challenge impeding the analysis of terabyte scale X-ray scattering data from the Linac Coherent Light Source (LCLS) is determining the number of clusters required for the execution of traditional clustering algorithms. Here, we demonstrate that the previous work using bi-cross validation to determine the number of singular vectors directly maps to the spectral clustering problem of estimating both the number of clusters and hyperparameter values. Here by applying this method to LCLS X-ray scattering data enables the identification of dropped shots without manually setting boundaries on detector fluence and provides a path toward identifying rare and anomalous events.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Scalable Tensor Methods for Nonuniform Hypergraphs

While multilinear algebra appears natural for studying the multiway interactions modeled by hypergraphs, tensor methods for general hypergraphs have been stymied by theoretical and practical barriers. A recently proposed adjacency tensor is applicable to nonuniform hypergraphs, but is prohibitively costly to form and analyze in practice. We develop tensor times same vector (TTSV) algorithms for this tensor which improve complexity from $O(n^r)$ to a low-degree polynomial in $r$, where $n$ is the number of vertices and $r$ is the maximum hyperedge size. Our algorithms are implicit, avoiding formation of the order $r$ adjacency tensor. Here, we demonstrate the flexibility and utility of our approach in practice by developing tensor-based hypergraph centrality and clustering algorithms. We also show these tensor measures offer complementary information to analogous graph-reduction approaches on data, and are also able to detect higher-order structure that many existing matrix-based approaches provably cannot.

97 MATHEMATICS AND COMPUTING↗

Data-driven evaluation of HVAC operation and savings in commercial buildings

Commercial buildings consumed 36% of electricity, or 1.35 trillion kWh, in the United States in 2017, and almost 30% of this energy was wasted. Much of this loss can be attributed to inefficient heating ventilation and air con­ditioning (HVAC) systems. By improving the operational conditions of HVAC, significant savings can be achieved. However, most buildings and building equipment do not use costly sub-meters to monitor and address performance issues, and on-site auditing can be expensive and insufficient. Alternatively in this study, we propose a data-driven method to identify savings opportunities using only whole building meter data and without setting foot in the building. For this purpose, we introduced two algorithms that virtually quantify the value of a thermostat setpoint setback and HVAC rescheduling. Additionally, we developed novel methods for detecting occupancy patterns and quantifying the baseload of the HVAC operation. Using a clustering algorithm, we identified those buildings for which HVAC savings was significant and further categorized the buildings based on their potential for savings. A population study of over 432 commercial buildings demonstrated a median percentage energy savings of 1.6% from a baseload reduction and 2.1% from HVAC rescheduling. Additionally, results indicate that retail buildings have the highest potential for savings among the building types studied.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Partitioning of Large-Scale Power Electronics-Based Power Systems for Small-Signal Stability Analysis

The nodal admittance matrix (NAM)-based approach is suitable for analyzing the small-signal stability of large-scale power electronics-based power systems (PEPSs) as it preserves the system structure by utilizing the admittance matrix. Previously, NAM-based area partition has been proposed, which divides the system into various subareas and interconnections for easier analysis of the low-dimension matrix compared to the entire system-based high-dimension matrix. However, no partition algorithm has been presented for the NAM-based area partition method. This paper focuses on implementing the spectral partitioning algorithm for partitioning large-scale PEPSs into a low-dimension matrix to reduce the computation complexity of the analysis. These spectral components facilitate data transformation into a new space, enabling the application of traditional clustering methods like k-means. To evaluate the performance of the partitioning method, the subareas and interconnections obtained from the spectral clustering algorithm are incorporated into the NAM-based area partition method for a large system with 140 buses. The computational times of the original method, where the NAM-based criterion is directly applied to the entire system, are compared with those of the NAM-based partition method in MATLAB. PSCAD simulations of the whole system and the obtained subareas are conducted to validate the effectiveness of the proposed algorithm.

Nupur, Nupur↗

Sub-10 nm Probing of Ferroelectricity in Heterogeneous Materials by Machine Learning Enabled Contact Kelvin Probe Force Microscopy

Reducing the dimensions of ferroelectric materials down to the nanoscale has strong implications on the ferroelectric polarization pattern and on the ability to switch the polarization. As the size of ferroelectric domains shrinks to the nanometer scale, the heterogeneity of the polarization pattern becomes increasingly pronounced, enabling a large variety of possible polar textures in nanocrystalline and nanocomposite materials. Critical to the understanding of fundamental physics of such materials and hence their applications in electronic nanodevices is the ability to investigate their ferroelectric polarization at the nanoscale in a nondestructive way. We show that contact Kelvin probe force microscopy (cKPFM) combined with a k-means response clustering algorithm enables to measure the ferroelectric response at a mapping resolution of 8 nm. In a BaTiO 3 thin film on silicon composed of tetragonal and hexagonal nanocrystals, we determine a nanoscale lateral distribution of discrete ferroelectric response clusters, fully consistent with the nanostructure determined by transmission electron microscopy. Moreover, we apply this data clustering method to the cKPFM responses measured at different temperatures, which allows us to follow the corresponding change in the polarization pattern as the Curie temperature is approached and across the phase transition. This work opens up perspectives for mapping complex ferroelectric polarization textures such as curled/swirled polar textures that can be stabilized in epitaxial heterostructures and more generally for mapping the polar domain distribution of any spatially highly heterogeneous ferroelectric materials.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Nearby stellar substructures in the Galactic halo from DESI Milky Way Survey Year 1 Data Release

We report five nearby ($d_{\mathrm{helio}} < 5$ kpc) stellar substructures in the Galactic halo from a subset of 138 661 stars in the Dark Energy Spectroscopic Instrument (DESI) Milky Way Survey Year 1 Data Release. With an unsupervised clustering algorithm, HDBSCAN*, these substructures are independently identified in Integrals of Motion ($E_{\rm tot}$, $L_{\rm z}$, $\log {J_r}$, $\log {J_z}$) space and Galactocentric cylindrical velocity space ($V_{R}$, $V_{\phi }$, $V_{z}$). We associate all identified clusters with known nearby substructures (Helmi streams, M18-Cand10/MMH-1, Sequoia, Antaeus, and ED-2) previously reported in various studies. With metallicities precisely measured by DESI, we confirm that the Helmi streams, M18-Cand10, and ED-2 are chemically distinct from local halo stars. We have characterized the chemodynamic properties of each dynamic group, including their metallicity dispersions, to associate them with their progenitor types (globular cluster or dwarf galaxy). Our approach for searching substructures with HDBSCAN* reliably detects real substructures in the Galactic halo, suggesting that applying the same method can lead to the discovery of new substructures in future DESI data. With more stars from future DESI data releases and improved astrometry from the upcoming Gaia Data Release 4, we will have a more detailed blueprint of the Galactic halo, offering a significant improvement in our understanding of the formation and evolutionary history of the Milky Way Galaxy.

dynamics↗

Strong chemical tagging with APOGEE: 21 candidate star clusters that have dissolved across the Milky Way disc

ABSTRACT Chemically tagging groups of stars born in the same birth cluster is a major goal of spectroscopic surveys. To investigate the feasibility of such strong chemical tagging, we perform a blind chemical tagging experiment on abundances measured from APOGEE survey spectra. We apply a density-based clustering algorithm to the 8D chemical space defined by [Mg/Fe], [Al/Fe], [Si/Fe], [K/Fe], [Ti/Fe], [Mn/Fe], [Fe/H], and [Ni/Fe], abundances ratios which together span multiple nucleosynthetic channels. In a high-quality sample of 182 538 giant stars, we detect 21 candidate clusters with more than 15 members. Our candidate clusters are more chemically homogeneous than a population of non-member stars with similar [Mg/Fe] and [Fe/H], even in abundances not used for tagging. Group members are consistent with having the same age and fall along a single stellar-population track in log g versus Teff space. Each group’s members are distributed over multiple kpc, and the spread in their radial and azimuthal actions increases with age. We qualitatively reproduce this increase using N-body simulations of cluster dissolution in Galactic potentials that include transient winding spiral arms. Observing our candidate birth clusters with high-resolution spectroscopy in other wavebands to investigate their chemical homogeneity in other nucleosynthetic groups will be essential to confirming the efficacy of strong chemical tagging. Our initially spatially compact but now widely dispersed candidate clusters will provide novel limits on chemical evolution and orbital diffusion in the Galactic disc, and constraints on star formation in loosely bound groups.

Price-Jones, Natalie↗

Examination of Radiation Belt Dynamics During Substorm Clusters: Activity Drivers and Dependencies of Trapped Flux Enhancements

Here, dynamical variations of radiation belt trapped electron fluxes are examined to better understand the variability of enhancements linked to substorm clusters. Analysis is undertaken using the Substorm Onsets and Phases from Indices of the Electrojet substorm cluster algorithm for event detection. Observations from low earth orbit are complemented by additional measurements from medium earth orbit to allow a major expansion in the energy range considered, from medium energy energetic electrons up to ultra-relativistic electrons. The number of substorms identified inside a cluster does not depend strongly on solar wind drivers or geomagnetic indices either before, during, or after the cluster start time. Clusters of substorms linked to moderate (100 nT < AE ≤ 300 nT) or strong AE (AE ≥ 300 nT) disturbances are associated with radiation belt flux enhancements, including up to ultra-relativistic energies by the strongest substorms (as measured by strong southward Bz and high AE). These clusters reliably occur during times of high speed solar winds streams with associated increased magnetospheric convection. However, substorm clusters associated with quiet AE disturbances (AE ≤ 100 nT) lead to no significant chorus whistler mode intensity enhancements, or increases in energetic, relativistic, or ultra-relativistic electron flux in the outer radiation belts. In these cases the solar wind speed is low, and the geomagnetic Kp index indicates a lack of magnetospheric convection. Our study clearly indicates that clusters of substorms occurring outside of high speed wind streams are not by themselves sufficient to drive acceleration, which may be due to the lack of pre-cluster convection.

79 ASTRONOMY AND ASTROPHYSICS↗