Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “clustering algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Online Voltage Event Detection Using Synchrophasor Data with Structured Sparsity-Inducing Norms

This paper develops an accurate and computationally efficient data-driven framework to detect voltage events from PMU data streams. It develops an innovative Proximal Bilateral Random Projection (PBRP) algorithm to quickly decompose the PMU data matrix into a low-rank matrix, a row-sparse event-pattern matrix and a noise matrix. Here, the row-sparse pattern matrix significantly distinguishes events from normal behavior. These matrices are then fed into a clustering algorithm to separate voltage events from normal operating conditions. Large-scale numerical study results on real-world PMU data show that the proposed algorithm is computationally more efficient and achieves higher F scores than state-of-the-art benchmarks.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Enhancing Cluster Identification in Atom Probe Tomography Data Using Transfer Learning

Atom Probe Tomography (APT) is a powerful technique for visualizing the atomic-scale distribution of solutes in materials, but quantitative cluster analysis of APT datasets remains a challenge due to the need for subjective parameter selection in clustering algorithms. While distance-based and density-based methods such as HDBSCAN are widely used, their performance is highly sensitive to user-defined parameters, which undermines reproducibility and accuracy. This study proposes an image-based, deep learning-aided workflow for automating parameter selection and cluster detection in APT data analysis. By projecting 3D APT point clouds onto 2D planes, we leverage pretrained convolutional neural networks (ConvNeXt-Tiny and ResNet-50) through transfer learning to predict the number of clusters present in synthetic datasets. The output is used to guide K-means clustering and estimate HDBSCAN parameters, specifically minimum cluster size and minimum sample points. This approach reduces reliance on manual parameter tuning, improving consistency and scalability. The methodology demonstrates the feasibility of using image-based deep learning for interpreting complex spatial patterns in APT data, enabling faster and more objective analysis. The complete workflow and code are made publicly available to support reproducibility and future research.

Density-based clustering↗

Building Stock Segmentation Cluster Development: Technical Reference Document

The building stock in the United States (U.S.) varies significantly as a function of several macro variables such as: climate, building type, vintage, and density. These variables change across the U.S. and can also significantly impact energy usage of the individual buildings and overall stock. For example, the square foot density and building type varies by several orders of magnitude from Manhattan to the eastern plains of Colorado. The diversity in energy use of the building stock of different areas of the U.S. is significant, and as a result, analyses that require localized results need to consider the relevant geography and the current makeup of the building stock. This document discusses the development and implementation of a stock clustering algorithm that produces a technically rigorous, consistent, and repeatable collection of geographies which are used as the basis for localized analysis. This framework considers the impact of built environment density, diversity, and climate in creating groupings of counties that create a far more nuanced analysis framework than national averages. Clustering of counties together represents a similarity of building characteristics and climate zone.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Joint Management and Optimization of Residential Natural Gas and Electricity Distribution Networks Coupled via Fuel Cells

The interesting properties of natural gas as well as the growing electric power demand worldwide have led to increasing attention to natural-gas-based distributed generation applications in electric distribution systems. This paper goes over the interdependency between a residential natural gas network and an electric distribution network that are coupled via fuel cells. The modeling of the gas network is introduced first, and then the algorithm for gas flow study is presented. The optimal placement and sizing of fuel cell based distributed generation systems are formulated to minimize the losses in both the gas and electric distribution networks, subject to their model constraints. In addition to this, in order to capture the probabilistic nature of the optimization problem under study, the K-means clustering algorithm is applied to the gas and electricity demands to determine hourly load states and their corresponding probabilities. Furthermore, simulation studies are carried out on an integrated system consisting of the IEEE 69-bus distribution feeder and a radial 27-node natural gas network to verify the developed optimization model and the proposed method.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Assessing MP2 frozen natural orbitals in relativistic correlated electronic structure calculations

The high computational scaling with the basis set size and the number of correlated electrons is a bottleneck limiting applications of coupled cluster algorithms, in particular for calculations based on two- or four-component relativistic Hamiltonians, which often employ uncontracted basis sets. This problem may be alleviated by replacing canonical Hartree–Fock virtual orbitals by natural orbitals (NOs). Here, in this paper, we describe the implementation of a module for generating NOs for correlated wavefunctions and, in particular, second order Møller–Plesset perturbation frozen natural orbitals (MP2FNOs) as a component of our novel implementation of relativistic coupled cluster theory for massively parallel architectures [Pototschnig et al. J. Chem. Theory Comput. 17, 5509, (2021)]. Our implementation can manipulate complex or quaternion density matrices, thus allowing for the generation of both Kramers-restricted and Kramers-unrestricted MP2FNOs. Furthermore, NOs are re-expressed in the parent atomic orbital (AO) basis, allowing for generating coupled cluster singles and doubles NOs in the AO basis for further analysis. By investigating the truncation errors of MP2FNOs for both the correlation energy and molecular properties—electric field gradients at the nuclei, electric dipole and quadrupole moments for hydrogen halides HX (X = F–Ts), and parity-violating energy differences for H 2 Z 2 (Z = O–Se)—we find MP2FNOs accelerate the convergence of the correlation energy in a roughly uniform manner across the Periodic Table. It is possible to obtain reliable estimates for both energies and the molecular properties considered with virtual molecular orbital spaces truncated to about half the size of the full spaces.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Artificial Neural Networks for In-Cycle Prediction of Knock Events

Downsized turbocharged engines have been increasingly popular in modern light-duty vehicles due to their fuel efficiency benefits. However, high power density in such engines is achieved thanks to high in-cylinder pressure and temperature conditions that increase knock propensity. Next-cycle control has been studied as a method to reduce the damaging effects of knock by operating the engine in a low knock probability condition. This exploratory study looks at the feasibility of in-cycle knock prediction as a tool for advanced knock control algorithms. A methodology is proposed to 1) choose in-cycle features of the pressure trace that highly correlate with knock events and 2) train artificial neural networks to predict in-cycle knock events before knock onset. The methodology was validated at different operating conditions and different levels of generalization. Precision and recall were used as metrics to evaluate the binary classifier. However, the Fowlkes-Mallows (FM) index was used to compare the result of the clustering algorithm at different operating conditions. The results showed a maximum FM index of 0.7 when the prediction was done at knock onset and a minimum FM index of 0.45 when the prediction was done at spark timing.

42 ENGINEERING↗

An interregional optimization approach for time series aggregation in continent-scale electricity system models

Modeling electric power systems with high shares of weather-dependent resources requires tradeoffs between temporal, spatial, and operational resolution. Many studies perform time series aggregation using clustering algorithms to reduce the temporal dimension, but when modeling continent-scale electricity systems that are large enough to contain multiple independent weather systems, this approach requires large numbers of representative periods to minimize errors in regional wind and solar capacity factors. Here, a new optimization-based approach for representative period selection and weighting is introduced that minimizes regional errors in average renewable capacity factors and electricity demand. The method delivers higher regional fidelity with fewer representative periods than alternative clustering methods when applied to wind, solar, and demand profiles for the contiguous United States. When representative periods are selected from multiple weather years, the optimized method reproduces regional averages with lower error than a complete 365-day time series from any single weather year. The method identifies only representative (as opposed to outlying) periods but can be combined with an iterative "stress period" identification approach to guide efficient decision-making considering both average and high-risk weather conditions.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Augmenting Graph Convolution with Distance Preserving Embedding for Improved Learning

Graph convolution incorporates topological information of a graph into learning. Message passing corresponds to traversal of a local neighborhood in classical graph algorithms. We show that incorporating additional global structures, such as shortest paths, through distance preserving embedding can improve performance. Our approach, Gavotte, significantly improves the performance of a range of popular graph neu-ral networks such as GCN, GA T,Graph SAGE, and GCNII for transductive learning. Gavotte also improves the performance of graph neural networks for full-supervised tasks, albeit to a smaller degree. As high-quality embeddings are generated by Gavotte as a by-product, we leverage clustering algorithms on these embed dings to augment the training set and introduce Gavotte+. Our results of Gavotte+ on datasets with very few labels demonstrate the advantage of augmenting graph convolution with distance preserving embedding.

Cong, Guojing↗

The Future Electron-Ion Collider

Generalized Parton Distributions include rich information and became a powerful tool for studying hadron structure. Deeply Virtual Compton Scattering is the golden channel to access GPDs. The development of high luminosity and high-acceptance detectors (Electron-Ion Collider) allows physicists to overcome the difficulty of DVCS measurements. The outstanding performance of the accelerator performs the eA collisions at the center of mass energy from 20 to 140 GeV, while the luminosity at the scale of 1034 cm?2s?1. The new type of high-density crystal, PWO-II, produced by CRYTUR in 2×2×20 cm3 will be utilized in the calorimeter. The transparency of PWO-II crystals is > 70% at 620 nm, > 60% at 420 nm, and >35% at 360 nm. The crystals also show good radiation hardness under 30 Gy radiation exposure. Ten PWO-II crystals have a uniform light yield at 30 p.e./MeV, providing sufficient light yield to reduce the fluctuation of the energy measurements. Near 3000 crystals were constructed inside the 12-sided polygon supporting structure in the simulation. The island clustering algorithm was used for reconstructed the energy. The energy resolution study by the particle gun shows the stochastic term and constant term are 1.8% and 1.2%, respectively. The spatial resolution of NEEMC varies from 5% to 15% of the crystal?s width depending on the particle?s incident angle. The pion rejection of NEEMC can reach 103 with electron efficiency > 85% when the particle?s energy is larger than 1GeV. The Pi0- identified efficiency study can be interpreted as finding the local maxima in the single cluster caused by two high energy close photons. The study results show that efficiency is nearly 100% for pi0 energy is smaller than 10 GeV, and efficiency drops to 30% with 20 GeV pi0. The new type of 3x3 pixelated AC-LGAD were wire bonded to ALTIROC for performance test. The cross-talk between the adjacent channels is about 20% for VPA and 10% for TZ. A series of the TDC characteristic measurements show that the jitter for both preamplifiers is about 20 ps, and the time-walk effect is mild for inject charge > 12 pF. Furthermore, the beta source radiation results quantify the sharing scale of the 3x3 pixels AC-LGAD (? 20%). The electronics simulation study results show no significant changes in the spatial resolution, whether ADC resolution is 8, 10, or 12 bits. The 8-bit ADC is decided to use in the EICROC as it has a smaller size and power consumption than the 10-bit and 12-bit ADC. The ECCE is one of the full detector proposals for EIC. The new type of DVCS generator, called TOPEG, is used to generate the high acceptance beam configurations of 18×110 GeV2 electron and 4He. The acceptance study shows a -1.8 < ? < -1.4 gap between the BEMC and EEMC. The 10?x,y geometry cut is applied on the Roman Pots, so the acceptance of Roman Pots quickly drops to 0% when the polar angle of the recoiled 4He < 2 mrad. This primary ECCE simulation study suggests extending the acceptance of BEMC longitudinally as the structure limits the radial size of EEMC. The acceptance of the Roman Pots is still challenging. In 2023, The overall design of ePIC was finalized by merging two full detector proposals (ECCE and ATHENA) after a series of intensive simulation studies.

Wang, Pu-Kai↗

Improving topological cluster reconstruction using calorimeter cell timing in ATLAS

Clusters of topologically connected calorimeter cells around cells with large absolute signal-to-noise ratio (topo-clusters) are the basis for calorimeter signal recon struction in the ATLAS experiment. Topological cell clus tering has proven performant in LHC Runs 1 and 2. It is, however, susceptible to out-of-time pile-up of signals from soft collisions outside the 25 ns proton-bunch-crossing window associated with the event’s hard collision. To reduce this effect, a calorimeter-cell timing criterion was added to the signal-to-noise ratio requirement in the clustering algorithm. Multiple versions of this criterion were tested by reconstructing hadronic signals in simulated events and Run 2 ATLAS data. The preferred version is found to reduce the out-of-time pile-up jet multiplicity by ~50% for jet p T ~ 20 GeV and by ~80% for jet p T ≳ 50 GeV, while not disrupting the reconstruction of hadronic signals of interest, and improving the jet energy resolution by up to 5% for 20 < p T < 30 GeV. Pile-up is also suppressed for other physics objects based on topo-clusters (electrons, photons, τ-leptons), reducing the overall event size on disk by about 6% in early Run 3 pile up conditions. Offline reconstruction for Run 3 includes the timing requirement.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

A New CIERRA Gridded Product Climatology for LIS/OTD

The GLM-CIERRA processing generates two products: A level-2 cluster feature dataset that merges flashes that should have been a single flash following the clustering algorithm flash definition (for GLM, this is due to the LCFA thresholds); A set of level-3 gridded products including Flash Extent Density, Convective Probability, etc. The CIERRA Level-3 products are generated on a 0.1 degree grid. This is suboptimal given the spatial variations in GLM pixel size AND the size of gridpoints in the output grid. As a stepping-stone towards developing improved GLM-CIERRA grids, I am testing pixel matching techniques on the LIS/OTD data. The goal is to maintain accuracy while vastly improving computational efficiency. The CIERRA reclustering codes are also being applied to LIS / OTD to merge flashes split by the “first fit” clustering technique. The end result will be standardized grids at the nominal resolution of each instrument (5 km for LIS, 10 km for OTD) that take into account the pixel geometry for each event that comprises a given reclustered flash.

54 ENVIRONMENTAL SCIENCES↗

Subhalos in Galaxy Clusters: Coherent Accretion and Internal Orbits

Subhalo dynamics in galaxy cluster host halos govern the observed distribution and properties of cluster member galaxies. We use the IllustrisTNG simulation to investigate the accretion and orbits of subhalos found in cluster-size halos. We find that the median change in the major axis direction of cluster-size host halos is approximately 80° between a ∼ 0.1 and the present day. We identify coherent regions in the angular distribution of subhalo accretion, and ∼68% of accreted subhalos enter their host halo through ∼38% of the surface area at the virial radius. The majority of galaxy clusters in the sample have ∼2 such coherent regions. We further measure angular orbits of subhalos with respect to the host major axis and use a clustering algorithm to identify distinct orbit modes with varying oscillation timescales. The orbit modes correlate with subhalo accretion conditions. Subhalos in orbit modes with shorter oscillations tend to have lower peak masses and accretion directions somewhat more aligned with the major axis. One orbit mode, exhibiting the least oscillatory behavior, largely consists of subhalos that accrete near the plane perpendicular to the host halo major axis. Our findings are consistent with expectations from inflow from major filament structures and internal dynamical friction: most subhalos accrete through coherent regions, and more massive subhalos experience fewer orbits after accretion. Our work offers a unique quantification of subhalo dynamics that can be connected to how the intracluster medium strips and quenches cluster galaxies.

galaxy clusters↗

Advances in ArborX to support exascale applications

ArborX is a performance portable geometric search library developed as part of the Exascale Computing Project (ECP). In this paper, we explore a collaboration between ArborX and a cosmological simulation code HACC. Large cosmological simulations on exascale platforms encounter a bottleneck due to the in-situ analysis requirements of halo finding, a problem of identifying dense clusters of dark matter (halos). This problem is solved by using a density-based DBSCAN clustering algorithm. With each MPI rank handling hundreds of millions of particles, it is imperative for the DBSCAN implementation to be efficient. In addition, the requirement to support exascale supercomputers from different vendors necessitates performance portability of the algorithm. We describe how this challenge problem guided ArborX development, and enhanced the performance and the scope of the library. We explore the improvements in the basic algorithms for the underlying search index to improve the performance, and describe several implementations of DBSCAN in ArborX. Further, we report the history of the changes in ArborX and their effect on the time to solve a representative benchmark problem, as well as demonstrate the real world impact on production end-to-end cosmology simulations.

97 MATHEMATICS AND COMPUTING↗

Studying the hadron structure with PANDA and CLAS using machine learning techniques

The hadron spectroscopy and structure are currently very active fields of research to study the non-perturbative regime of quantum chronodynamics. The first one studies the complex structure of excited hadrons by looking at their decay products, while the latter uses lepton scattering on nucleons. Both methods require reconstruction algorithms with great efficiency and good particle identification and background rejection rates. This work aims to provide these by either improving the existing methods or developing new ones. The first part of this document presents a feasibility study of a predicted hybrid charmonium state for the PANDA experiment. Lattice QCD calculations predict the ground state hybrid charmonium to be a spin exotic with quantum numbers of JP C = 1?+ at a mass of around 4.3 GeV with a width to be around 20 MeV. A machine learning based data analysis scheme is proposed to further improve the signal efficiency and the background reduction, alongside with improvements of the analysis software (PandaRoot), that are vital for this study. These improvements include a reworked clustering algorithm for the electromagnetic calorimeter (EMC) and an optimized monte carlo matching for neutral particles. The second part of this document is about studying the proton structure. A multidimensional study of the structure function ratio Fsin(?)LU /FUU has been performed for K±, based on the measurement of beam-spin asymmetries. It uses the high statistics data recorded with the CLAS12 spectrometer at Jefferson Laboratory. Fsin(?)LU is a twist-3 quantity that provides information about the quark gluon correlations in the proton. This document will present for the first time a simultaneous analysis of two kaon channels over a large kinematic range of z, xB , PT and Q2 with virtualities Q2 ranging from 1 GeV2 up to 8 GeV2 using machine learning techniques for improved particle identification.

Kripko, Aron↗

Machine Learning-Enabled Quantitative Analysis of Optically Obscure Scratches on Nickel-Plated Additively Manufactured (AM) Samples

Additively manufactured metal components often have rough and uneven surfaces, necessitating post-processing and surface polishing. Hardness is a critical characteristic that affects overall component properties, including wear. This study employed K-means unsupervised machine learning to explore the relationship between the relative surface hardness and scratch width of electroless nickel plating on additively manufactured composite components. The Taguchi design of experiment (TDOE) L9 orthogonal array facilitated experimentation with various factors and levels. Initially, a digital light microscope was used for 3D surface mapping and scratch width quantification. However, the microscope struggled with the reflections from the shiny Ni-plating and scatter from small scratches. To overcome this, a scanning electron microscope (SEM) generated grayscale images and 3D height maps of the scratched Ni-plating, thus enabling the precise characterization of scratch widths. Optical identification of the scratch regions and quantification were accomplished using Python code with a K-means machine-learning clustering algorithm. The TDOE yielded distinct Ni-plating hardness levels for the nine samples, while an increased scratch force showed a non-linear impact on scratch widths. The enhanced surface quality resulting from Ni coatings will have significant implications in various industrial applications, and it will play a pivotal role in future metal and alloy surface engineering.

36 MATERIALS SCIENCE↗

Characterizing Signatures of Geothermal Exploration Data with Machine Learning Techniques: An Application to the Nevada Play Fairway Analysis

We are introducing machine learning methods to the play fairway analysis to generate geothermal potential maps to support the evaluation of geothermal resource potential and the exploration for undiscovered blind geothermal systems in the Nevada Great Basin region. Our project aims to identify new ways to combine the play fairway data and empirically organize relationships between feature weights and labels in an improved workflow. As a means of doing this, we introduce machine learning methods to evaluate the influence of certain geological and geophysical features/feature sets in predicting geothermal favorability. This report highlights promising approaches based on supervised and unsupervised learning methods. First, we demonstrate a filter method applied to supervised classification modeling. The supervised filter method is based on permutation analysis to evaluate every possible feature combination/drop out scenario and rank feature influence based on the performance variance of supervised classification models. Additionally, we present an unsupervised factor analysis based on principal component analysis coupled with a semi-supervised kmeans clustering algorithm. This analysis allows us to identify the optimal number of groups/clusters for training sites and structural settings to identify feature patterns including correlation, variance, and latent and dominant feature relationships. The results from these methods offer a promising avenue for identifying favorable sources of predictive information to identify the locations of blind geothermal systems and furthering our understanding of complex geothermal feature and label relationships in the Great Basin region and beyond.

15 GEOTHERMAL ENERGY↗

Photon Reconstruction in the Belle II Calorimeter Using Graph Neural Networks

We present the study of a fuzzy clustering algorithm for the Belle II electromagnetic calorimeter using Graph Neural Networks. We use a realistic detector simulation including simulated beam backgrounds and focus on the reconstruction of both isolated and overlapping photons. We find significant improvements of the energy resolution compared to the currently used reconstruction algorithm for both isolated and overlapping photons of more than 30% for photons with energies E γ < 0.5 GeV and high levels of beam backgrounds. Overall, the GNN reconstruction improves the resolution and reduces the tails of the reconstructed energy distribution and therefore is a promising option for the upcoming high luminosity running of Belle II.

Calorimeter↗

Unsupervised Image-Based Classification of Corrosion Severity in Automobile Engine Connecting Rods

Corrosion in engine connecting rods is a critical issue in the automotive industry, potentially leading to catastrophic engine failure, monetary losses, and safety hazards. The labor shortage in the industry further emphasizes the need for fast, accurate, and automated corrosion detection methods to ensure appropriate surface treatments can be applied to restore component integrity. We present an unsupervised image-based framework for classifying corrosion severity in automobile engine connecting rods using short-wave infrared (SWIR) and telecentric grayscale imaging. We employ the structural similarity index measure (SSIM) as a dissimilarity metric and the k-medians clustering algorithm for classification. Our algorithm achieves an overall accuracy of 80.64% for SWIR images, with 100% accuracy in classifying highly corroded samples. For grayscale images, the method attains an overall accuracy of 77.42%, with 90.91% accuracy for highly corroded samples. The method’s ability to work with different imaging modalities and its high accuracy in identifying severe corrosion cases make it a promising tool for automated corrosion assessment in the automotive industry, potentially improving efficiency and safety in engine component maintenance.

42 ENGINEERING↗