Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “clustering algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 469 records · Page 26

A search for novae in M 31 globular clusters

By combining a local sky-fitting algorithm with a Fourier point-spread-function matching technique, nova outbursts have been searched for inside 54 of the globular clusters contained on the Ciardullo et al. (1987 and 1990) H-alpha survey frames of M 31. Over a mean effective survey time of about 2.0 years, no cluster exhibited a magnitude increase indicative of a nova explosion. If the cataclysmic variables (CVs) contained within globular clusters are similar to those found in the field, then these data imply that the overdensity of CVs within globulars is at least several times less than that of the high-luminosity X-ray sources. If tidal capture is responsible for the high density of hard binaries within globulars, then the probability of capturing condensed objects inside globular clusters may depend strongly on the mass of the remnant.

Ciardullo, Robin↗

ORCA: Outlier detection and Robust Clustering for Attributed graphs

Here, a framework is proposed to simultaneously cluster objects and detect anomalies in attributed graph data. Our objective function along with the carefully constructed constraints promotes interpretability of both the clustering and anomaly detection components, as well as scalability of our method. In addition, we developed an algorithm called Outlier detection and Robust Clustering for Attributed graphs (ORCA) within this framework. ORCA is fast and convergent under mild conditions, produces high quality clustering results, and discovers anomalies that can be mapped back naturally to the features of the input data. The efficacy and efficiency of ORCA is demonstrated on real world datasets against multiple state-of-the-art techniques.

97 MATHEMATICS AND COMPUTING↗

Electric vehicle supply equipment location and capacity allocation for fixed-route networks

Electric vehicle (EV) supply equipment location and allocation (EVSELCA) problems for freight vehicles are becoming more important because of the trending electrification shift. Some previous works address EV charger location and vehicle routing problems simultaneously by generating vehicle routes from scratch. Although such routes can be efficient, introducing new routes may violate practical constraints, such as drive schedules, and satisfying electrification requirements can require dramatically altering existing routes. To address the challenges in the prevailing adoption scheme, we approach the problem from a fixed -route perspective. We develop a mixed -integer linear program, a clustering approach, and a metaheuristic solution method using a genetic algorithm (GA) to solve the EVSELCA problem. The clustering approach simplifies the problem by grouping customers into clusters, while the GA generates solutions that are shown to be nearly optimal for small problem cases. A case study examines how charger costs, energy costs, the value of time (VOT), and battery capacity impact the cost of the EVSELCA. Charger equipment costs were found to be the most significant component in the objective function, leading to a substantial reduction in cost when decreased. VOT costs exhibited a significant decrease with rising energy costs. Further, an increase in VOT resulted in a notable rise in the number of fast chargers. Longer EV ranges decrease total costs up to a certain point, beyond which the decrease in total costs is negligible.

33 ADVANCED PROPULSION SYSTEMS↗

Constraining Galaxy-Halo connection using machine learning

We investigate the potential of machine learning (ML) methods to model small-scale galaxy clustering for constraining Halo Occupation Distribution (HOD) parameters. Our analysis reveals that while many ML algorithms report good statistical fits, they often yield likelihood contours that are significantly biased in both mean values and variances relative to the true model parameters. This highlights the importance of careful data processing and algorithm selection in ML applications for galaxy clustering, as even seemingly robust methods can lead to biased results if not applied correctly. ML tools offer a promising approach to exploring the HOD parameter space with significantly reduced computational costs compared to traditional brute-force methods if their robustness is established. Using our ANN-based pipeline, we successfully recreate some standard results from recent literature. Properly restricting the HOD parameter space, transforming the training data, and carefully selecting ML algorithms are essential for achieving unbiased and robust predictions. Among the methods tested, artificial neural networks (ANNs) outperform random forests (RF) and ridge regression in predicting clustering statistics, when the HOD prior space is appropriately restricted. We demonstrate these findings using the projected two-point correlation function (w p (r p )), angular multipoles of the correlation function (ξ ℓ (r)), and the void probability function (VPF) of Luminous Red Galaxies from Dark Energy Spectroscopic Instrument mocks. Our results show that while combining w p (r p ) and VPF improves parameter constraints, adding the multipoles ξ 0 , ξ 2 , and ξ 4 to w p (r p ) does not significantly improve the constraints.

cosmology↗

Parallel Implementation of the Recursive Approximation of an Unsupervised Hierarchical Segmentation Algorithm

The hierarchical image segmentation algorithm (referred to as HSEG) is a hybrid of hierarchical step-wise optimization (HSWO) and constrained spectral clustering that produces a hierarchical set of image segmentations. HSWO is an iterative approach to region grooving segmentation in which the optimal image segmentation is found at N(sub R) regions, given a segmentation at N(sub R+1) regions. HSEG's addition of constrained spectral clustering makes it a computationally intensive algorithm, for all but, the smallest of images. To counteract this, a computationally efficient recursive approximation of HSEG (called RHSEG) has been devised. Further improvements in processing speed are obtained through a parallel implementation of RHSEG. This chapter describes this parallel implementation and demonstrates its computational efficiency on a Landsat Thematic Mapper test scene.

Tilton, James C.↗

Critical points of the random cluster model with Newman–Ziff sampling

Here, we present a method for computing transition points of the random cluster model using a generalization of the Newman–Ziff algorithm, a celebrated technique in numerical percolation, to the random cluster model. The new method is straightforward to implement and works for real cluster weight q > 0. Furthermore, results for an arbitrary number of values of q can be found at once within a single simulation. Because the algorithm used to sweep through bond configurations is identical to that of Newman and Ziff, which was conceived for percolation, the method loses accuracy for large lattices when q > 1. However, by sampling the critical polynomial, accurate estimates of critical points in two dimensions can be found using relatively small lattice sizes, which we demonstrate here by computing critical points for non-integer values of q on the square lattice, to compare with the exact solution, and on the unsolved non-planar square matching lattice. The latter results would be much more difficult to obtain using other techniques.

97 MATHEMATICS AND COMPUTING↗

Identifying Vehicle Signals in Continuous Seismic Data Using Unsupervised Machine-Learning Techniques

Seismic sensors deployed near roadways effectively capture ground vibrations generated by passing vehicles. Although both traditional and machine‐learning algorithms have been utilized for analyzing such signals, independent validation of detected vehicle events remains limited. We applied two unsupervised machine‐learning algorithms, uniform manifold approximation and projection for dimension reduction, and hierarchical density‐based spatial clustering of applications with noise, to continuous seismic data collected along a road on the main campus of Oak Ridge National Laboratory. The algorithms identified seven distinct cluster labels across the entire dataset. By comparing these cluster labels with precipitation records from a nearby weather station and image‐derived labels from a local camera system, we identified one cluster associated with rainfall and another with vehicle activity. Our algorithms identified a greater number of vehicle‐related labels compared to the camera‐derived labels because seismic data are unaffected by poor lighting conditions. The arrival times of the newly detected vehicle signals corresponded well with the road’s speed limit, supporting our findings. Our algorithm outperformed the short‐term average/long‐term average method and k‐means clustering. Our results suggest that seismic data, when analyzed with machine‐learning algorithms, can complement existing vehicle monitoring systems, particularly under challenging environmental conditions.

Chai, Chengping [Oak Ridge National Laboratory (OR↗

Optimization of Support Vector Machine (SVM) for Object Classification

The Support Vector Machine (SVM) is a powerful algorithm, useful in classifying data into species. The SVMs implemented in this research were used as classifiers for the final stage in a Multistage Automatic Target Recognition (ATR) system. A single kernel SVM known as SVMlight, and a modified version known as a SVM with K-Means Clustering were used. These SVM algorithms were tested as classifiers under varying conditions. Image noise levels varied, and the orientation of the targets changed. The classifiers were then optimized to demonstrate their maximum potential as classifiers. Results demonstrate the reliability of SVM as a method for classification. From trial to trial, SVM produces consistent results.

support vector machice (SVM)↗

Testing of the Support Vector Machine for Binary-Class Classification

The Support Vector Machine is a powerful algorithm, useful in classifying data in to species. The Support Vector Machines implemented in this research were used as classifiers for the final stage in a Multistage Autonomous Target Recognition system. A single kernel SVM known as SVMlight, and a modified version known as a Support Vector Machine with K-Means Clustering were used. These SVM algorithms were tested as classifiers under varying conditions. Image noise levels varied, and the orientation of the targets changed. The classifiers were then optimized to demonstrate their maximum potential as classifiers. Results demonstrate the reliability of SMV as a method for classification. From trial to trial, SVM produces consistent results

autonomous target recognition systemr↗

Automated characterization of spatial and dynamical heterogeneity in supercooled liquids via implementation of machine learning

Abstract A computational approach by an implementation of the principle component analysis (PCA) with K -means and Gaussian mixture (GM) clustering methods from machine learning algorithms to identify structural and dynamical heterogeneities of supercooled liquids is developed. In this method, a collection of the average weighted coordination numbers ( W C N s ‾ ) of particles calculated from particles’ positions are used as an order parameter to build a low-dimensional representation of feature (structural) space for K -means clustering to sort the particles in the system into few meso-states using PCA. Nano-domains or aggregated clusters are also formed in configurational (real) space from a direct mapping using associated meso-states’ particle identities with some misclassified interfacial particles. These classification uncertainties can be improved by a co-learning strategy which utilizes the probabilistic GM clustering and the information transfer between the structural space and configurational space iteratively until convergence. A final classification of meso-states in structural space and domains in configurational space are stable over long times and measured to have dynamical heterogeneities. Armed with such a classification protocol, various studies over the thermodynamic and dynamical properties of these domains indicate that the observed heterogeneity is the result of liquid–liquid phase separation after quenching to a supercooled state.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

The composite sequential clustering technique for analysis of multispectral scanner data

The clustering technique consists of two parts: (1) a sequential statistical clustering which is essentially a sequential variance analysis, and (2) a generalized K-means clustering. In this composite clustering technique, the output of (1) is a set of initial clusters which are input to (2) for further improvement by an iterative scheme. This unsupervised composite technique was employed for automatic classification of two sets of remote multispectral earth resource observations. The classification accuracy by the unsupervised technique is found to be comparable to that by traditional supervised maximum likelihood classification techniques. The mathematical algorithms for the composite sequential clustering program and a detailed computer program description with job setup are given.

Su, M. Y.↗

Normalized Cut Algorithm for Automated Assignment of Protein Domains

We present a novel computational method for automatic assignment of protein domains from structural data. At the core of our algorithm lies a recently proposed clustering technique that has been very successful for image-partitioning applications. This grap.,l-theory based clustering method uses the notion of a normalized cut to partition. an undirected graph into its strongly-connected components. Computer implementation of our method tested on the standard comparison set of proteins from the literature shows a high success rate (84%), better than most existing alternative In addition, several other features of our algorithm, such as reliance on few adjustable parameters, linear run-time with respect to the size of the protein and reduced complexity compared to other graph-theory based algorithms, would make it an attractive tool for structural biologists.

Samanta, M. P.↗

Elucidating the Molecular Origins of the Transference Number in Battery Electrolytes Using Computer Simulations

The rate at which rechargeable batteries can be charged and discharged is governed by the selective transport of the working ions through the electrolyte. Conductivity, the parameter commonly used to characterize ion transport in electrolytes, reflects the mobility of both cations and anions. The transference number, a parameter introduced over a century ago, sheds light on the relative rates of cation and anion transport. This parameter is, not surprisingly, affected by cation–cation, anion–anion, and cation–anion correlations. In addition, it is affected by correlations between the ions and neutral solvent molecules. Computer simulations have the potential to provide insights into the nature of these correlations. We review the dominant theoretical approaches used to predict the transference number from simulations by using a model univalent lithium electrolyte. In electrolytes of low concentration, one can obtain a quantitative model by assuming that the solution is made up of discrete ion-containing clusters–neutral ion pairs, negatively and positively charged triplets, neutral quadruplets, and so on. These clusters can be identified in simulations using simple algorithms, provided their lifetimes are sufficiently long. In concentrated electrolytes, more clusters are short-lived and more rigorous approaches that account for all correlations are necessary to quantify transference. Elucidating the molecular origin of the transference number in this limit remains an unmet challenge.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Message Passing vs. Shared Address Space on a Cluster of SMPs

The convergence of scalable computer architectures using clusters of PCs (or PC-SMPs) with commodity networking has become an attractive platform for high end scientific computing. Currently, message-passing and shared address space (SAS) are the two leading programming paradigms for these systems. Message-passing has been standardized with MPI, and is the most common and mature programming approach. However message-passing code development can be extremely difficult, especially for irregular structured computations. SAS offers substantial ease of programming, but may suffer from performance limitations due to poor spatial locality, and high protocol overhead. In this paper, we compare the performance of and programming effort, required for six applications under both programming models on a 32 CPU PC-SMP cluster. Our application suite consists of codes that typically do not exhibit high efficiency under shared memory programming. due to their high communication to computation ratios and complex communication patterns. Results indicate that SAS can achieve about half the parallel efficiency of MPI for most of our applications: however, on certain classes of problems SAS performance is competitive with MPI. We also present new algorithms for improving the PC cluster performance of MPI collective operations.

Shan, Hongzhang↗

Reconstruction of Thermal Protection System Aeroheating using a Green’s Function Approach

Inverse heat transfer (IHT) techniques are often used to reconstruct the surface heating conditions on spacecraft thermal protection systems (TPS) during atmospheric entry. Current IHT techniques for entry spacecraft applications, however, demand substantial computational resources, and are impractical for analyses such as uncertainty quantification and real-time health monitoring. In this paper, a Green’s function sensor fusion approach is used to reconstruct the TPS surface aeroheating conditions on experimental spaceflight and ground test systems from collocated temperature and heat flux sensors embedded in the TPS. The algorithm leverages Green’s functions to model the heat conduction within the spacecraft TPS and stabilizes the recovery of the surface heating condition using the direct heat flux sensor measurement. The algorithm is validated using arc-jet ground test data and applied to the reconstruction of the Mars 2020 backshell heating during Martian atmospheric entry. The performance of the algorithm is benchmarked against a current state-of-the-art IHT framework, FIAT_Opt. The Green’s function-based reconstruction algorithm recovers the net hot-wall heat flux absorbed by the TPS and the incident heat flux from the atmospheric entry environment in close agreement with FIAT_Opt. Notably, computation of the surface heating condition is completed in three orders of magnitude less time with the Green’s function sensor fusion approach using a consumer-grade PC, versus with FIAT_Opt running on a high performance computer cluster. The efficiency of the algorithm is leveraged to compute the uncertainty contributions of input parameters to the total uncertainty in reconstructed Mars 2020 backshell heating for the full atmospheric entry heat pulse. The sensitivity analysis uncovers that, at different times throughout the entry heat pulse, uncertainties in the TPS specific heat, thermal conductivity, and emissivity are all dominant drivers of the reconstruction uncertainty. These results demonstrate Green’s functions and sensor-fusion techniques as promising IHT approaches to reconstruct atmospheric entry environments from TPS-embedded measurements, and highlight how these techniques may give access to post-flight analyses previously hindered by the prohibitive cost of current methods.

Kenneth McAfee↗

Reconstruction of Thermal Protection System Aeroheating using a Green’s Function Approach

Inverse heat transfer (IHT) techniques are often used to reconstruct the surface heating conditions on spacecraft thermal protection systems (TPS) during atmospheric entry. Current IHT techniques for entry spacecraft applications, however, demand substantial computational resources, and are impractical for analyses such as uncertainty quantification and real-time health monitoring. In this paper, a Green’s function sensor fusion approach is used to reconstruct the TPS surface aeroheating conditions on experimental spaceflight and ground test systems from collocated temperature and heat flux sensors embedded in the TPS. The algorithm leverages Green’s functions to model the heat conduction within the spacecraft TPS and stabilizes the recovery of the surface heating condition using the direct heat flux sensor measurement. The algorithm is validated using arc-jet ground test data and applied to the reconstruction of the Mars 2020 backshell heating during Martian atmospheric entry. The performance of the algorithm is benchmarked against a current state-of-the-art IHT framework, FIAT_Opt. The Green’s function-based reconstruction algorithm recovers the net hot-wall heat flux absorbed by the TPS and the incident heat flux from the atmospheric entry environment in close agreement with FIAT_Opt. Notably, computation of the surface heating condition is completed in three orders of magnitude less time with the Green’s function sensor fusion approach using a consumer-grade PC, versus with FIAT_Opt running on a high performance computer cluster. The efficiency of the algorithm is leveraged to compute the uncertainty contributions of input parameters to the total uncertainty in reconstructed Mars 2020 backshell heating for the full atmospheric entry heat pulse. The sensitivity analysis uncovers that, at different times throughout the entry heat pulse, uncertainties in the TPS specific heat, thermal conductivity, and emissivity are all dominant drivers of the reconstruction uncertainty. These results demonstrate Green’s functions and sensor-fusion techniques as promising IHT approaches to reconstruct atmospheric entry environments from TPS-embedded measurements, and highlight how these techniques may give access to post-flight analyses previously hindered by the prohibitive cost of current methods.

Kenneth McAfee↗

Massively parallel GPU enabled third-order cluster perturbation excitation energies for cost-effective large scale excitation energy calculations

We present here a massively parallel implementation of the recently developed CPS(D-3) excitation energy model that is based on cluster perturbation theory. The new algorithm extends the one developed in Baudin et al. [J. Chem. Phys., 150, 134110 (2019)] to leverage multiple nodes and utilize graphical processing units for the acceleration of heavy tensor contractions. Furthermore, we show that the extended algorithm scales efficiently with increasing amounts of computational resources and that the developed code enables CPS(D-3) excitation energy calculations on large molecular systems with a low time-to-solution. More specifically, calculations on systems with over 100 atoms and 1000 basis functions are possible in a few hours of wall clock time. This establishes CPS(D-3) excitation energies as a computationally efficient alternative to those obtained from the coupled-cluster singles and doubles model.

38 RADIATION CHEMISTRY, RADIOCHEMISTRY, AND NUCLEA↗

Artificial Diversity and Defense Security (ADDSec)

Artificial Diversity and Defense Security (ADDSec) machine learning algorithms are used to classify and cluster threats so that an appropriate response can be initiated as a mitigation strategy. The package includes an ensemble of machine learning algorithms such as Support Vector Machines, naïve bayes, logistic regression, and random forest that evolve with the data to recognize anomalous behavior at the host and network levels. Inputs into the machine learning algorithms include end host system calls, system utilization, packet captures, and syslog messages. The machine learning algorithms can be retrained based on user defined intervals or on the number of packets received. ADDSEC's threat responses include Internet Protocol (IP) Address randomization, application port number randomization, and application library randomization. The IP randomization implementation is built on top of a Software Defined Networking (SDN) framework. The SDN controller installs flows on each of the SDN switches with randomized source and destination IP addresses. The application port numbers are randomized using iptables. The application library randomization is created with a LLVM compiler. All randomization schemes are transparent to the endpoints on the network. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525. SAND2021-3379 O

Cox, RebeccaE.↗