Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “clustering algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21

Recursive Hierarchical Image Segmentation by Region Growing and Constrained Spectral Clustering

This paper describes an algorithm for hierarchical image segmentation (referred to as HSEG) and its recursive formulation (referred to as RHSEG). The HSEG algorithm is a hybrid of region growing and constrained spectral clustering that produces a hierarchical set of image segmentations based on detected convergence points. In the main, HSEG employs the hierarchical stepwise optimization (HS WO) approach to region growing, which seeks to produce segmentations that are more optimized than those produced by more classic approaches to region growing. In addition, HSEG optionally interjects between HSWO region growing iterations merges between spatially non-adjacent regions (i.e., spectrally based merging or clustering) constrained by a threshold derived from the previous HSWO region growing iteration. While the addition of constrained spectral clustering improves the segmentation results, especially for larger images, it also significantly increases HSEG's computational requirements. To counteract this, a computationally efficient recursive, divide-and-conquer, implementation of HSEG (RHSEG) has been devised and is described herein. Included in this description is special code that is required to avoid processing artifacts caused by RHSEG s recursive subdivision of the image data. Implementations for single processor and for multiple processor computer systems are described. Results with Landsat TM data are included comparing HSEG with classic region growing. Finally, an application to image information mining and knowledge discovery is discussed.

Tilton, James C.↗

Search for interacting galaxy clusters from SDSS DR-17 employing optimized friends-of-friends algorithm and multimessenger tracers

ABSTRACT In the theoretical framework of hierarchical structure formation, galaxy clusters evolve through continuous accretion and mergers of substructures. Cosmological simulations have revealed the best picture of the universe as a 3D filamentary network of dark-matter distribution called the cosmic web. Galaxy clusters are found to form at the nodes of this network and are the regions of high merging activity. Such mergers being highly energetic, contain a wealth of information about the dynamical evolution of structures in the Universe. Observational validation of this scenario needs a colossal effort to identify numerous events from all-sky surveys. Therefore, such efforts are sparse in literature and tend to focus on individual systems. In this work, we present an improved search algorithm for identifying interacting galaxy clusters and have successfully produced a comprehensive list of systems from SDSS DR-17. By proposing a set of physically motivated criteria, we classified these interacting clusters into two broad classes, ‘merging’ and ‘pre-merging/postmerging’ systems. Interestingly, as predicted by simulations, we found that most cases show cluster interaction along the prominent cosmic filaments of galaxy distribution (i.e. the proxy for dark matter filaments), with the most violent ones at their nodes. Moreover, we traced the imprint of interactions through multiband signatures, such as diffuse cluster emissions in radio or X-rays. Although we could not find direct evidence of diffuse emission from connecting filaments and ridges; our catalogue of interacting clusters will ease locating such faintest emissions as data from sensitive telescopes such as eROSITA or SKA, becomes accessible.

Oak, Tejas↗

The HectoMAP Cluster Survey: Spectroscopically Identified Clusters and their Brightest Cluster Galaxies (BCGs)

We apply a friends-of-friends (FoF) algorithm to identify galaxy clusters and we use the catalog to explore the evolutionary synergy between brightest cluster galaxies (BCGs) and their host clusters. We base the cluster catalog on the dense HectoMAP redshift survey (2000 redshifts deg -2 ). The HectoMAP FoF catalog includes 346 clusters with 10 or more spectroscopic members within the range 0.05 < z < 0.55 and with a median z = 0.29. We list these clusters and their members. We also include central velocity dispersions (σ*, BCG ) for the FoF cluster BCGs, a distinctive feature of the HectoMAP FoF catalog. HectoMAP clusters with higher galaxy number density (80 systems) are all genuine clusters with a strong concentration and a prominent BCG in Subaru/Hyper Suprime-Cam images. The phase-space diagrams show the expected elongation along the line of sight. Lower-density systems include some low reliability systems. We establish a connection between BCGs and their host clusters by demonstrating that σ*, BCG /σ cl decreases as a function of cluster velocity dispersion (σ cl ), in contrast, numerical simulations predict a constant σ*, BCG /σ cl . Sets of clusters at two different redshifts show that BCG evolution in massive systems is slow over the redshift range z < 0.4. The data strongly suggest that minor mergers may play an important role in BCG evolution in clusters with σ cl ≳ 300 km s -1 . For lower mass systems (σ cl < 300 km s -1 ), major mergers may play a significant role. The coordinated evolution of BCGs and their host clusters provides an interesting test of simulations in high-density regions of the universe.

79 ASTRONOMY AND ASTROPHYSICS↗

High performance FPGA embedded system for machine learning based tracking and trigger in sPhenix and EIC

We present a comprehensive end-to-end pipeline to classify triggers versus background events in this paper. This pipeline makes online decisions to select signal data and enables the intelligent trigger system for efficient data collection in the Data Acquisition System (DAQ) of the upcoming sPHENIX and future EIC (Electron-Ion Collider) experiments. Starting from the coordinates of pixel hits that are lightened by passing particles in the detector, the pipeline applies three-stage of event processing (hits clustering, track reconstruction, and trigger detection) and labels all processed events with the binary tag of trigger versus background events. The pipeline consists of deterministic algorithms such as clustering pixels to reduce event size, tracking reconstruction to predict candidate edges, and advanced graph neural network-based models for recognizing the entire jet pattern. In particular, we apply the message-passing graph neural network to predict links between hits and reconstruct tracks and a hierarchical pooling algorithm (DiffPool) to make the graph-level trigger detection. We obtain an impressive performance (≥70% accuracy) for trigger detection with only 3200 neuron weights in the end-to-end pipeline. We deploy the end-to-end pipeline into a field-programmable gate array (FPGA) and accelerate the three stages with speedup factors of 1152, 280, and 21, respectively.

Instruments & Instrumentation↗

Characterization of surficial geologic units on Venus from Pioneer Venus radar data: A progress report

A classification database using the reflectivity (derived from the altimetry data), rms slope, and the first principal component of altimetry and topographic slope is presented. The resultant clustered data is examined qualitatively as well as quantitatively, to establish the statistical integrity of each cluster by use of an interactive, ternary plotting algorithm. This algorithm plots, for a cluster, the position of each of its pixels within a ternary diagram whose apices represent reflectivity, rms slope, and the first principal component. The digital values in these three databases are normalized such that unity is represented by a value of 255 in each database. The frequencies of each plotted point within the ternary diagram are recorded in order to establish the mode of each cluster. The pixels of each cluster are displayed as one separate color; their ternary plot will show not only the interrelations between clusters, but also the presence of any anomalous points within a cluster. Existing lunar and terrestrial analog radar data is used to establish fields within this ternary diagram that are indicative of as many different geologic materials and tectonics settings as possible. The resultant fields are used to determine empirically the geologic significance of the clusters resulting from the cluster analysis.

Davis, P. A.↗

Automated System for Early Breast Cancer Detection in Mammograms

The increasing demand on mammographic screening for early breast cancer detection, and the subtlety of early breast cancer signs on mammograms, suggest an automated image processing system that can serve as a diagnostic aid in radiology clinics. We present a fully automated algorithm for detecting clusters of microcalcifications that are the most common signs of early, potentially curable breast cancer. By using the contour map of the mammogram, the algorithm circumvents some of the difficulties encountered with standard image processing methods. The clinical implementation of an automated instrument based on this algorithm is also discussed.

Bankman, Isaac N.↗

Analyzing acoustic emission data to identify cracking modes in cement paste using an artificial neural network

This research is focused on the identification of cracking mechanisms for cement paste using acoustic emission data, recorded from compression and notched four-point bending tests. A procedure is developed for analyzing the data by employing an agglomerative hierarchical clustering method, an artificial neural network, and a ray-tracing source location algorithm. An agglomerative hierarchical clustering method is utilized to cluster the AE data from a compression test using frequency-dependent features. A neural network is trained using the compression test data and applied to the AE data emitted during the four-point bending test. The clustered data from the four-point bending test is localized using a ray-tracing algorithm. Based on the occurrence and locations of the clustered events and signal feature analyses, potential cracking mechanisms are identified and assigned.

36 MATERIALS SCIENCE↗

Implementation of Relativistic Coupled Cluster Theory for Massively Parallel GPU-Accelerated Computing Architectures

In this paper, we report reimplementation of the core algorithms of relativistic coupled cluster theory aimed at modern heterogeneous high-performance computational infrastructures. The code is designed for parallel execution on many compute nodes with optional GPU coprocessing, accomplished via the new ExaTENSOR back end. The resulting ExaCorr module is primarily intended for calculations of molecules with one or more heavy elements, as relativistic effects on the electronic structure are included from the outset. In the current work, we thereby focus on exact two-component methods and demonstrate the accuracy and performance of the software. The module can be used as a stand-alone program requiring a set of molecular orbital coefficients as the starting point, but it is also interfaced to the DIRAC program that can be used to generate these. We therefore also briefly discuss an improvement of the parallel computing aspects of the relativistic self-consistent field algorithm of the DIRAC program.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Parallel Wavefront Analysis for a 4D Interferometer

This software provides a programming interface for automating data collection with a PhaseCam interferometer from 4D Technology, and distributing the image-processing algorithm across a cluster of general-purpose computers. Multiple instances of 4Sight (4D Technology s proprietary software) run on a networked cluster of computers. Each connects to a single server (the controller) and waits for instructions. The controller directs the interferometer to several images, then assigns each image to a different computer for processing. When the image processing is finished, the server directs one of the computers to collate and combine the processed images, saving the resulting measurement in a file on a disk. The available software captures approximately 100 images and analyzes them immediately. This software separates the capture and analysis processes, so that analysis can be done at a different time and faster by running the algorithm in parallel across several processors. The PhaseCam family of interferometers can measure an optical system in milliseconds, but it takes many seconds to process the data so that it is usable. In characterizing an adaptive optics system, like the next generation of astronomical observatories, thousands of measurements are required, and the processing time quickly becomes excessive. A programming interface distributes data processing for a PhaseCam interferometer across a Windows computing cluster. A scriptable controller program coordinates data acquisition from the interferometer, storage on networked hard disks, and parallel processing. Idle time of the interferometer is minimized. This architecture is implemented in Python and JavaScript, and may be altered to fit a customer s needs.

Rao, Shanti R.↗

NETRA: A parallel architecture for integrated vision systems 2: Algorithms and performance evaluation

In part 1 architecture of NETRA is presented. A performance evaluation of NETRA using several common vision algorithms is also presented. Performance of algorithms when they are mapped on one cluster is described. It is shown that SIMD, MIMD, and systolic algorithms can be easily mapped onto processor clusters, and almost linear speedups are possible. For some algorithms, analytical performance results are compared with implementation performance results. It is observed that the analysis is very accurate. Performance analysis of parallel algorithms when mapped across clusters is presented. Mappings across clusters illustrate the importance and use of shared as well as distributed memory in achieving high performance. The parameters for evaluation are derived from the characteristics of the parallel algorithms, and these parameters are used to evaluate the alternative communication strategies in NETRA. Furthermore, the effect of communication interference from other processors in the system on the execution of an algorithm is studied. Using the analysis, performance of many algorithms with different characteristics is presented. It is observed that if communication speeds are matched with the computation speeds, good speedups are possible when algorithms are mapped across clusters.

Choudhary, Alok N.↗

Cluster Analysis of Spectroscopic Line Profiles and EUV Emission in RMHD Simulations and Observations of the Solar Atmosphere

Spatially-resolved observations from the IRIS, SDO/AIA, and other space mission and ground-based telescopes, coupled with realistic 3D RMHD simulations, are a powerful tool for analysis of processes in the solar atmosphere. To better understand the dynamical and thermodynamic properties in the simulation data and their connection to observations, it is essential to determine similarities in the behaviors of the synthesized and observed emission. However, the complexity of observational data and physical processes makes comparison of observations and modeling results difficult. In this work, we show the initial results of application of K-Means clustering (unsupervised machine learning) algorithm to two different problems: 1) recognition of the typical spectroscopic line profiles observed by IRIS during solar flares and their typical dynamic behavior; 2) recognition of shocks and heating events in synthetic AIA emission data obtained from StellarBox quiet-Sun simulations. The average silhouette width technique for the KMeans algorithm is utilized in different ways to obtain optimal numbers of clusters. We discuss application of the emission clustering to visualizations of the computational volume, understanding its evolutionary trends and behavior patterns, and inversion (reconstruction) of physical properties of the solar atmosphere from synthesizes emission data.

Sadykov, Viacheslav↗

Improving the Accuracy of Clustering Electric Utility Net Load Data using Dynamic Time Warping

Identifying patterns in electric utility net load data in a time-series format is very useful in preparing the operation for next day. Machine learning algorithms have been used in other domains and those concepts are applied in this paper on real-world net load measurement data. Clustering is the practice of grouping data with similar characteristics as determined by the distance measure. The K-means clustering algorithm is utilized here with actual electric utility data. The paper uses the standard distance measure, Euclidean distance (ED), and compares its performance against the dynamic time warping (DTW) measure. An actual case study with real data is presented, and DTW distance measure-based method observed to result better accuracy compared to the ED based method for substation net load measurements predominantly with residential customers.

clustering↗

Toward Discord : Code for Simulating Continuous Spin Systems

A new computational tool to simulate classical spin systems with frustrated crystal structures is presented. Complementary single- and cluster-spin flip algorithms are implemented to calculate the diffuse scattering patterns, spin-pair correlations, and thermodynamic quantities. Test cases of geometrically frustrated kagome, pyrochlore, and cubic systems are detailed. Two recent scientific cases are also shown here. This new method, together with recent developments of the rmc-discord package (https://github.com/zjmorgan/rmc-discord), represent integrated and strategic step in a complete forward and reverse Monte Carlo framework discord.

36 MATERIALS SCIENCE↗

Selecting representative geological realizations to model subsurface CO 2 storage under uncertainty

Carbon capture and storage (CCS) is one of the quickest and most effective solutions for reducing carbon emissions. The majority of subsurface storage occurs in saline aquifers, for which geological information is lacking which in turn results in geological uncertainty. To evaluate uncertainty in CO 2 injection projections, the use of multiple geological realizations (GRs) has been practiced very commonly. In this approach, hundreds or thousands of high-resolution GRs is used that quickly becomes computationally expensive. This issue can be addressed with representative geological realizations (RGRs) that preserve the uncertainty domain of the ensemble GRs. Here, in this study, we propose the use of unsupervised machine learning (UML) frameworks, including dissimilarity measurement, dimensionality reduction, clustering and sampling algorithms ta select a predetermined number of RGRs. We compare the simulation outputs of the RGR sets and the ensemble using the Kolmogorov–Smirnov (KS) test to select the best UML. The UML frameworks and their associated selection processes are evaluated using a saline aquifer with a single CO 2 injection well and 200 GRs with varying uncertain petrophysical characteristics. The best UML framework is selected to use only 5% of the GRs while maintaining the uncertainty domain of the ensemble GRs. In addition, the best UML framework is tested using a saline aquifer with three CO 2 injection wells and varied GRs. The results show that our proposed UML framework can be used to choose RGRs, capturing the whole uncertainty domain. Our approach leads to a significant reduction in the computational cost associated with scenario testing, decision-making, and development planning for CO 2 storage sites under geological uncertainty.

58 GEOSCIENCES↗

Landfalling Droughts: Global Tracking of Moisture Deficits From the Oceans Onto Land

Abstract Droughts threaten food, energy, and water security, causing death and displacement of millions of people and billions of dollars in damages. However, there are still important gaps in the understanding of drought mechanisms and behaviors, inhibiting the accuracy of early‐warning systems designed to protect communities worldwide. We use an object‐tracking algorithm to track clusters of precipitation‐minus‐evaporation moisture deficits across land and ocean areas of the globe from 1981–2018. This analysis reveals a new type of “landfalling drought” that originates over the ocean and “migrates” onto land. We find that 16% of droughts that affected the continents worldwide from 1981–2018 were landfalling droughts. These droughts were significantly larger (220–425%) and more intense (4–30%)—and grew (253–285%) and intensified (9–28%) faster—than droughts that developed solely over the land or ocean. To identify potential underlying mechanisms, we analyze moisture transport associated with landfalling droughts over western North America. We find that landfalling droughts in this region are associated with anomalously anticyclonic atmospheric pressure patterns that reduce moisture fluxes over the Pacific Ocean toward the continent. By advancing understanding of the spatiotemporal evolution of droughts, our findings offer the potential to improve seasonal‐scale prediction and long‐term projection of global drought risks.

Herrera‐Estrada, Julio E.↗

DeepAstroUDA: semi-supervised universal domain adaptation for cross-survey galaxy morphology classification and anomaly detection

Abstract Artificial intelligence methods show great promise in increasing the quality and speed of work with large astronomical datasets, but the high complexity of these methods leads to the extraction of dataset-specific, non-robust features. Therefore, such methods do not generalize well across multiple datasets. We present a universal domain adaptation method, DeepAstroUDA , as an approach to overcome this challenge. This algorithm performs semi-supervised domain adaptation (DA) and can be applied to datasets with different data distributions and class overlaps. Non-overlapping classes can be present in any of the two datasets (the labeled source domain, or the unlabeled target domain), and the method can even be used in the presence of unknown classes. We apply our method to three examples of galaxy morphology classification tasks of different complexities (three-class and ten-class problems), with anomaly detection: (1) datasets created after different numbers of observing years from a single survey (Legacy Survey of Space and Time mock data of one and ten years of observations); (2) data from different surveys (Sloan Digital Sky Survey (SDSS) and DECaLS); and (3) data from observing fields with different depths within one survey (wide field and Stripe 82 deep field of SDSS). For the first time, we demonstrate the successful use of DA between very discrepant observational datasets. DeepAstroUDA is capable of bridging the gap between two astronomical surveys, increasing classification accuracy in both domains (up to 40 % on the unlabeled data), and making model performance consistent across datasets. Furthermore, our method also performs well as an anomaly detection algorithm and successfully clusters unknown class samples even in the unlabeled target dataset.

79 ASTRONOMY AND ASTROPHYSICS↗

Using modularity to segment binary code

We consider the problem of recovering program structure from compiled binary code. We first extract the call graph and layout of functions in memory from the compiled code and represent this information in a graphical format. We then employ Louvain's modularity algorithm to identify clusters of functions that are considered to be related. We find that the quality and properties of clusters extracted by our technique are greatly impacted by the relative importance we assign to the call graph and the ordering of functions in memory.

97 MATHEMATICS AND COMPUTING↗

A VXS [VITA41] Trigger Processor for the 12GEV Experimental Programs at Jefferson Lab

The VXS_Trigger_Processor [VTP] was developed and commissioned for CLAS12 in the fall of 2016. This board is a VITA41 switch card and it collects data from a variety of front-end TDC and Flash ADC modules. The VTP has since been used in several experiments at Jefferson Lab serving as the L1 trigger module for a variety of detector types, such as stacked calorimeters, strip calorimeters, time-of-flight, Cerenkov, hodoscopes, drift chambers, and silicon strips. Trigger algorithms implemented include cluster finding (1D, 2D), drift chamber segment and road finding, geometry matching between various detectors, particle counting, and general global trigger bit processing. The VTP is also capable of reading out each front-end crate with up to 40Gbps Ethernet which is an enormous increase compared to the currently used 200MB/s VME bus. Recent progress has been made to show that a firmware and software upgrade can enable existing Jefferson Lab front-end crates to operate in a streaming DAQ mode. In February 2020, tests will be performed on a full calorimeter and matched hodoscope which are components of the CLAS12 Forward Tagger detector system with beam in Hall B. This paper details the hardware performance, triggered, and streaming applications that have been implemented using the VTP for several experiments at Jefferson Lab.

ABBOTT, David↗