Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “clustering algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 451 records · Page 25

Pattern recognition for Space Applications Center director's discretionary fund

Results and conclusions are presented on the application of recent developments in pattern recognition to spacecraft star mapping systems. Sensor data for two representative starfields are processed by an adaptive shape-seeking version of the Fc-V algorithm with good results. Cluster validity measures are evaluated, but not found especially useful to this application. Recommendations are given two system configurations worthy of additional study,

Singley, M. E.↗

Predicting thunderstorm evolution using ground-based lightning detection networks

Lightning measurements acquired principally by a ground-based network of magnetic direction finders are used to diagnose and predict the existence, temporal evolution, and decay of thunderstorms over a wide range of space and time scales extending over four orders of magnitude. The non-linear growth and decay of thunderstorms and their accompanying cloud-to-ground lightning activity is described by the three parameter logistic growth model. The growth rate is shown to be a function of the storm size and duration, and the limiting value of the total lightning activity is related to the available energy in the environment. A new technique is described for removing systematic bearing errors from direction finder data where radar echoes are used to constrain site error correction and optimization (best point estimate) algorithms. A nearest neighbor pattern recognition algorithm is employed to cluster the discrete lightning discharges into storm cells and the advantages and limitations of different clustering strategies for storm identification and tracking are examined.

Goodman, Steven J.↗

Knowledge Driven Image Mining with Mixture Density Mercer Kernels

This paper presents a new methodology for automatic knowledge driven image mining based on the theory of Mercer Kernels; which are highly nonlinear symmetric positive definite mappings from the original image space to a very high, possibly infinite dimensional feature space. In that high dimensional feature space, linear clustering, prediction, and classification algorithms can be applied and the results can be mapped back down to the original image space. Thus, highly nonlinear structure in the image can be recovered through the use of well-known linear mathematics in the feature space. This process has a number of advantages over traditional methods in that it allows for nonlinear interactions to be modelled with only a marginal increase in computational costs. In this paper, we present the theory of Mercer Kernels, describe its use in image mining, discuss a new method to generate Mercer Kernels directly from data, and compare the results with existing algorithms on data from the MODIS (Moderate Resolution Spectral Radiometer) instrument taken over the Arctic region. We also discuss the potential application of these methods on the Intelligent Archive, a NASA initiative for developing a tagged image data warehouse for the Earth Sciences.

Srivastava, Ashok N.↗

Knowledge Driven Image Mining with Mixture Density Mercer Kernals

This paper presents a new methodology for automatic knowledge driven image mining based on the theory of Mercer Kernels, which are highly nonlinear symmetric positive definite mappings from the original image space to a very high, possibly infinite dimensional feature space. In that high dimensional feature space, linear clustering, prediction, and classification algorithms can be applied and the results can be mapped back down to the original image space. Thus, highly nonlinear structure in the image can be recovered through the use of well-known linear mathematics in the feature space. This process has a number of advantages over traditional methods in that it allows for nonlinear interactions to be modelled with only a marginal increase in computational costs. In this paper we present the theory of Mercer Kernels; describe its use in image mining, discuss a new method to generate Mercer Kernels directly from data, and compare the results with existing algorithms on data from the MODIS (Moderate Resolution Spectral Radiometer) instrument taken over the Arctic region. We also discuss the potential application of these methods on the Intelligent Archive, a NASA initiative for developing a tagged image data warehouse for the Earth Sciences.

Srivastava, Ashok N.↗

Wavefront Sensing and Control Architecture for the Spherical Primary Optical Telescope (SPOT)

Testbed results are presented demonstrating high-speed image-based wavefront sensing and control for a spherical primary optical telescope (SPOT). The testbed incorporates a phase retrieval camera coupled to a 3-Mirror Vertex testbed (3MV) at the NASA Goddard Space Flight Center. Actuator calibration based on the Hough transform is discussed as well as several supercomputing archtectures for image-based wavefront sensing. Timing results are also presented based on various algorithm implementations using a cluster of 64 TigerShare TSlOl DSP's (digital-signal processors).

Dean, Bruce H.↗

The Software Correlator of the Chinese VLBI Network

The software correlator of the Chinese VLBI Network (CVN) has played an irreplaceable role in the CVN routine data processing, e.g., in the Chinese lunar exploration project. This correlator will be upgraded to process geodetic and astronomical observation data. In the future, with several new stations joining the network, CVN will carry out crustal movement observations, quick UT1 measurements, astrophysical observations, and deep space exploration activities. For the geodetic or astronomical observations, we need a wide-band 10-station correlator. For spacecraft tracking, a realtime and highly reliable correlator is essential. To meet the scientific and navigation requirements of CVN, two parallel software correlators in the multiprocessor environments are under development. A high speed, 10-station prototype correlator using the mixed Pthreads and MPI (Massage Passing Interface) parallel algorithm on a computer cluster platform is being developed. Another real-time software correlator for spacecraft tracking adopts the thread-parallel technology, and it runs on the SMP (Symmetric Multiple Processor) servers. Both correlators have the characteristic of flexible structure and scalability.

Zheng, Weimin↗

Coronado Ecological Conservation: Assessing Vegetation Change Due to Border Wall Construction and Shifting Social Trails

Species monitoring is essential for mitigating the impacts of plant invasion, such as radical changes in an area’s ecosystem, degraded soil health, increased wildfire severity, landslides, and increased flooding. For this project, NASA DEVELOP partnered with the National Park Service (NPS) to investigate invasive species in disturbed lands: specifically, areas affected by off-trail travel and U.S.-Mexico border construction activities. The team assessed how construction has impacted the distribution of Lehmann’s lovegrass and Russian thistle invasives throughout Coronado National Memorial, AZ from 1986-2022. Using data from Landsat 5 and 8, Sentinel-2, NAIP, and PlanetScope, the team computed NDVI, NDMI, MSAVI2, EVI, and Tasseled Cap Wetness, Brightness, and Greenness transformations as vegetation health indicators to input into various machine learning algorithms. To minimize noise, the team conducted Principal Component Analysis on vegetation indices and spectral bands before running k-means clustering and random forest classification algorithms. Between all datasets, the team found that the median area fully overtaken by invasive plants was 5.37% of the park’s total area in 2022. The NPS will use end products to help increase restoration efforts in disturbed areas with high concentrations of invasive plants, and this project can serve as a jumping off point for future invasive species monitoring. The NPS’s collection of ground data for 2022-2023, in conjunction with future data collection, will notably improve the accuracy of classification models, leading to more precise monitoring of invasive species spread over time.

Coronado National Memorial↗

Parallel computational fluid dynamics - Implementations and results

The present volume on parallel CFD discusses implementations on parallel machines, numerical algorithms for parallel CFD, and performance evaluation and computer science issues. Attention is given to a parallel algorithm for compressible flows through rotor-stator combinations, a massively parallel Euler solver for unstructured grids, a fast scheme to analyze 3D disk airflow on a parallel computer, and a block implicit multigrid solution of the Euler equations. Topics addressed include a 3D ADI algorithm on distributed memory multiprocessors, clustered element-by-element computations for fluid flow, hypercube FFT and the Fourier pseudospectral method, and an investigation of parallel iterative algorithms for CFD. Also discussed are fluid dynamics using interface methods on parallel processors, sorting for particle flow simulation on the connection machine, a large grain mapping method, and efforts toward a Teraflops capability for CFD.

Simon, Horst D.↗

Coronado Ecological Conservation: Assessing Vegetation Change Due to Border Wall Construction and Shifting Social Trails

Species monitoring is essential in mitigating the impacts of plant invasion, such as radical changes in an area’s ecosystem, degraded soil health, increased wildfire severity, landslides, and increased flooding. NASA DEVELOP partnered with the National Park Service (NPS) to investigate invasive species in disturbed lands: specifically, areas affected by off-trail walking and US-Mexico border construction activities. The team assessed how construction has impacted the distribution of Lehmann’s lovegrass and Russian thistle invasives throughout Coronado National Memorial, AZ from 1986 to 2022. Using data from Landsat 5 and 8, Sentinel-2, the National Agriculture Imagery Program, and PlanetScope, the team computed vegetation indices including the Normalized Difference Vegetation Index, Normalized Difference Moisture Index, Modified Soil Adjusted Vegetation Index 2, Enhanced Vegetation Index, and Tasseled Cap Wetness, Brightness, and Greenness transformations as vegetation health indicators to input into various machine learning algorithms. To minimize noise, the team conducted Principal Component Analysis on the vegetation indices and spectral bands before running k-means++ clustering and random forest classification algorithms. Between all datasets, we found the median area fully overtaken by invasive plants was 5.37% of the park’s total area in 2022. The NPS will use the end products to help increase restoration efforts in disturbed areas with high concentrations of invasive plants. The NPS’s collection of ground data for 2022–2023, in conjunction with future data collection, will notably improve the accuracy of classification models, leading to more precise monitoring of invasive spread over time.

Carson Schuetze↗

The WaZP galaxy cluster sample of the dark energy survey year 1

We present a new (2+1)D galaxy cluster finder based on photometric redshifts called Wavelet Z Photometric (WaZP) applied to DES first year (Y1A1) data. The results are compared to clusters detected by the South Pole Telescope (SPT) survey and the redMaPPer cluster finder, the latter based on the same photometric data. WaZP searches for clusters in wavelet-based density maps of galaxies selected in photometric redshift space without any assumption on the cluster galaxy populations. The comparison to other cluster samples was performed with a matching algorithm based on angular proximity and redshift difference of the clusters. It led to the development of a new approach to match two optical cluster samples, following an iterative approach to minimize incorrect associations. The WaZP cluster finder applied to DES Y1A1 galaxy survey (1511.13 deg^2 up to = 23 mag) led to the detection of 60 547 galaxy clusters with redshifts 0.05 < z < 0.9 and richness N_gals ≥ 5. Considering the overlapping regions and redshift ranges between the DES Y1A1 and SPT cluster surveys, all sz based SPT clusters are recovered by the WaZP sample. The comparison between WaZP and redMaPPer cluster samples showed an excellent overall agreement for clusters with richness N_gals (λ for redMaPPer) greater than 25 (20), with 95 per cent recovery on both directions. Based on the cluster cross-match, we explore the relative fragmentation of the two cluster samples and investigate the possible signatures of unmatched clusters.

79 ASTRONOMY AND ASTROPHYSICS↗

Scalable Pattern Matching in Metadata Graphs via Constraint Checking

Pattern matching is a fundamental tool for answering complex graph queries. Unfortunately, existing solutions have limited capabilities: They do not scale to process large graphs and/or support only a restricted set of search templates or usage scenarios. Moreover, the algorithms at the core of the existing techniques are not suitable for today’s graph processing infrastructures relying on horizontal scalability and shared-nothing clusters, as most of these algorithms are inherently sequential and difficult to parallelize. In this article we present an algorithmic pipeline that bases pattern matching on constraint checking. The key intuition is that each vertex and edge participating in a match has to meet a set of constraints implicitly specified by the search template. These constraints can be verified independently and typically are less expensive to compute than searching the full template. The pipeline we propose generates these constraints and iterates over them to eliminate all the vertices and edges that do not participate in any match, thus reducing the background graph to a subgraph that is the union of all template matches—the complete set of all vertices and edges that participate in at least one match. Additional analysis can be performed on this annotated, reduced graph, such as full match enumeration, match counting, or computing vertex/edge centrality. Furthermore, a vertex-centric formulation for constraint checking algorithms exists, and this makes it possible to harness existing high-performance, vertex-centric graph processing frameworks. This technique (i) enables highly scalable pattern matching in metadata (labeled) graphs; (ii) supports arbitrary patterns with 100% precision; (iii) enables tradeoffs between precision and time-to-solution, while always selects all vertices and edges that participate in matches, thus offering 100% recall; and (iv) supports a set of popular data analytics scenarios. We implement our approach on top of HavoqGT, an open-source asynchronous graph processing framework, and demonstrate its advantages through strong and weak scaling experiments on massive scale real-world (up to 257 billion edges) and synthetic (up to 4.4 trillion edges) labeled graphs, respectively, and at scales (1,024 nodes / 36,864 cores), orders of magnitude larger than used in the past for similar problems. This article serves two purposes: First, it synthesises the knowledge accumulated during a long-term project. Second, it presents new system features, usage scenarios, optimizations, and comparisons with related work that strengthen the confidence that pattern matching based on iterative pruning via constraint checking is an effective and scalable approach in practice. The new contributions include the following: (i) We demonstrate the ability of the constraint checking approach to efficiently support two additional search scenarios that often emerge in practice, interactive incremental search and exploratory search. (ii) We empirically compare our solution with two additional state-of-the-art systems, Arabsque and TriAD. (iii) We show the ability of our solution to accommodate a more diverse range of datasets with varying properties, e.g., scale, skewness, label distribution, and match frequency. (iv) We introduce or extend a number of system features (e.g., work aggregation, load balancing, and the ability to cap the generated traffic) and design optimizations and demonstrate their advantages with respect to improving performance and scalability. (v) We present bottleneck analysis and insights into artifacts that influence performance. (vi) We present a theoretical complexity argument that motivates the performance gains we observe.

97 MATHEMATICS AND COMPUTING↗

Jacobian-scaled K-means clustering for physics-informed segmentation of reacting flows

This work introduces Jacobian-scaled K-means (JSK-means) clustering, which is a physicsinformed clustering strategy centered on the K-means framework. The method allows for the injection of underlying physical knowledge into the clustering procedure through a distance function modification: instead of leveraging conventional Euclidean distance vectors, the JSKmeans procedure operates on distance vectors scaled by matrices obtained from dynamical system Jacobians evaluated at the cluster centroids. The goal of this work is to show how the JSKmeans algorithm - without modifying the input dataset - produces clusters that capture regions of dynamical similarity, in that the clusters are redistributed towards high-sensitivity regions in phase space and are described by similarity in the source terms of samples instead of the samples themselves. The algorithm is demonstrated on a complex reacting flow simulation dataset (a channel detonation configuration), where the dynamics in the thermochemical composition space are known through the highly nonlinear and stiff Arrhenius-based chemical source terms. Interpretations of cluster partitions in both physical space and composition space reveal how JSK-means shifts clusters produced by standard K-means towards regions of high chemical sensitivity (e.g., towards regions of peak heat release rate near the detonation reaction zone). Furthermore, the findings presented here illustrate the benefits of utilizing Jacobian-scaled distances in clustering techniques, and the JSK-means method in particular displays promising potential for improving former partition-based modeling strategies in reacting flow (and other multi-physics) applications.

Clustering↗

Small bimetallic clusters Agn-1M (M = Au, Co, Cu, Ni, Pd, Pt; n = 3, 9, 15): Density functional theory and genetic algorithm

We investigated the effect of size and composition on the properties of bimetallic nanoclusters. The geometric structures, stabilities, and electronic properties of size-selected Ag n-1 M (M = Au, Co, Cu, Ni, Pd, Pt; n = 3, 9, 15) bimetallic nanoclusters are systematically analyzed using spin-polarized density functional theory (DFT) within the generalized gradient approximation (GGA). We determine the most stable geometries for these clusters using a genetic algorithm (GA) in combination with DFT. Our results show that doping pure silver clusters with an M atom (transition metal), referred to as a “guest atom”, increases the stability as compared to pure Ag n (n = 3, 9, 15) clusters. The results for various properties including formation energy per atom, electronic structure, magnetic moments, and vibrational density of states (VDOS) are evaluated as a function of both size and composition of the system. The adsorption of selected bimetallic clusters on hydroxylated alumina substrate shows weak binding and minor changes in geometric properties except for Ag 8 Pt.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Weak lensing mass-richness relation of redMaPPer clusters in LSST DESC DC2 simulations

Cluster scaling relations are key ingredients in cluster abundance-based cosmological studies. In optical cluster cosmology, where clusters are detected through their richness, cluster-weak gravitational lensing has proven to be a powerful tool to constrain the cluster mass-richness relation. This work is conducted as part of the Dark Energy Science Collaboration (DESC), which aims to analyze the Legacy Survey of Space and Time (LSST) of the Vera C. Rubin Observatory, starting in 2026. Cluster properties inferred from weak lensing, such as mass, suffer from several sources of bias. In this paper, we aim to test the impact of modeling choices and observational systematics in cluster lensing on the inference of the mass-richness relation. We constrained the mass-richness relation of 3600 clusters detected by the redMaPPer algorithm in the cosmoDC2 extragalactic mock catalog of the LSST DESC DC2 simulation, covering 440 deg 2 , using number count measurements and either stacked weak lensing profiles or mean cluster masses in several intervals of richness (20 ≤ λ ≤ 200) and redshift (0.2 ≤ z ≤ 1). We provide the first constraints on the redMaPPer cluster mass-richness relation detected in cosmoDC2. We find that for an LSST-like source galaxy density, our constraints are robust to changes in the concentration-mass relation, as well as the dark matter density profile modeling choices, when source redshifts and shapes are perfectly known. We find that photometric redshift uncertainties can introduce bias at the 1σ level, which could be mitigated by an overall correction factor fitted jointly with the scaling parameters. We find that including positive shear-richness covariance in the fit shifts the results by up to 0.5σ. Our constraints also offer a fair comparison to a fiducial mass-richness relation, obtained from matching cosmoDC2 halo masses to redMaPPer-detected cluster richness results.

galaxy clusters↗

Poisson hurdle model-based method for clustering microbiome features

Abstract Motivation High-throughput sequencing technologies have greatly facilitated microbiome research and have generated a large volume of microbiome data with the potential to answer key questions regarding microbiome assembly, structure and function. Cluster analysis aims to group features that behave similarly across treatments, and such grouping helps to highlight the functional relationships among features and may provide biological insights into microbiome networks. However, clustering microbiome data are challenging due to the sparsity and high dimensionality. Results We propose a model-based clustering method based on Poisson hurdle models for sparse microbiome count data. We describe an expectation–maximization algorithm and a modified version using simulated annealing to conduct the cluster analysis. Moreover, we provide algorithms for initialization and choosing the number of clusters. Simulation results demonstrate that our proposed methods provide better clustering results than alternative methods under a variety of settings. We also apply the proposed method to a sorghum rhizosphere microbiome dataset that results in interesting biological findings. Availability and implementation R package is freely available for download at https://cran.r-project.org/package=PHclust. Supplementary information Supplementary data are available at Bioinformatics online.

59 BASIC BIOLOGICAL SCIENCES↗

Diagnosing Open Cluster Stock 2: Member Candidates and Mass Distribution with Gaia DR2 and LAMOST

We identify 1325 member candidates of the open cluster (OC) Stock 2 using data from Gaia DR2. We use the algorithms Clusterix 2.0 and HDBSCAN to select cluster candidates and further refine the final cluster membership by defining neighbors in 5D phase space (X {sub cp}, Y {sub cp}, Z{sub cp},κ⋅μ{sub α}{sup ∗}/ϖ, κ · μ {sub δ}/ϖ). Among these candidates, less than half have G, G {sub BP}, and G {sub RP} extinctions from Gaia. When Gaia extinctions are unavailable, we compute extiction using empirical formulas and E(B − V) = 0.350. We analyze the spatial distribution and mass profile of Stock 2. Our results reveal Stock 2 is still a bound OC and we find evidence of mass segregation. By comparing initial mass functions, the present-day mass function indicates that Stock 2 is a massive stellar cluster with a mass of 4000 M {sub ⊙}. The core radius and tidal radius, calculated via the radial density profile and total mass, are 3.97 pc and 22.65 pc, respectively. Common stars between our selected member candidates and the Large Sky Area Multi-Object Fiber Spectroscopic Telescope DR7 medium-resolution catalog give a metalliclity of [Fe/H] = −0.040 ± 0.147.

47 OTHER INSTRUMENTATION↗