Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “clustering algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Investigation of correlation classification techniques

A two-step classification algorithm for processing multispectral scanner data was developed and tested. The first step is a single pass clustering algorithm that assigns each pixel, based on its spectral signature, to a particular cluster. The output of that step is a cluster tape in which a single integer is associated with each pixel. The cluster tape is used as the input to the second step, where ground truth information is used to classify each cluster using an iterative method of potentials. Once the clusters have been assigned to classes the cluster tape is read pixel-by-pixel and an output tape is produced in which each pixel is assigned to its proper class. In addition to the digital classification programs, a method of using correlation clustering to process multispectral scanner data in real time by means of an interactive color video display is also described.

Haskell, R. E.↗

A Census of Young Stellar Objects in Two Line-of-Sight Star-Forming Regions Toward IRAS 22147+5948 in the Outer Galaxy

Context. Star formation in the outer Galaxy, namely, outside of the Solar circle, has not been extensively studied in part due to the low CO brightness of the molecular clouds linked with the negative metallicity gradient. Recent infrared surveys provide an overview of dust emission in large sections of the Galaxy, but they suffer from cloud confusion and poor spatial resolution at far-infrared wavelengths. Aims. We aim to develop a methodology to identify and classify young stellar objects (YSOs) in star-forming regions in the outer Galaxy and use it to resolve a long-standing disparity in terms of the distance and evolutionary status of IRAS 22147+5948. Methods. We used a support vector machine learning algorithm to complement standard color–color and color–magnitude diagrams in our search for YSOs in the IRAS 22147 region, based on publicly available data from the Spitzer Mapping of the Outer Galaxy survey. The agglomerative hierarchical clustering algorithm was used to identify clusters. Then the physical properties of individual YSOs were calculated. The distances were determined using CO 1–0 from the Five College Radio Astronomy Observatory survey. Results. We identified 13 Class I and 13 Class II YSO candidates using the color–color diagrams, along with an additional 2 and 21 sources, respectively, using the applied machine learning techniques. The spectral energy distributions of 23 sources were modeled with a star and a passive disk, corresponding to Class II objects. The models of three sources include envelopes that are typical for Class I objects. The objects were grouped into two clusters located at a distance of 2:2 kpc and 5 clusters at 5:6 kpc. The spatial extent of CO, radio continuum, and dust emission confirms the origin of YSOs in two distinct star-forming regions along a similar line of sight. Conclusions. The outer Galaxy may serve as a unique laboratory for exploring star formation across environments, on the condition that complementary methods and ancillary data are used to properly account for cloud confusion and distance uncertainties.

Agata Karska↗

Investigation of the application of remote sensing technology to environmental monitoring

Activities and results are reported of a project to investigate the application of remote sensing technology developed for the LACIE, AgRISTARS, Forestry and other NASA remote sensing projects for the environmental monitoring of strip mining, industrial pollution, and acid rain. Following a remote sensing workshop for EPA personnel, the EOD clustering algorithm CLASSY was selected for evaluation by EPA as a possible candidate technology. LANDSAT data acquired for a North Dakota test sight was clustered in order to compare CLASSY with other algorithms.

Rader, M. L.↗

Global Weather States and Their Properties from Passive and Active Satellite Cloud Retrievals

In this study, the authors apply a clustering algorithm to International Satellite Cloud Climatology Project (ISCCP) cloud optical thickness-cloud top pressure histograms in order to derive weather states (WSs) for the global domain. The cloud property distribution within each WS is examined and the geographical variability of each WS is mapped. Once the global WSs are derived, a combination of CloudSat and Cloud-Aerosol Lidar and Infrared Pathfinder Satellite Observations (CALIPSO) vertical cloud structure retrievals is used to derive the vertical distribution of the cloud field within each WS. Finally, the dynamic environment and the radiative signature of the WSs are derived and their variability is examined. The cluster analysis produces a comprehensive description of global atmospheric conditions through the derivation of 11 WSs, each representing a distinct cloud structure characterized by the horizontal distribution of cloud optical depth and cloud top pressure. Matching those distinct WSs with cloud vertical profiles derived from CloudSat and CALIPSO retrievals shows that the ISCCP WSs exhibit unique distributions of vertical layering that correspond well to the horizontal structure of cloud properties. Matching the derived WSs with vertical velocity measurements shows a normal progression in dynamic regime when moving from the most convective to the least convective WS. Time trend analysis of the WSs shows a sharp increase of the fair-weather WS in the 1990s and a flattening of that increase in the 2000s. The fact that the fair-weather WS is the one with the lowest cloud radiative cooling capability implies that this behavior has contributed excess radiative warming to the global radiative budget during the 1990s.

histograms↗

Serial crystallography with multi-stage merging of thousands of images

KAMO and BLEND provide particularly effective tools to automatically manage the merging of large numbers of data sets from serial crystallography. The requirement for manual intervention in the process can be reduced by extending BLEND to support additional clustering options such as the use of more accurate cell distance metrics and the use of reflection-intensity correlation coefficients to infer `distances' among sets of reflections. This increases the sensitivity to differences in unit-cell parameters and allows clustering to assemble nearly complete data sets on the basis of intensity or amplitude differences. If the data sets are already sufficiently complete to permit it, one applies KAMO once and clusters the data using intensities only. When starting from incomplete data sets, one applies KAMO twice, first using unit-cell parameters. In this step, either the simple cell vector distance of the original BLEND or the more sensitive NCDist is used. This step tends to find clusters of sufficient size such that, when merged, each cluster is sufficiently complete to allow reflection intensities or amplitudes to be compared. One then uses KAMO again using the correlation between reflections with a common hkl to merge clusters in a way that is sensitive to structural differences that may not have perturbed the unit-cell parameters sufficiently to make meaningful clusters. Many groups have developed effective clustering algorithms that use a measurable physical parameter from each diffraction still or wedge to cluster the data into categories which then can be merged, one hopes, to yield the electron density from a single protein form. Since these physical parameters are often largely independent of one another, it should be possible to greatly improve the efficacy of data-clustering software by using a multi-stage partitioning strategy. Here, one possible approach to multi-stage data clustering is demonstrated. The strategy is to use unit-cell clustering until the merged data are sufficiently complete and then to use intensity-based clustering. Using this strategy, it is demonstrated that it is possible to accurately cluster data sets from crystals that have subtle differences.

36 MATERIALS SCIENCE↗

Optical processing of imaging spectrometer data

The data-processing problems associated with imaging spectrometer data are reviewed; new algorithms and optical processing solutions are advanced for this computationally intensive application. Optical decision net, directed graph, and neural net solutions are considered. Decision nets and mineral element determination of nonmixture data are emphasized here. A new Fisher/minimum-variance clustering algorithm is advanced, initialization using minimum-variance clustering is found to be preferred and fast. Tests on a 500-class problem show the excellent performance of this algorithm.

Liu, Shiaw-Dong↗

Mitigate: An Adaptive Network Data Anonymization Tool Using Condensation-Based Differential Privacy

Modern network devices collect a large amount of data that can be analyzed to identify bottlenecks, anomalies, cyber-attacks, etc. Therefore, there is often a need to analyze such collections of network data quite often by an external expert or by the research community. However, these collections of data contain sensitive, proprietary information. In order for the network data to be shared, it must first be anonymized. The overall objective of this project is to develop an innovative privacy management tool to anonymize network data and achieve sufficient privacy, acceptable data utility, and efficient data analysis at the same time. No existing anonymization methods can achieve all of these at the same time. The core of this technology is a differential private clustering algorithm that provides strong privacy protection, preserves data properties important for subsequent analysis, and allows the party receiving the anonymized data to conduct analysis directly on anonymized data without the need of decryption or any extra processing. The research carried out was to design, implement and verify a solution to this problem by completing the following tasks: 1) developing the core technology; 2) developing a context based method that automatically recommends fields that must be anonymized; 3) conducted experiments showing superior results using our approach compared to existing tools, and 4) developed an intuitive but basic user interface. The research that was conducted generated novel algorithmic techniques that utilize state-of-the-art methods such as condensation, differential privacy preservation, clustering, automated tuning based on contextual awareness, and recommendation techniques to specify columns to users for anonymization leading to optimal privacy that allows research analysis on the dataset. Experiments were conducted to evaluate the efficacy of these novel algorithmic techniques by performing analysis on original non-anonymized datasets, then conducting analysis on the same yet anonymized datasets and comparing the results of the analyses. Overall, the anonymized analysis results were within 1% of the original results, verifying that the generated technology not only guarantees a high level of privacy but also enables research analysis as if it were conducted on the original dataset. Potential applications of this technology include anonymization of any type of structured network datasets that contain sensitive identifiers, such as IP addresses, that can be used in multiple applications. For example, to create an AI or machine learning model for cyber security, e.g., to detect attacks, or for performance analysis, e.g., identify bottlenecks or predict performance. In addition, a market analysis that was conducted for potential applications of this technology identified a broader range of applications of our anonymization technology beyond the network sector that includes healthcare, banking, insurance, securities, finance (FISB), data brokering, cloud services, ad sales, and government.

97 MATHEMATICS AND COMPUTING↗

Comparing Synoptic Pattern Evolution for Flash‐Flood‐Producing and Non‐Flash‐Flood‐Producing Mesoscale Convective Systems in the United States

Understanding how the short-term evolution of synoptic weather patterns influence Mesoscale Convective Systems (MCSs) is essential, as these systems are responsible for over half of central U.S. flash floods, leading to substantial socioeconomic and water resource management impacts. This study analyzes long-term MCS data, flash flood reports, and atmospheric reanalyses from 2007 to 2017 using a machine learning clustering algorithm to examine how the synoptic weather patterns evolve prior to MCS initiation. While the clusters reflect seasonal and regional differences in MCS occurrence, they do not consistently distinguish between MCSs that do and do not produce flash floods. Systems in the southern Great Plains are more flood-prone when a synoptic-scale forcing, located near the system, drives strong water vapor transport from the nearby moisture source. More generally under different synoptic weather patterns, a broader precipitating area is the most dominant factor governing MCS flash flood potential.

atmospheric dynamics↗

Finding the earth in ERTS

Itek's Optical System Division recently completed an investigation funded by NASA to develop interpretation methods and algorithms suitable for recognition of earth resources by machines using multispectral data from ERTS. Through the algorithms developed (and described here) it is now possible to automatically recognize terrain types. The clustering algorithm guarantees high accuracy in the recognition process with almost complete automation. Interestingly, the machine recognition seems to be more accurate than a human photointerpreter who has been restricted to using only ERTS-1 color composites. That is, machine recognition appears to be more sensitive, it can operate much closer, to the resolution limit of the ERTS-1 imagery than the human photointerpreter.

Gramenopoulos, N.↗

Intelligent Experiments through Real-Time AI: Fast Data Processing and Autonomous Detector Control for High-Energy Nuclear Experiments

The aim of this project is to develop software and hardware for fast real-time data processing and autonomous detector control and calibration for the sPHENIX and the future EIC experiments. Below summarizes Georgia Tech team efforts in the past year: 1. We developed a real-time clustering algorithm and FPGA-based pipeline architecture for processing fired pixel data from ALPIDE sensors in sPHENIX experiments. Our Columnar Clustering Co-Design introduces a hardware-aware, stream-friendly approach that segments pixel data by column pairs using a Column Pair Clustering (CPC) strategy, followed by Cluster Stitching to merge adjacent subclusters. Implemented in Vitis HLS, the pipeline comprises five stages—read-in, subclustering, stitching, analysis, and write-out—connected by tagged HLS streams with custom end-of-event signaling for robust synchronization. We designed a pipelined dataflow model optimized for throughput, low latency, and minimal buffering, enabling scalable clustering across events of arbitrary size. Our system maintains spatial precision via center-of-mass and shape key extraction and efficiently handles edge cases such as fragmented or nested clusters. Compared against DBSCAN in both software and hardware, our approach demonstrates competitive performance under FPGA constraints. 2. We also conducted a comprehensive algorithm-to-hardware co-design of connected component analysis tailored for sPHENIX experiments, focusing on real-time, low-latency processing using FPGAs and High-Level Synthesis (HLS). Starting from a Python-based particle tracking pipeline, the team translated the core logic—graph traversal via DFS and Union-Find—into an HLS-compatible C++ model, replacing dynamic memory and recursion with static arrays and pipelined control flow. The final design includes a fully streamed and dataflow-compatible Union-Find kernel optimized across five iterations, incorporating loop pipelining, array partitioning, AXI/FIFO interface tuning, and function flattening. Experimental results show up to 14.8× speedup over the CPU baseline, reducing per-graph latency to 1.58 μs and demonstrating strong resource efficiency with only ~7k LUTs and zero BRAM usage. The design maintains functional correctness against the Python reference using a Python-based C-simulation framework and Mean Squared Error metrics. This work validates the potential of HLS-driven FPGA designs for edge-level HEP data acquisition, laying a scalable foundation for future integration with real-time detector pipelines and multi-graph processing systems.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

On evaluating clustering procedures for use in classification

The problem of evaluating clustering algorithms and their respective computer programs for use in a preprocessing step for classification is addressed. In clustering for classification the probability of correct classification is suggested as the ultimate measure of accuracy on training data. A means of implementing this criterion and a measure of cluster purity are discussed. Examples are given. A procedure for cluster labeling that is based on cluster purity and sample size is presented.

Pore, M. D.↗

Visualizing Time-Varying Distribution Data in EOS Application

In this research, we have developed several novel visualization methods for spatial probability density function data. Our focus has been on 2D spatial datasets, where each pixel is a random variable, and has multiple samples which are the results of experiments on that random variable. We developed novel clustering algorithms as a means to reduce the information contained in these datasets; and investigated different ways of interpreting and clustering the data.

Shen, Han-Wei↗

Unsupervised learning for identifying events in active target experiments

This article presents novel applications of unsupervised machine learning methods to the problem of event separation in an active target detector, the Active-Target Time Projection Chamber (AT-TPC). The overarching goal is to group similar events in the early stages of the data analysis, thereby improving efficiency by limiting the computationally expensive processing of unnecessary events. The application of unsupervised clustering algorithms to the analysis of two-dimensional projections of particle tracks from a resonant proton scattering experiment on 46 Ar is introduced. We explore the performance of autoencoder neural networks and a pre-trained VGG16 Simonyan and Zisserman (2015) convolutional neural network. We study clustering performance on both data from a simulated 46 Ar experiment, and real events from the AT-TPC detector. We find that a -means algorithm applied to simulated data in the VGG16 latent space forms almost perfect clusters. Additionally, the VGG16+-means approach finds high purity clusters of proton events for real experimental data. Here, we also explore the application of clustering the latent space of autoencoder neural networks for event separation. While these networks show strong performance, they suffer from high variability in their results.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Orbit Clustering Based on Transfer Cost

We propose using cluster analysis to perform quick screening for combinatorial global optimization problems. The key missing component currently preventing cluster analysis from use in this context is the lack of a useable metric function that defines the cost to transfer between two orbits. We study several proposed metrics and clustering algorithms, including k-means and the expectation maximization algorithm. We also show that proven heuristic methods such as the Q-law can be modified to work with cluster analysis.

combinatorial optimization↗

Improving and Expanding NASA Software Cost Estimation Methods

Estimators and analysts are increasingly being tasked to develop better models and reliable cost estimates in support of program planning and execution. While there has been extensive work on improving parametric methods for cost estimation, there is very little focus on the use of cost models based on analogy and clustering algorithms. In this paper we summarize the results of our research in developing an analogy method for estimating NASA spacecraft flight software using spectral clustering on system characteristics (symbolic nonnumerical data) and evaluate its performance by comparing it to a number of the most commonly used estimation methods. The strengths and weaknesses of each method based on their performance are also discussed. The paper concludes with an overview of the analogy estimation tool (ASCoT) developed for use within NASA that implements the recommended analogy algorithm.

Hihn, Jairus↗

A Decentralized Approach for Modeling Organized Convection Based on Thermal Populations on Microgrids

Abstract In this study, a spectral model for convective transport is coupled to a thermal population model on a two‐dimensional horizontal “microgrid,” covering the typical gridbox size of general circulation models. The goal is to explore new ways of representing impacts of spatial organization in cumulus cloud fields. The thermals are considered the smallest building block of convection, with thermal life cycle and movement represented through binomial functions. Thermals interact through two simple rules, reflecting pulsating growth and environmental deformation. Long‐lived thermal clusters thus form on the microgrid, exhibiting scale growth and spacing that represent simple forms of spatial organization and memory. Size distributions of cluster number are diagnosed from the microgrid through an online clustering algorithm, and provided as input to a spectral multiplume eddy‐diffusivity mass flux scheme. This yields a decentralized transport system, in that the thermal clusters acting as independent but interacting nodes that carry information about spatial structure. The main objectives of this study are (a) to seek proof of concept of this approach, and (b) to gain insight into impacts of spatial organization on convective transport. Single‐column model experiments demonstrate satisfactory skill in reproducing two observed cases of continental shallow convection. Metrics expressing self‐organization and spatial organization match well with large‐eddy simulation results. We find that in this coupled system, spatial organization impacts convective transport primarily through the scale break in the size distribution of cluster number. The rooting of saturated plumes in the subcloud mixed layer plays a key role in this process.

54 ENVIRONMENTAL SCIENCES↗

Scalable Tensor Methods for Nonuniform Hypergraphs

While multilinear algebra appears natural for studying the multiway interactions modeled by hypergraphs, tensor methods for general hypergraphs have been stymied by theoretical and practical barriers. A recently proposed adjacency tensor is applicable to nonuniform hypergraphs, but is prohibitively costly to form and analyze in practice. We develop tensor times same vector (TTSV) algorithms for this tensor which improve complexity from $O(n^r)$ to a low-degree polynomial in $r$, where $n$ is the number of vertices and $r$ is the maximum hyperedge size. Our algorithms are implicit, avoiding formation of the order $r$ adjacency tensor. Here, we demonstrate the flexibility and utility of our approach in practice by developing tensor-based hypergraph centrality and clustering algorithms. We also show these tensor measures offer complementary information to analogous graph-reduction approaches on data, and are also able to detect higher-order structure that many existing matrix-based approaches provably cannot.

97 MATHEMATICS AND COMPUTING↗

New output improvements for CLASSY

Additional output data and formats for the CLASSY clustering algorithm were developed. Four such aids to the CLASSY user are described. These are: (1) statistical measures; (2) special map types; (3) formats for standard output; and (4) special cluster display method.

Rassbach, M. E.↗