Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “clustering algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Cluster Analysis of IRIS Spectroscopic Line Profiles and SDO/AIA EUV Emission in Observations and RMHD Simulations of the Solar Atmosphere

Spatially-resolved observations from the IRIS and SDO/AIA satellites, especially when coupled with realistic 3D RMHD simulations, are a powerful tool for analysis of processes in the solar chromosphere, transition region, and corona. However, the complexity of the data makes understanding the observations and modeling results difficult. In this work, we apply unsupervised clustering algorithms for analysis of observational and synthetic chromospheric Mg II h&k 2796Å&2803Å and transition region C II 1334Å&1335Å line profiles observed by IRIS, and extreme ultraviolet (EUV) emission observed by SDO/AIA, for various types of problems. The synthetic line profiles are computed for simulations of the quiescent solar atmosphere (using the StellarBox and RH1.5 codes). The K-Means clustering algorithm is applied, and the selection of an optimal number of clusters is supported by the average silhouette width technique. We discuss applications of the line profile clustering method to 1) visualization of computational and observational spectroscopic imaging data; 2) understanding of evolutionary trends and behavior patterns of quiet Sun emission and during solar flares; and 3) recognition of heating events and shock waves.

Sadykov, Viacheslav↗

Performance Analysis of Data Processing in Distributed File Systems with Near Data Processing

In the era of big data, the escalating volume and velocity of data generation pose significant challenges in data processing. Traditional systems like Spark and Hadoop manage the increasing amount and velocity of data by improving data placement and processing speeds. However, they face inherent limitations due to the essential data movement required for processing. In this paper, we explore the Skyhook framework, a novel extension of the Ceph distributed system, which significantly reduces the need for data movement. We present an extensive case study using the Skyhook framework, applying it with the TPC-H and K-means clustering algorithms. More specifically, we leverage the TPC-H benchmark to distinguish between CPU-intensive and I/O-intensive tasks. We explore the integration of K-means clustering into SQL, coupled with a near-data processing system to offload the computational burden of the K-means clustering algorithm to storage nodes. We conduct a comprehensive performance evaluation of distributed data processing applications across three processing approaches: traditional layout (baseline), optimized layout, and near-data processing. Additionally, we introduce the use of the FIO tool to simulate real-world system workloads, enabling the measurement of performance metrics such as average latency and CPU utilization. Our research is a significant advance in understanding how to optimize data processing systems to meet the demands of the modern data landscape.

Hou, Shiyue↗

Multipath Mitigation via Clustering for Position Estimation Refinement in Urban Environments

Position estimation using global navigation satellite systems (GNSS) suffers from poor accuracy within urban canyons due to significant signal disruption caused by tall buildings. This issue can be attributed to the GNSS signals reflecting off buildings resulting in severe multipath reflections which degrade the receiver's performance. In this paper, we introduce an innovative approach to filter GNSS satellite measurements to improve the accuracy of the estimated position by leveraging a clustering algorithm. This approach utilizes a predictive GNSS availability service to filter out non-line-of-sight measurements. Then, a subset of line-of-sight satellite measurement combinations are evaluated using a clustering algorithm. When combined, results show these techniques can reduce the mean horizontal error measured in an urban canyon by nearly an order of magnitude, from ~ 18 meters to ~ 2 meters when using a single point positioning solver.

GPS↗

Multipath Mitigation via Clustering for Position Estimation Refinement in Urban Environments

Position estimation using global navigation satellite systems (GNSS) suffers from poor accuracy within urban canyons due to significant signal disruption caused by tall buildings. This issue can be attributed to the GNSS signals reflecting off buildings resulting in severe multipath reflections which degrade the receiver's performance. In this paper, we introduce an innovative approach to filter GNSS satellite measurements to improve the accuracy of the estimated position by leveraging a clustering algorithm. This approach utilizes a predictive GNSS availability service to filter out non-line-of-sight measurements. Then, a subset of line-of-sight satellite measurement combinations are evaluated using a clustering algorithm. When combined, results show these techniques can reduce the mean horizontal error measured in an urban canyon by nearly an order of magnitude, from ~ 18 meters to ~ 2 meters when using a single point positioning solver.

GPS↗

X-ray nano-imaging of defects in thin film catalysts via cluster analysis

Functional properties of transition-metal oxides strongly depend on crystallographic defects; crystallographic lattice deviations can affect ionic diffusion and adsorbate binding energies. Scanning x-ray nanodiffraction enables imaging of local structural distortions across an extended spatial region of thin samples. Yet, localized lattice distortions remain challenging to detect and localize using nanodiffraction, due to their weak diffuse scattering. Here, in this study, we apply an unsupervised machine learning clustering algorithm to isolate the low-intensity diffuse scattering in as-grown and alkaline-treated thin epitaxially strained SrIrO 3 films. We pinpoint the defect locations, find additional strain variation in the morphology of electrochemically cycled SrIrO 3 , and interpret the defect type by analyzing the diffraction profile through clustering. Our findings demonstrate the use of a machine learning clustering algorithm for identifying and characterizing hard-to-find crystallographic defects in thin films of electrocatalysts and highlight the potential to study electrochemical reactions at defect sites in operando experiments.

42 ENGINEERING↗

Computer-aided analysis of Landsat-1 MSS data - A comparison of three approaches, including a 'modified clustering' approach

Three approaches for analyzing Landsat-1 data from Ludwig Mountain in the San Juan Mountain range in Colorado are considered. In the 'supervised' approach the analyst selects areas of known spectral cover types and specifies these to the computer as training fields. Statistics are obtained for each cover type category and the data are classified. Such classifications are called 'supervised' because the analyst has defined specific areas of known cover types. The second approach uses a clustering algorithm which divides the entire training area into a number of spectrally distinct classes. Because the analyst need not define particular portions of the data for use but has only to specify the number of spectral classes into which the data is to be divided, this classification is called 'nonsupervised'. A hybrid method which selects training areas of known cover type but then uses the clustering algorithm to refine the data into a number of unimodal spectral classes is called the 'modified-supervised' approach.

Fleming, M. D.↗

Two generalizations of Kohonen clustering

The relationship between the sequential hard c-means (SHCM), learning vector quantization (LVQ), and fuzzy c-means (FCM) clustering algorithms is discussed. LVQ and SHCM suffer from several major problems. For example, they depend heavily on initialization. If the initial values of the cluster centers are outside the convex hull of the input data, such algorithms, even if they terminate, may not produce meaningful results in terms of prototypes for cluster representation. This is due in part to the fact that they update only the winning prototype for every input vector. The impact and interaction of these two families with Kohonen's self-organizing feature mapping (SOFM), which is not a clustering method, but which often leads ideas to clustering algorithms is discussed. Then two generalizations of LVQ that are explicitly designed as clustering algorithms are presented; these algorithms are referred to as generalized LVQ = GLVQ; and fuzzy LVQ = FLVQ. Learning rules are derived to optimize an objective function whose goal is to produce 'good clusters'. GLVQ/FLVQ (may) update every node in the clustering net for each input vector. Neither GLVQ nor FLVQ depends upon a choice for the update neighborhood or learning rate distribution - these are taken care of automatically. Segmentation of a gray tone image is used as a typical application of these algorithms to illustrate the performance of GLVQ/FLVQ.

Bezdek, James C.↗

Clustering Days with Similar Airport Weather Conditions

On any given day, traffic flow managers must often rely on past experience and intuition when developing traffic flow management initiatives that mitigate imbalances between the aircraft demand and the weather impacted airport capacity. The goal of this study was to build on recent efforts to apply data mining classification and clustering algorithms to vast archives of historical weather and air traffic data to identify patterns and past decisions that can ultimately inform day-of-operations decision-making. More specifically, this study identified similar weather impacted days at select U.S. airports, and analyzed the traffic management initiatives implemented on these representative days. The identification of the similar days was accomplished by applying a decision tree algorithm to the hourly Localized Aviation Model Output Statistics Program observations and the arrival delays for Newark Liberty International Airport. The branches from the trained decision tree were subsequently pruned to identify four weather conditions that resulted in medium to high delays for the arrivals scheduled to Newark in 2012. Using these weather conditions, four, daily airport-level Weather Impacted Traffic Index values were calculated using the Localized Aviation Model Output Statistics Program observations and the 2012 scheduled arrival counts from the FAAs Aviation System Performance Metric system. The four, daily Weather Impacted Traffic Index values for 2012 were subsequently clustered using an Expectation Maximization clustering algorithm, and nine unique types of weather days at Newark were identified. By far the most prominent type of day at Newark was a day associated with relatively good weather conditions, where there was little convective activity, winds were low, ceilings and visibility were high and there was little precipitation. Moderate levels of convective activity characterized the next most prominent type of day. Days with persistently high winds or low ceiling and visibility levels were relatively rare in 2012. Lastly, the frequency at which Ground Delay Programs, Ground Stops and Miles-in-Trail restrictions were implemented on each of the typical types of days at Newark were analyzed. Based on the results, it does appear as if the usage of Miles-in-Trail, Ground Delay Program and Ground Stop restrictions correlates well with the severity of the weather associated with each unique type of weather impacted day at Newark. Furthermore, the results demonstrate that it is feasible to use historical weather and air traffic archives to provide guidance on the types of traffic management restrictions to implement in response to the weather conditions impacting an airport.

traffic flow management↗

Clustering Days with Similar Airport Weather Conditions

On any given day, traffic flow managers must often rely on past experience and intuition when developing traffic flow management initiatives that mitigate imbalances between the aircraft demand and the weather impacted airport capacity. The goal of this study was to build on recent efforts to apply data mining classification and clustering algorithms to vast archives of historical weather and air traffic data to identify patterns and past decisions that can ultimately inform day-of-operations decision-making. More specifically, this study identified similar weather impacted days at select U.S. airports, and analyzed the traffic management initiatives implemented on these representative days. The identification of the similar days was accomplished by applying a decision tree algorithm to the hourly Localized Aviation Model Output Statistics Program observations and the arrival delays for Newark Liberty International Airport. The branches from the trained decision tree were subsequently pruned to identify four weather conditions that resulted in medium to high delays for the arrivals scheduled to Newark in 2012. Using these weather conditions, four, daily airport-level Weather Impacted Traffic Index values were calculated using the Localized Aviation Model Output Statistics Program observations and the 2012 scheduled arrival counts from the FAAs Aviation System Performance Metric system. The four, daily Weather Impacted Traffic Index values for 2012 were subsequently clustered using an Expectation Maximization clustering algorithm, and nine unique types of weather days at Newark were identified. By far the most prominent type of day at Newark was a day associated with relatively good weather conditions, where there was little convective activity, winds were low, ceilings and visibility were high and there was little precipitation. Moderate levels of convective activity characterized the next most prominent type of day. Days with persistently high winds or low ceiling and visibility levels were relatively rare in 2012. Lastly, the frequency at which Ground Delay Programs, Ground Stops and Miles-in-Trail restrictions were implemented on each of the typical types of days at Newark were analyzed. Based on the results, it does appear as if the usage of Miles-in-Trail, Ground Delay Program and Ground Stop restrictions correlates well with the severity of the weather associated with each unique type of weather impacted day at Newark. Furthermore, the results demonstrate that it is feasible to use historical weather and air traffic archives to provide guidance on the types of traffic management restrictions to implement in response to the weather conditions impacting an airport.

weather↗

Cloud-based Testbed for Adaptive Under-Frequency Load Shedding with High DER Penetration

Increasing penetration of distributed energy resources and behind-the-meter renewables may soon disrupt the efficacy of critical protection schemes, such as under-frequency load shedding (UFLS). Improved data exchange and coordination across the transmission-distribution boundary will be required to maintain reliability of bulk electric system. Standards-based data integration platforms using agreed-upon semantic vocabularies, such as the Common Information Model, will be key to enabling adaptive protection schemes requiring synthesized data from both the bulk power system and behind-the-meter resources. This paper introduces a cloud-based open-source data integration environment and UFLS clustering algorithm being developed to enable adaptive relay coordination between transmission and distribution utilities in the state of Vermont.

Anderson, Alexander A.↗

An Efficient, FPGA-Based, Cluster Detection Algorithm Implementation for a Strip Detector Readout System in a Time Projection Chamber Polarimeter

A fundamental challenge in a spaceborne application of a gas-based Time Projection Chamber (TPC) for observation of X-ray polarization is handling the large amount of data collected. The TPC polarimeter described uses the APV-25 Application Specific Integrated Circuit (ASIC) to readout a strip detector. Two dimensional photoelectron track images are created with a time projection technique and used to determine the polarization of the incident X-rays. The detector produces a 128x30 pixel image per photon interaction with each pixel registering 12 bits of collected charge. This creates challenging requirements for data storage and downlink bandwidth with only a modest incidence of photons and can have a significant impact on the overall mission cost. An approach is described for locating and isolating the photoelectron track within the detector image, yielding a much smaller data product, typically between 8x8 pixels and 20x20 pixels. This approach is implemented using a Microsemi RT-ProASIC3-3000 Field-Programmable Gate Array (FPGA), clocked at 20 MHz and utilizing 10.7k logic gates (14% of FPGA), 20 Block RAMs (17% of FPGA), and no external RAM. Results will be presented, demonstrating successful photoelectron track cluster detection with minimal impact to detector dead-time.

Polarimeter↗

In Situ Transmission Electron Microscopy of High-Temperature Inconel-625 Corrosion by Molten Chloride Salts

This paper describes an approach to monitor high temperature molten chloride (MgCl 2 -NaCl-KCl) salt corrosion of Inconel-625 alloy in real time at high spatial resolution. The approach is based on a micro-environmental-cell assembly integrated into a transmission-electron-microscope goniometer to examine in situ the salt-alloy interface during corrosion, employing real time electron diffraction and imaging. It establishes procedures to minimize incorporation of H 2 O or O 2 from atmosphere in the chloride salts during sample fabrication and corrosion, which is critical to understanding the fundamental corrosion mechanisms. A clustering algorithm and a 2D Gaussian fit function are used to determine diffraction spot intensities in in situ diffraction patterns, to quantify alloy corrosion. This facilitates quantitative observation of the evolution of individual grains, in contrast to conventional macroscopic corrosion rate quantification. The isothermal corrosion rate of Inconel-625 in an anhydrous, unoxidized salt-stack is 220 ± 30 μm year -1 at 700 °C and 350 ± 20 μm year -1 at 800 °C. However, the corrosion rate at 700 °C increases five-fold to 1000 ± 170 μm year -1 when the salt stack is air-exposed, indicating the dominant effects of hydrated or oxidized impurities on corrosion acceleration. Furthermore, real time imaging of the microstructure evolution suggests that corrosion is initiated at grain boundaries.

14 SOLAR ENERGY↗

Development of Multimodal Few-Shot Analytics for Electron Micrographs

Recent advances in materials data analytics have provided new avenues for determining process-structure-property (PSP) linkages in a variety of materials. Machine learning techniques including few-shot learning have increased the efficiency of classifying microscopy images for the purposes of material characterization. Attempts at creating a multimodal approach can provide further improvements to current models and help extract more salient features from data. In this vein, raw spectrum data was taken to provide an additional modality to our current pyCHIP classifier. Modifications in segmentation also show potential in improving the accuracy of the pyCHIP classifier. Classifier output was analyzed using network graphs and unsupervised clustering algorithms such as spectral clustering to detect better segmentation methods than the current “chipping” approach. We suggest that the chip selection process can be automated in the future using a combination of these techniques to enable high-throughput analyses.

36 MATERIALS SCIENCE↗

ICAP - An Interactive Cluster Analysis Procedure for analyzing remotely sensed data

An Interactive Cluster Analysis Procedure (ICAP) was developed to derive classifier training statistics from remotely sensed data. ICAP differs from conventional clustering algorithms by allowing the analyst to optimize the cluster configuration by inspection, rather than by manipulating process parameters. Control of the clustering process alternates between the algorithm, which creates new centroids and forms clusters, and the analyst, who can evaluate and elect to modify the cluster structure. Clusters can be deleted, or lumped together pairwise, or new centroids can be added. A summary of the cluster statistics can be requested to facilitate cluster manipulation. The principal advantage of this approach is that it allows prior information (when available) to be used directly in the analysis, since the analyst interacts with ICAP in a straightforward manner, using basic terms with which he is more likely to be familiar. Results from testing ICAP showed that an informed use of ICAP can improve classification, as compared to an existing cluster analysis procedure.

Wharton, S. W.↗

Data-Driven Performance Optimization of Gamma Spectrometers With Many Channels

In gamma spectrometers with variable spectroscopic performance across many channels (e.g., many pixels or voxels), a tradeoff exists between including data from successively worse-performing readout channels and increasing efficiency. Brute-force calculation of the optimal set of included channels is exponentially infeasible as the number of channels grows, and approximate methods are required. In this work, we present a data-driven framework for attempting to find near-optimal sets of included detector channels. The framework leverages non-negative matrix factorization (NMF) to learn the behavior of gamma spectra across the detector and clusters similarly-performing detector channels together. Performance comparisons are then made between spectra with channel clusters removed, which is more feasible than brute force. The framework is general and can be applied to arbitrary, user-defined performance metrics depending on the application. We apply this framework to optimizing gamma spectra measured by H3D M400 CdZnTe (CZT) spectrometers, which exhibit variable performance across their crystal volumes. In particular, we show several examples optimizing various performance metrics for uranium and plutonium gamma spectra in non-destructive assay (NDA) for nuclear safeguards, and explore trends in performance versus parameters such as clustering algorithm type. We also compare the NMF + clustering pipeline to several non-machine-learning (ML) algorithms, including several greedy algorithms. Although, we find that the NMF + clustering pipeline tends to find the best-performing set of detector voxels, significantly improving over the unoptimized spectra, but that a greedy accumulation of spectra segmented by detector depth can, in some cases, give similar performance improvements in much less computation time.

Energy resolution↗

Leveraging the digital thread for physics-based prediction of microstructure heterogeneity in additively manufactured parts

A major limitation of additive manufacturing (AM) processes is that local conditions of material deposition frequently lead to unintentional heterogeneities in microstructure and properties within a single component, despite nominally uniform process conditions. Up to now, there has been no way to a priori determine the distribution of these heterogeneities, requiring expensive trial-and-error approaches to fabrication, testing, and characterization. Here, a physics-based framework for creating a digital representation of the laser powder bed fusion (PBF) process is proposed to predict the variation in solidification behavior that leads to heterogeneous microstructures in an as-built part. By leveraging in situ process data stored in the part’s digital thread, the scan path and process parameters were input into a heat transfer model which predicted solidification data at the melt pool scale. A two-step unsupervised clustering algorithm was used to first cluster the local solidification conditions (12.5µm 3 voxels) and then to cluster the regional behavior on the scale of multiple scan passes and print layers (250µm 3 super-voxels). This process was used to identify regions with similar solidification characteristics for multiple locations in a Stainless Steel 316-L component. The corresponding as-built part was sectioned and characterized using electron backscatter diffraction (EBSD). Quantitative analysis of the pole figures confirmed that the predicted regions of heterogeneity in the solidification conditions corresponded with differences in the observed microstructure. In conclusion, this work shows a viable path for estimating the microstructural heterogeneity for additively manufactured parts to either limit microstructural variation throughout a part or to enable functionality-based variation of the microstructure.

36 MATERIALS SCIENCE↗

ArborX 2.0

ArborX library tackles a problem of efficiently finding geometric objects that are close in space. Variations of this problem, such as finding the nearest neighbors of a point, or finding all objects within a certain distance, are inherent components of applications in many fields. The data may be large so that solving the problem efficiently may require significant computational resources, such as multiple processors or accelerators such as general purpose GPUs. ArborX' main advantage in its ability to solve large problems efficiently utilizing a combination of distributed and on-node parallelism. ArborX can be run efficiently on a wide variety of hardware, including GPUs from different vendors, which distinguishes it from other available libraries which typically choose only few of these. The other advantage is that it supports both types of user problems: spatial problems (useful for intersections and finding objects within certain distance), and nearest neighbor problems. ArborX also supports flexible interface in its interaction with a user. Particularly, it allows a user to call user's own function on a positive match, a functionality not rarely available in other libraries. ArborX implements construction and traversal algorithms using efficient tree structures, such as bounding volume hierarchy (BVH). At its core, ArborX uses linear BVH for its low construction cost and sufficient quality. ArborX implements both spatial and nearest-neighbor traversal algorithms. ArborX also provides several clustering algorithms (minimum spanning tree, DBSCAN, HDBSCAN*), interpolation using minimum least squares and ray tracing. ArborX is written using C++, and is parallelized using the message passing interface (MPI) for the distributed communication, and the Kokkos library for on-node parallelism. This approach allows ArborX to be run on a wide variety of hardware, from common laptops and desktops to supercomputers while using the same codebase.

Prokopenko, Andrey [Oak Ridge National Laboratory ↗

Constrained spectral clustering under a local proximity structure assumption

This work focuses on incorporating pairwise constraints into a spectral clustering algorithm. A new constrained spectral clustering method is proposed, as well as an active constraint acquisition technique and a heuristic for parameter selection. We demonstrate that our constrained spectral clustering method, CSC, works well when the data exhibits what we term local proximity structure.

domain knowledge↗