Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “clustering algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

An unsupervised classification approach for analysis of Landsat data to monitor land reclamation in Belmont county, Ohio

Two unsupervised classification procedures for analyzing Landsat data used to monitor land reclamation in a surface mining area in east central Ohio are compared for agreement with data collected from the corresponding locations on the ground. One procedure is based on a traditional unsupervised-clustering/maximum-likelihood algorithm sequence that assumes spectral groupings in the Landsat data in n-dimensional space; the other is based on a nontraditional unsupervised-clustering/canonical-transformation/clustering algorithm sequence that not only assumes spectral groupings in n-dimensional space but also includes an additional feature-extraction technique. It is found that the nontraditional procedure provides an appreciable improvement in spectral groupings and apparently increases the level of accuracy in the classification of land cover categories.

Brumfield, J. O.↗

Performance tests of signature extension algorithms

Comparative tests were performed on seven signature extension algorithms to evaluate their effectiveness in correcting for changes in atmospheric haze and sun angle in a LANDSAT scene. Four of the algorithms were cluster matching, and two were maximum likelihood algorithms. The seventh algorithm determined the haze level in both training and recognition segments and used a set of tables calculated from an atmospheric model to determine the affine transformation that corrects the training signatures for changes in sun angle and haze level. Three of the algorithms were tested on a simulated data set, and all of the algorithms were tested on consecutive-day data.

Abotteen, R. A.↗

Performance tests of signature extension algorithms

Comparative tests were performed on seven signature extension algorithms to evaluate their effectiveness in correcting for changes in atmospheric haze and sun angle in a Landsat scene. Four of the algorithms were cluster matching, and two were maximum likelihood algorithms. The seventh algorithm determined the haze level in both training and recognition segments and used a set of tables calculated from an atmospheric model to determine the affine transformation that corrects the training signatures for changes in sun angle and haze level. Three of the algorithms were tested on a simulated data set, and all of the algorithms were tested on consecutive-day data. The classification performance on the data sets using the algorithms is presented, along with results of statistical tests on the accuracy and proportion estimates. The three algorithms tested on the simulated data produced significant improvements over the results obtained using untransformed signatures. For the consecutive-day data, the tested algorithms produced improvements in most but not all cases. The tests indicated also that no statistically significant differences were noted among the algorithms.

Abotteen, R.↗

Unsupervised learning-enabled pulsed infrared thermographic microscopy of subsurface defects in stainless steel

Metallic structures produced with laser powder bed fusion (LPBF) additive manufacturing method (AM) frequently contain microscopic porosity defects, with typical approximate size distribution from one to 100 microns. Presence of such defects could lead to premature failure of the structure. In principle, structural integrity assessment of LPBF metals can be accomplished with nondestructive evaluation (NDE). Pulsed infrared thermography (PIT) is a non-contact, one-sided NDE method that allows for imaging of internal defects in arbitrary size and shape metallic structures using heat transfer. PIT imaging is performed using compact instrumentation consisting of a flash lamp for deposition of a heat pulse, and a fast frame infrared (IR) camera for measuring surface temperature transients. However, limitations of imaging resolution with PIT include blurring due to heat diffusion, sensitivity limit of the IR camera. We demonstrate enhancement of PIT imaging capability with unsupervised learning (UL), which enables PIT microscopy of subsurface defects in high strength corrosion resistant stainless steel 316 alloy. PIT images were processed with UL spatial–temporal separation-based clustering segmentation (STSCS) algorithm, refined by morphology image processing methods to enhance visibility of defects. The STSCS algorithm starts with wavelet decomposition to spatially de-noise thermograms, followed by UL principal component analysis (PCA), fine-tuning optimization, and neural learning-based independent component analysis (ICA) algorithms to temporally compress de-noised thermograms. The compressed thermograms were further processed with UL-based graph thresholding K-means clustering algorithm for defects segmentation. The STSCS algorithm also includes online learning feature for efficient re-training of the model with new data. For this study, metallic specimens with calibrated microscopic flat bottom hole defects, with diameters in the range from 203 to 76 µm, were produced using electro discharge machining (EDM) drilling. While the raw thermograms do not show any material defects, using STSCS algorithm to process PIT images reveals defects as small as 101 µm in diameter. To the best of our knowledge, this is the smallest reported size of a sub-surface defect in a metal imaged with PIT, which demonstrates the PIT capability of detecting defects in the size range relevant to quality control requirements of LPBF-printed high-strength metals.

36 MATERIALS SCIENCE↗

The Optimization of Trained and Untrained Image Classification Algorithms for Use on Large Spatial Datasets

The HARVIST project seeks to automatically provide an accurate, interactive interface to predict crop yield over the entire United States. In order to accomplish this goal, large images must be quickly and automatically classified by crop type. Current trained and untrained classification algorithms, while accurate, are highly inefficient when operating on large datasets. This project sought to develop new variants of two standard trained and untrained classification algorithms that are optimized to take advantage of the spatial nature of image data. The first algorithm, harvist-cluster, utilizes divide-and-conquer techniques to precluster an image in the hopes of increasing overall clustering speed. The second algorithm, harvistSVM, utilizes support vector machines (SVMs), a type of trained classifier. It seeks to increase classification speed by applying a "meta-SVM" to a quick (but inaccurate) SVM to approximate a slower, yet more accurate, SVM. Speedups were achieved by tuning the algorithm to quickly identify when the quick SVM was incorrect, and then reclassifying low-confidence pixels as necessary. Comparing the classification speeds of both algorithms to known baselines showed a slight speedup for large values of k (the number of clusters) for harvist-cluster, and a significant speedup for harvistSVM. Future work aims to automate the parameter tuning process required for harvistSVM, and further improve classification accuracy and speed. Additionally, this research will move documents created in Canvas into ArcGIS. The launch of the Mars Reconnaissance Orbiter (MRO) will provide a wealth of image data such as global maps of Martian weather and high resolution global images of Mars. The ability to store this new data in a georeferenced format will support future Mars missions by providing data for landing site selection and the search for water on Mars.

Kocurek, Michael J.↗

Anomaly detection for MPC forecast in Fleet of Water Heaters

Among residential devices, water heaters consume 20% of home energy use in the United States. Water heaters possess the capability to store energy within their reservoirs, enabling the ability to decouple energy use from hot water use. This capability can be used to reduce energy usage and costs while also supporting grid services. This requires accurate forecasting of the parameters of the water heater such as upper and lower temperatures. In this study, we analyzed the performance and behavior of a water heater model used in the real-world to predict a control mechanism that is implemented in a smart residential neighborhood. The model forecasts are accurate in most cases but not all. In such scenarios, error correction of the model is necessary to further improve model predictive control accuracy. Anomaly detection is the first step of error correction. This study complements existing research by grouping time series data into two clusters one with anomalies and another without anomalies. To achieve this task, we explored and compared multiple unsupervised machine learning algorithms to perform clustering. Among these algorithms, Ward clustering has the lowest running time and identified the highest number of anomalies for the upper temperature limit. The proposed approach is tested based on the data collected in a neighborhood with 46 townhomes located in Atlanta, GA.

Lebakula, Viswadeep↗

Measurement of the Splashback Feature Around SZ-Selected Galaxy Clusters With DES, SPT, and ACT

We present a detection of the splashback feature around galaxy clusters selected using the Sunyaev–Zel’dovich (SZ) signal. Recent measurements of the splashback feature around optically selected galaxy clusters have found that the splashback radius, rsp, is smaller than predicted by N-body simulations. A possible explanation for this discrepancy is that rsp inferred from the observed radial distribution of galaxies is affected by selection effects related to the optical cluster-finding algorithms. We test this possibility by measuring the splashback feature in clusters selected via the SZ effect in data from the South Pole Telescope SZ survey and the Atacama Cosmology Telescope Polarimeter survey. The measurement is accomplished by correlating these cluster samples with galaxies detected in the Dark Energy Survey Year 3data. The SZ observable used to select clusters in this analysis is expected to have a tighter correlation with halo mass and to be more immune to projection effects and aperture-induced biases, potentially ameliorating causes of systematic error for optically selected clusters. We find that the measured rsp for SZ-selected clusters is consistent with the expectations from simulations, although the small number of SZ-selected clusters makes a precise comparison difficult. In agreement with previous work, when using optically selected red MaPPer clusters with similar mass and redshift distributions,rspis∼2σsmaller than in the simulations. These results motivate detailed investigations of selection biases in optically selected cluster catalogues and exploration of the splashback feature around larger samples of SZ-selected clusters. Additionally, we investigate trends in the galaxy profile and splashback feature as a function of galaxy colour, finding that blue galaxies have profiles close to a power law with no discernible splashback feature, which is consistent with them being on their first in fall into the cluster.

T Shin↗

Clustering of noisy image data using an adaptive neuro-fuzzy system

Identification of outliers or noise in a real data set is often quite difficult. A recently developed adaptive fuzzy leader clustering (AFLC) algorithm has been modified to separate the outliers from real data sets while finding the clusters within the data sets. The capability of this modified AFLC algorithm to identify the outliers in a number of real data sets indicates the potential strength of this algorithm in correct classification of noisy real data.

Pemmaraju, Surya↗

Photometric Redshifts and Galaxy Clusters for DES DR2, DESI DR9, and HSC-SSP PDR3 Data

Photometric redshift (photoz) is a fundamental parameter for multi-wavelength photometric surveys, while galaxy clusters are important cosmological probes and ideal objects for exploring the dense environmental impact on galaxy evolution. We extend our previous work on estimating photoz and detecting galaxy clusters to the latest data releases of the Dark Energy Spectroscopic Instrument (DESI) imaging surveys, Dark Energy Survey (DES) and Hyper Suprime-Cam Subaru Strategic Program (HSC-SSP) imaging surveys and make corresponding catalogs publicly available for more extensive scientific applications. The photoz catalogs include accurate measurements of photoz and stellar mass for about 320, 293 and 134 million galaxies with r < 23, i < 24 and i < 25 in DESI DR9, DES DR2 and HSC-SSP PDR3 data, respectively. The photoz accuracy is about 0.017, 0.024 and 0.029 and the general redshift coverage is z < 1, z < 1.2 and z < 1.6, respectively for those three surveys. Furthermore, the uncertainty of the logarithmic stellar mass that is inferred from stellar population synthesis fitting is about 0.2 dex. With the above photoz catalogs, galaxy clusters are detected using a fast cluster-finding algorithm. A total of 532,810, 86,963 and 36,566 galaxy clusters with the number of members larger than 10 is discovered for DESI, DES and HSC-SSP, respectively. Their photoz accuracy is at the level of 0.01. The total mass of our clusters is also estimated by using the calibration relations between the optical richness and the mass measurement from X-ray and radio observations. The photoz and cluster catalogs are available at ScienceDB (https://www.doi.org/10.11922/sciencedb.o00069.00003) and PaperData Repository (https://doi.org/10.12149/101089).

79 ASTRONOMY AND ASTROPHYSICS↗

Optimization of Canister Loading Patterns in Dual Purpose Canisters for Criticality Suppression

This report documents work performed in support of the US Department of Energy Office of Nuclear Energy (NE) Spent Fuel and Waste Disposition, Spent Fuel and Waste Science and Technology, under work breakdown structure element 1.08.01.03.05, “Direct Disposal of Dual Purpose Canisters.” In particular, this report fulfills milestone M3SF-22PN010305094, “Application of AI/ML techniques to optimize DPC loading,” within work package SF-22PN01030509, “Direct Disposal of Dual Purpose Canisters - PNNL.” This report continues the process of examining the potential for using loading optimization as a disposal criticality suppression technique. This latest update to the report added: 1. A validation of the artificial neural network (ANN) reactivity prediction tool against as-loaded dual-purpose canisters from the UNF-ST&DARDS database 2. Refinement of the algorithm to incorporate an additional DPC design to improve the predictions 3. Investigation of a storage and transportation focused loading optimization algorithm. A key result from this year’s work is that the ANN performance was significantly improved by including the additional model for the NUHOMS canisters and it is apparent that fuel type specific modeling considerations should be incorporated into the ANN in future work. The GRASP-enabled adaptive multi-objective memetic algorithm with partial clustering (GAMMA-PC) algorithm was evaluated as a tool for storage and transportation oriented optimization. The GAMMA-PC optimization routine considers decay heat at loading, the time between loading and when the canister is eligible for transportation and minimizes the number of casks loaded. Future work will develop a dose minimization optimization routine.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Survey of adaptive image coding techniques

The general problem of image data compression is discussed briefly with attention given to the use of Karhunen-Loeve transforms, suboptimal systems, and block quantization. A survey is then conducted encompassing the four categories of adaptive systems: (1) adaptive transform coding (adaptive sampling, adaptive quantization, etc.), (2) adaptive predictive coding (adaptive delta modulation, adaptive DPCM encoding, etc.), (3) adaptive cluster coding (blob algorithms and the multispectral cluster coding technique), and (4) adaptive entropy coding.

Habibi, A.↗

Modelling galaxy cluster triaxiality in stacked cluster weak lensing analyses

Counts of galaxy clusters offer a high-precision probe of cosmology, but control of systematic errors will determine the accuracy of this measurement. Using Buzzard simulations, we quantify one such systematic, the triaxiality distribution of clusters identified with the redMaPPer optical cluster finding algorithm, which was used in the Dark Energy Survey Year-1 (DES Y1) cluster cosmology analysis. We test whether redMaPPer selection biases the clusters’ shape and orientation and find that it only biases orientation, preferentially selecting clusters with their major axes oriented along the line of sight. Modelling the richness–mass relation as log-linear, we find that the log-richness amplitude ln (A) is boosted from the lowest to highest orientation bin with a significance of 14σ, while the orientation dependence of the richness-mass slope and intrinsic scatter is minimal. We also find that the weak lensing shear-profile ratios of cluster-associated dark haloes in different orientation bins resemble a ‘bottleneck’ shape that can be quantified with a Cauchy function. We test the correlation of orientation with two other leading systematics in cluster cosmology – miscentering and projection – and find a null correlation. The resulting mass bias predicted from our templates confirms the DES Y1 finding that triaxiality is a leading source of bias in cluster cosmology. However, the richness-dependence of the bias confirms that triaxiality does not fully resolve the tension at low-richness between DES Y1 cluster cosmology and other probes. Our model can be used for quantifying the impact of triaxiality bias on cosmological constraints for upcoming weak lensing surveys of galaxy clusters.

79 ASTRONOMY AND ASTROPHYSICS↗

Scheduling Tasks In Parallel Processing

Algorithms sought to minimize time and cost of computation. Report describes research on scheduling of computations tasks in system of multiple identical data processors operating in parallel. Computational intractability requires use of suboptimal heuristic algorithms. First algorithm called "list heuristic", variation of classical list scheduling. Second algorithm called "cluster heuristic" applied to tightly coupled tasks and consists of four phases. Third algorithm called "exchange heuristic", iterative-improvement algorithm beginning with initial feasible assignment of tasks to processors and periods of time. Fourth algorithm is iterative one for optimal assignment of tasks and based on concept called "simulated annealing" because of mathematical resemblance to aspects of physical annealing processes.

Price, Camille C.↗

The XMM cluster survey: exploring scaling relations and completeness of the dark energy survey year 3 redMaPPer cluster catalogue

ABSTRACT We cross-match and compare characteristics of galaxy clusters identified in observations from two sky surveys using two completely different techniques. One sample is optically selected from the analysis of 3 years of Dark Energy Survey observations using the redMaPPer cluster detection algorithm. The second is X-ray selected from XMM observations analysed by the XMM Cluster Survey. The samples comprise a total area of 57.4 deg2, bounded by the area of four contiguous XMM survey regions that overlap the DES footprint. We find that the X-ray-selected sample is fully matched with entries in the redMaPPer catalogue, above λ > 20 and within 0.1 <$z$ <0.9. Conversely, only 38 per cent of the redMaPPer catalogue is matched to an X-ray extended source. Next, using 120 optically clusters and 184 X-ray-selected clusters, we investigate the form of the X-ray luminosity–temperature (LX –TX ), luminosity–richness (LX –λ), and temperature–richness (TX –λ) scaling relations. We find that the fitted forms of the LX –TX relations are consistent between the two selection methods and also with other studies in the literature. However, we find tentative evidence for a steepening of the slope of the relation for low richness systems in the X-ray-selected sample. When considering the scaling of richness with X-ray properties, we again find consistency in the relations (i.e. LX –λ and TX –λ) between the optical and X-ray-selected samples. This is contrary to previous similar works that find a significant increase in the scatter of the luminosity scaling relation for X-ray-selected samples compared to optically selected samples.

79 ASTRONOMY AND ASTROPHYSICS↗

Improve Data Mining and Knowledge Discovery Through the Use of MatLab

Data mining is widely used to mine business, engineering, and scientific data. Data mining uses pattern based queries, searches, or other analyses of one or more electronic databases/datasets in order to discover or locate a predictive pattern or anomaly indicative of system failure, criminal or terrorist activity, etc. There are various algorithms, techniques and methods used to mine data; including neural networks, genetic algorithms, decision trees, nearest neighbor method, rule induction association analysis, slice and dice, segmentation, and clustering. These algorithms, techniques and methods used to detect patterns in a dataset, have been used in the development of numerous open source and commercially available products and technology for data mining. Data mining is best realized when latent information in a large quantity of data stored is discovered. No one technique solves all data mining problems; challenges are to select algorithms or methods appropriate to strengthen data/text mining and trending within given datasets. In recent years, throughout industry, academia and government agencies, thousands of data systems have been designed and tailored to serve specific engineering and business needs. Many of these systems use databases with relational algebra and structured query language to categorize and retrieve data. In these systems, data analyses are limited and require prior explicit knowledge of metadata and database relations; lacking exploratory data mining and discoveries of latent information. This presentation introduces MatLab(R) (MATrix LABoratory), an engineering and scientific data analyses tool to perform data mining. MatLab was originally intended to perform purely numerical calculations (a glorified calculator). Now, in addition to having hundreds of mathematical functions, it is a programming language with hundreds built in standard functions and numerous available toolboxes. MatLab's ease of data processing, visualization and its enormous availability of built in functionalities and toolboxes make it suitable to perform numerical computations and simulations as well as a data mining tool. Engineers and scientists can take advantage of the readily available functions/toolboxes to gain wider insight in their perspective data mining experiments.

Shaykhian, Gholam Ali↗

Improve Data Mining and Knowledge Discovery through the use of MatLab

Data mining is widely used to mine business, engineering, and scientific data. Data mining uses pattern based queries, searches, or other analyses of one or more electronic databases/datasets in order to discover or locate a predictive pattern or anomaly indicative of system failure, criminal or terrorist activity, etc. There are various algorithms, techniques and methods used to mine data; including neural networks, genetic algorithms, decision trees, nearest neighbor method, rule induction association analysis, slice and dice, segmentation, and clustering. These algorithms, techniques and methods used to detect patterns in a dataset, have been used in the development of numerous open source and commercially available products and technology for data mining. Data mining is best realized when latent information in a large quantity of data stored is discovered. No one technique solves all data mining problems; challenges are to select algorithms or methods appropriate to strengthen data/text mining and trending within given datasets. In recent years, throughout industry, academia and government agencies, thousands of data systems have been designed and tailored to serve specific engineering and business needs. Many of these systems use databases with relational algebra and structured query language to categorize and retrieve data. In these systems, data analyses are limited and require prior explicit knowledge of metadata and database relations; lacking exploratory data mining and discoveries of latent information. This presentation introduces MatLab(TradeMark)(MATrix LABoratory), an engineering and scientific data analyses tool to perform data mining. MatLab was originally intended to perform purely numerical calculations (a glorified calculator). Now, in addition to having hundreds of mathematical functions, it is a programming language with hundreds built in standard functions and numerous available toolboxes. MatLab's ease of data processing, visualization and its enormous availability of built in functionalities and toolboxes make it suitable to perform numerical computations and simulations as well as a data mining tool. Engineers and scientists can take advantage of the readily available functions/toolboxes to gain wider insight in their perspective data mining experiments.

Shaykahian, Gholan Ali↗

GOES-R AWG GLM Val Tool Development

We are developing tools needed to enable the validation of the Geostationary Lightning Mapper (GLM). In order to develop and test these tools, we have need of a robust, high-fidelity set of GLM proxy data. Many steps have been taken to ensure that the proxy data are high quality. LIS is the closest analog that exists for GLM, so it has been used extensively in developing the GLM proxy. We have verified the proxy data both statistically and algorithmically. The proxy data are pixel (event) data, called Level 1B. These data were then clustered into flashes by the Lightning Cluster-Filter Algorithm (LCFA), generating proxy Level 2 data. These were then compared with the data used to generate the proxy, and both the proxy data and the LCFA were validated. We have developed tools to allow us to visualize and compare the GLM proxy data with several other sources of lightning and other meteorological data (the so-called shallow-dive tool). The shallow-dive tool shows storm-level data and can ingest many different ground-based lightning detection networks, including: NLDN, LMA, WWLLN, and ENTLN. These are presented in a way such that it can be seen if the GLM is properly detecting the lightning in location and time comparable to the ground-based networks. Currently in development is the deep-dive tool, which will allow us to dive into the GLM data, down to flash, group and event level. This will allow us to assess performance in comparison with other data sources, and tell us if there are detection, timing, or geolocation problems. These tools will be compatible with the GLM Level-2 data format, so they can be used beginning on Day 0.

Bateman, Monte↗