Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Cluster analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Large-Scale Circulation and Climate Variability

The causes of regional climate trends cannot be understood without considering the impact of variations in large-scale atmospheric circulation and an assessment of the role of internally generated climate variability. There are contributions to regional climate trends from changes in large-scale latitudinal circulation, which is generally organized into three cells in each hemisphere-Hadley cell, Ferrell cell and Polar cell-and which determines the location of subtropical dry zones and midlatitude jet streams. These circulation cells are expected to shift poleward during warmer periods, which could result in poleward shifts in precipitation patterns, affecting natural ecosystems, agriculture, and water resources. In addition, regional climate can be strongly affected by non-local responses to recurring patterns (or modes) of variability of the atmospheric circulation or the coupled atmosphere-ocean system. These modes of variability represent preferred spatial patterns and their temporal variation. They account for gross features in variance and for teleconnections which describe climate links between geographically separated regions. Modes of variability are often described as a product of a spatial climate pattern and an associated climate index time series that are identified based on statistical methods like Principal Component Analysis (PC analysis), which is also called Empirical Orthogonal Function Analysis (EOF analysis), and cluster analysis.

Perlwitz, J.↗

Orbit Clustering Based on Transfer Cost

We propose using cluster analysis to perform quick screening for combinatorial global optimization problems. The key missing component currently preventing cluster analysis from use in this context is the lack of a useable metric function that defines the cost to transfer between two orbits. We study several proposed metrics and clustering algorithms, including k-means and the expectation maximization algorithm. We also show that proven heuristic methods such as the Q-law can be modified to work with cluster analysis.

combinatorial optimization↗

Sulfur in Cometary Dust

The computer-intensive project consisted of the analysis and synthesis of existing data on composition of comet Halley dust particles. The main objective was to obtain a complete inventory of sulfur containing compounds in the comet Halley dust by building upon the existing classification of organic and inorganic compounds and applying a variety of statistical techniques for cluster and cross-correlational analyses. A student hired for this project wrote and tested the software to perform cluster analysis. The following tasks were carried out: (1) selecting the data from existing database for the proposed project; (2) finding access to a standard library of statistical routines for cluster analysis; (3) reformatting the data as necessary for input into the library routines; (4) performing cluster analysis and constructing hierarchical cluster trees using three methods to define the proximity of clusters; (5) presenting the output results in different formats to facilitate the interpretation of the obtained cluster trees; (6) selecting groups of data points common for all three trees as stable clusters. We have also considered the chemistry of sulfur in inorganic compounds.

Fomenkova, M. N.↗

Extracting galactic structure parameters from multivariated density estimation

Multivariate statistical analysis, including includes cluster analysis (unsupervised classification), discriminant analysis (supervised classification) and principle component analysis (dimensionlity reduction method), and nonparameter density estimation have been successfully used to search for meaningful associations in the 5-dimensional space of observables between observed points and the sets of simulated points generated from a synthetic approach of galaxy modelling. These methodologies can be applied as the new tools to obtain information about hidden structure otherwise unrecognizable, and place important constraints on the space distribution of various stellar populations in the Milky Way. In this paper, we concentrate on illustrating how to use nonparameter density estimation to substitute for the true densities in both of the simulating sample and real sample in the five-dimensional space. In order to fit model predicted densities to reality, we derive a set of equations which include n lines (where n is the total number of observed points) and m (where m: the numbers of predefined groups) unknown parameters. A least-square estimation will allow us to determine the density law of different groups and components in the Galaxy. The output from our software, which can be used in many research fields, will also give out the systematic error between the model and the observation by a Bayes rule.

Chen, B.↗

Designers' models of the human-computer interface

Understanding design models of the human-computer interface (HCI) may produce two types of benefits. First, interface development often requires input from two different types of experts: human factors specialists and software developers. Given the differences in their backgrounds and roles, human factors specialists and software developers may have different cognitive models of the HCI. Yet, they have to communicate about the interface as part of the design process. If they have different models, their interactions are likely to involve a certain amount of miscommunication. Second, the design process in general is likely to be guided by designers' cognitive models of the HCI, as well as by their knowledge of the user, tasks, and system. Designers do not start with a blank slate; rather they begin with a general model of the object they are designing. The author's approach to a design model of the HCI was to have three groups make judgments of categorical similarity about the components of an interface: human factors specialists with HCI design experience, software developers with HCI design experience, and a baseline group of computer users with no experience in HCI design. The components of the user interface included both display components such as windows, text, and graphics, and user interaction concepts, such as command language, editing, and help. The judgments of the three groups were analyzed using hierarchical cluster analysis and Pathfinder. These methods indicated, respectively, how the groups categorized the concepts, and network representations of the concepts for each group. The Pathfinder analysis provides greater information about local, pairwise relations among concepts, whereas the cluster analysis shows global, categorical relations to a greater extent.

Gillan, Douglas J.↗

Probabilistic Classification Using Elemental Abundance Distributions and Lossless Image Compression in Apollo 17 Lunar Dust Samples from Mare Serenitatis

We have previously outlined a strategy for the detection of fossils [Storrie-Lombardi and Hoover, 2004] and extant microbial life [Storrie-Lombaudi and Hoover, 20051 during robotic missions to Mars using co-registered structural and chemical signatures. Data inputs included image lossless compression indices to estimate relative textural complexity and elemental abundance distributions. Two exploratory classification algorithms (principal component analysis and hierarchical cluster analysis) provide an initial tentative classification of all targets. Nonlinear stochastic neural networks are then trained to produce a Bayesian estimate of algorithm classification accuracy. The strategy previously has been successful in distinguishing regions of biotic and abiotic alteration of basalt glass from unaltered samples. [Storrie-Lombardi and Fisk, 2004; Storrie-Lombardi and Fisk, 2004] Such investigations of abiotic versus biotic alteration of terrestrial mineralogy on Earth are compromised by .the difficulty finding mineralogy completely unaffected by the ubiquitous presence of microbial life on the planet. The renewed interest in lunar exploration offers an opportunity to investigate geological materials that may exhibit signs of aqueous alteration, but are highly unlikely to contain contaminating biological weathering signatures. We here present an extension of our earlier data set to include lunar dust samples obtained during the Apollo 17 mission. Apollo 17 landed in the Taurus-Littrow Valley in Mare Serenitatis. Most of the rock samples from this region of the lunar highlands are basalts comprised primarily of plagioclase and pyroxene and selected examples of orange and black volcanic glass. SEM images and elemental abundances (C6, N7, O8, Na11, Mg12, Al13, Si14, P15, S16, Cll7, K19, Ca20, Fe26) for a series of targets in the lunar dust samples are compared to the extant cyanobacteria, fossil trilobites, Orgueil meteorite, and terrestrial basalt targets previously discussed. The data set provides a first step in producing a quantitative probabilistic methodology for geobiological analysis of returned lunar samples or in situ exploration.

Storrie-Lombardi, Michael C.↗

Use of UV Sources for Detection and Identification of Explosives

Measurement of Raman and native fluorescence emission using ultraviolet (UV) sources (<400 nm) on targeted materials is suitable for both sensitive detection and accurate identification of explosive materials. When the UV emission data are analyzed using a combination of Principal Component Analysis (PCA) and cluster analysis, chemicals and biological samples can be differentiated based on the geometric arrangement of molecules, the number of repeating aromatic rings, associated functional groups (nitrogen, sulfur, hydroxyl, and methyl), microbial life cycles (spores vs. vegetative cells), and the number of conjugated bonds. Explosive materials can be separated from one another as well as from a range of possible background materials, which includes microbes, car doors, motor oil, and fingerprints on car doors, etc. Many explosives are comprised of similar atomic constituents found in potential background samples such as fingerprint oils/skin, motor oil, and soil. This technique is sensitive to chemical bonds between the elements that lead to the discriminating separability between backgrounds and explosive materials.

Hug, William↗

Lumping and splitting: Toward a classification of mineral natural kinds

How does one best subdivide nature into kinds? All classification systems require rules for lumping similar objects into the same category, while splitting differing objects into separate categories. Mineralogical classification systems are no exception. Our work in placing mineral species within their evolutionary contexts necessitates this lumping and splitting because we classify “mineral natural kinds” based on unique combinations of formational environments and continuous temperature-pressure-composition phase space. Consequently, we lump two minerals into a single natural kind only if they: (1) are part of a continuous solid solution; (2) are isostructural or members of a homologous series; and (3) form by the same process. A systematic survey based on these criteria suggests that 2310 (~41%) of 5659 IMA-approved mineral species can be lumped with one or more other mineral species, corresponding to 667 “root mineral kinds,” of which 353 lump pairs of mineral species, while 129 lump three species. Eight mineral groups, including cancrinite, eudialyte, hornblende, jahnsite, labuntsovite, satorite, tetradymite, and tourmaline, are represented by 20 or more lumped IMA-approved mineral species. A list of 5659 IMA-approved mineral species corresponds to 4016 root mineral kinds according to these lumping criteria. The evolutionary system of mineral classification assigns an IMA-approved mineral species to two or more mineral natural kinds under either of two splitting criteria: (1) if it forms in two or more distinct paragenetic environments, or (2) if cluster analysis of the attributes of numerous specimens reveals more than one discrete combination of chemical and physical attributes. A total of 2310 IMA-approved species are known to form by two or more paragenetic processes and thus correspond to multiple mineral natural kinds; however, adequate data resources are not yet in hand to perform cluster analysis on more than a handful of mineral species. We find that 1623 IMA-approved species (~29%) correspond exactly to mineral natural kinds; i.e., they are known from only one paragenetic environment and are not lumped with another species in our evolutionary classification. Greater complexity is associated with 587 IMA-approved species that are both lumped with one or more other species and occur in two or more paragenetic environments. In these instances, identification of mineral natural kinds may involve both lumping and splitting of the corresponding IMA-approved species on the basis of multiple criteria. Based on the numbers of root mineral kinds, their known varied modes of formation, and predictions of minerals that occur on Earth but are as yet undiscovered and described, we estimate that Earth holds more than 10000 mineral natural kinds.

Philosophy of mineralogy↗

Identification of a typical flight patterns

Method and system for analyzing aircraft data, including multiple selected flight parameters for a selected phase of a selected flight, and for determining when the selected phase of the selected flight is atypical, when compared with corresponding data for the same phase for other similar flights. A flight signature is computed using continuous- valued and discrete-valued flight parameters for the selected flight parameters and is optionally compared with a statistical distribution of other observed flight signatures, yielding a typicality scores for the same phase for other similar flights. A cluster analysis is optionally applied to the flight signatures to define an optimal collection of clusters. A level of atypicality for a selected flight is estimated, based upon an index associated with the cluster analysis.

Irving C Statler↗

Identification of atypical flight patterns

Method and system for analyzing aircraft data, including multiple selected flight parameters for a selected phase of a selected flight, and for determining when the selected phase of the selected flight is atypical, when compared with corresponding data for the same phase for other similar flights. A flight signature is computed using continuous-valued and discrete-valued flight parameters for the selected flight parameters and is optionally compared with a statistical distribution of other observed flight signatures, yielding atypicality scores for the same phase for other similar flights. A cluster analysis is optionally applied to the flight signatures to define an optimal collection of clusters. A level of atypicality for a selected flight is estimated, based upon an index associated with the cluster analysis.

Statler, Irving C.↗

An analysis of pilot error-related aircraft accidents

A multidisciplinary team approach to pilot error-related U.S. air carrier jet aircraft accident investigation records successfully reclaimed hidden human error information not shown in statistical studies. New analytic techniques were developed and applied to the data to discover and identify multiple elements of commonality and shared characteristics within this group of accidents. Three techniques of analysis were used: Critical element analysis, which demonstrated the importance of a subjective qualitative approach to raw accident data and surfaced information heretofore unavailable. Cluster analysis, which was an exploratory research tool that will lead to increased understanding and improved organization of facts, the discovery of new meaning in large data sets, and the generation of explanatory hypotheses. Pattern recognition, by which accidents can be categorized by pattern conformity after critical element identification by cluster analysis.

Kowalsky, N. B.↗

Determining the Optimal Number of Clusters with the Clustergram

Cluster analysis aids research in many different fields, from business to biology to aerospace. It consists of using statistical techniques to group objects in large sets of data into meaningful classes. However, this process of ordering data points presents much uncertainty because it involves several steps, many of which are subject to researcher judgment as well as inconsistencies depending on the specific data type and research goals. These steps include the method used to cluster the data, the variables on which the cluster analysis will be operating, the number of resulting clusters, and parts of the interpretation process. In most cases, the number of clusters must be guessed or estimated before employing the clustering method. Many remedies have been proposed, but none is unassailable and certainly not for all data types. Thus, the aim of current research for better techniques of determining the number of clusters is generally confined to demonstrating that the new technique excels other methods in performance for several disparate data types. Our research makes use of a new cluster-number-determination technique based on the clustergram: a graph that shows how the number of objects in the cluster and the cluster mean (the ordinate) change with the number of clusters (the abscissa). We use the features of the clustergram to make the best determination of the cluster-number.

Fluegemann, Joseph K.↗

Satellite Remote Sensing to Assess Cyanobacterial Bloom Frequency Across the United States at Multiple Spatial Scales

Cyanobacterial blooms can have negative effects on human health and local ecosystems. Field monitoring of cyanobacterial blooms can be costly, but satellite remote sensing has shown utility for more efficient spatial and temporal monitoring across the United States. Here, satellite imagery was used to assess the annual frequency of surface cyanobacterial blooms, defined for each satellite pixel as the percentage of images for that pixel throughout the year exhibiting detectable cyanobacteria. Cyanobacterial frequency was assessed across 2,196 large lakes in 46 states across the continental United States (CONUS) using imagery from the European Space Agency’s Ocean and Land Colour Imager for the years 2017 through 2019. In 2019, across all satellite pixels considered, annual bloom frequency had a median value of 4% and a maximum value of 100%, the latter indicating that for those satellite pixels, a cyanobacterial bloom was detected by the satellite sensor for every satellite image considered. In addition to annual pixel-scale cyanobacterial frequency, results were summarized at the lake- and state-scales by averaging annual pixel-scale results across each lake and state. For 2019, average annual lake-scale frequencies also had a maximum value of 100%, and Oregon and Ohio had the highest average annual state-scale frequencies at 65% and 52%. Pixel-scale frequency results can assist in identifying portions of a lake that are more prone to cyanobacterial blooms, while lake- and state-scale frequency results can assist in the prioritization of sampling resources and mitigation efforts. Satellite imagery is limited by the presence of snow and ice, as imagery collected in these conditions are quality flagged and discarded. Thus, annual bloom frequencies within nine climate regions were investigated to determine whether missing data biased results in climate regions more prone to snow and ice, given that their annual summaries would be weighted toward the summer months when cyanobacterial blooms tend to occur. Results were unbiased by the time period selected in most climate regions, but a large bias was observed for the Northwest Rockies and Plains climate region. Moderate biases were observed for the Ohio Valley and the Southeast climate regions. Finally, a clustering analysis was used to identify areas of high and low cyanobacterial frequency across CONUS based on average annual lake-scale cyanobacterial frequencies for 2019. Several clusters were identified that transcended state, watershed, and eco-regional boundaries. Combined with additional data, results from the clustering analysis may offer insight regarding large-scale drivers of cyanobacterial blooms.

remote sensing↗

Technical support for creating an artificial intelligence system for feature extraction and experimental design

Techniques for classifying objects into groups or clases go under many different names including, most commonly, cluster analysis. Mathematically, the general problem is to find a best mapping of objects into an index set consisting of class identifiers. When an a priori grouping of objects exists, the process of deriving the classification rules from samples of classified objects is known as discrimination. When such rules are applied to objects of unknown class, the process is denoted classification. The specific problem addressed involves the group classification of a set of objects that are each associated with a series of measurements (ratio, interval, ordinal, or nominal levels of measurement). Each measurement produces one variable in a multidimensional variable space. Cluster analysis techniques are reviewed and methods for incuding geographic location, distance measures, and spatial pattern (distribution) as parameters in clustering are examined. For the case of patterning, measures of spatial autocorrelation are discussed in terms of the kind of data (nominal, ordinal, or interval scaled) to which they may be applied.

Glick, B. J.↗

Preliminary Comparisons of the Information Content and Utility of TM Versus MSS Data

Comparisons were made between subscenes from the first TM scene acquired of the Washington, D.C. area and a MSS scene acquired approximately one year earlier. Three types of analyses were conducted to compare TM and MSS data: a water body analysis, a principal components analysis and a spectral clustering analysis. The water body analysis compared the capability of the TM to the MSS for detecting small uniform targets. Of the 59 ponds located on aerial photographs 34 (58%) were detected by the TM with six commission errors (15%) and 13 (22%) were detected by the MSS with three commission errors (19%). The smallest water body detected by the TM was 16 meters; the smallest detected by the MSS was 40 meters. For the principal components analysis, means and covariance matrices were calculated for each subscene, and principal components images generated and characterized. In the spectral clustering comparison each scene was independently clustered and the clusters were assigned to informational classes. The preliminary comparison indicated that TM data provides enhancements over MSS in terms of (1) small target detection and (2) data dimensionality (even with 4-band data). The extra dimension, partially resultant from TM band 1, appears useful for built-up/non-built-up area separation.

Markham, B. L.↗

Functional Groups Based on Leaf Physiology: Are they Spatially and Temporally Robust?

The functional grouping hypothesis, which suggests that complexity in ecosystem function can be simplified by grouping species with similar responses, was tested in the Florida scrub habitat. Functional groups were identified based on how species in fire maintained Florida scrub regulate exchange of carbon and water with the atmosphere as indicated by both instantaneous gas exchange measurements and integrated measures of function (%N, delta C-13, delta N-15, C-N ratio). Using cluster analysis, five distinct physiologically-based functional groups were identified in the fire maintained scrub. These functional groups were tested to determine if they were robust spatially, temporally, and with management regime. Analysis of Similarities (ANOSIM), a non-parametric multivariate analysis, indicated that these five physiologically-based groupings were not altered by plot differences (R = -0.115, p = 0.893) or by the three different management regimes; prescribed burn, mechanically treated and burn, and fire-suppressed (R = 0.018, p = 0.349). The physiological groupings also remained robust between the two climatically different years 1999 and 2000 (R = -0.027, p = 0.725). Easy-to-measure morphological characteristics indicating functional groups would be more practical for scaling and modeling ecosystem processes than detailed gas-exchange measurements, therefore we tested a variety of morphological characteristics as functional indicators. A combination of non-parametric multivariate techniques (Hierarchical cluster analysis, non-metric Multi-Dimensional Scaling, and ANOSIM) were used to compare the ability of life form, leaf thickness, and specific leaf area classifications to identify the physiologically-based functional groups. Life form classifications (ANOSIM; R = 0.629, p 0.001) were able to depict the physiological groupings more adequately than either specific leaf area (ANOSIM; R = 0.426, p = 0.001) or leaf thickness (ANOSIM; R 0.344, p 0.001). The ability of life forms to depict the physiological groupings was improved by separating the parasitic Ximenia americana from the shrub category (ANOSIM; R = 0.794, p = 0.001). Therefore, a life form classification including parasites was determined to be a good indicator of the physiological processes of scrub species, and would be a useful method of grouping for scaling physiological processes to the ecosystem level.

Foster, Tammy E.↗

Applications of Fuzzy Set Theory to Satellite Soundings

The introduction of an appropriate fuzzy setting for satellite soundings and its application to clustering methods via unimodal fuzzy sets in the future is proposed. Methods of hard clustering analysis and fuzzy partitioned clustering were applied on simulated data with very encouraging results. The proposed clustering technique is discussed. The notion of a unimodal fuzzy set was chosen to represent the partition of a data set for two reasons: (1) it detects all the locations in the vector space where highly concentrated clusters of points exist; and (2) the notion is general enough to represent clusters that exhibit quite general distributions of points. The technique detects all of the existing unimodal fuzzy sets and realizes the maximum separation among them. It is economical in memory space and computational time requirements and also detects groups that are fairly generally distributed in the feature space.

Munteanu, M. J.↗

Dynamics of cD clusters of galaxies. II: Analysis of seven Abell clusters

We have investigated the dynamics of the seven Abell clusters A193, A399, A401, A1795, A1809, A2063, and A2124, based on redshift data reported previously by us (Hill & Oegerle, (1993)). These papers present the initial results of a survey of cD cluster kinematics, with an emphasis on studying the nature of peculiar velocity cD galaxies and their parent clusters. In the current sample, we find no evidence for significant peculiar cD velocities, with respect to the global velocity distribution. However, the cD in A2063 has a significant (3 sigma) peculiar velocity with respect to galaxies in the inner 1.5 Mpc/h, which is likely due to the merger of a subcluster with A2063. We also find significant evidence for subclustering in A1795, and a marginally peculiar cD velocity with respect to galaxies within approximately 200 kpc/h of the cD. The available x-ray, optical, and galaxy redshift data strongly suggest that a subcluster has merged with A1795. We propose that the subclusters which merged with A1795 and A2063 were relatively small, with shallow potential wells, so that the cooling flows in these clusters were not disrupted. Two-body gravitational models of the A399/401 and A2063/MKW3S systems indicate that A399/401 is a bound pair with a total virial mass of approximately 4 x 10(exp 15) solar mass/h, while A2063 and MKW3S are very unlikely to be bound.

Oegerle, William R.↗