Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “pattern clustering”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Using Clustering to Establish Climate Regimes from PCM Output

A multivariate statistical clustering technique--based on the k-means algorithm of Hartigan has been used to extract patterns of climatological significance from 200 years of general circulation model (GCM) output. Originally developed and implemented on a Beowulf-style parallel computer constructed by Hoffman and Hargrove from surplus commodity desktop PCs, the high performance parallel clustering algorithm was previously applied to the derivation of ecoregions from map stacks of 9 and 25 geophysical conditions or variables for the conterminous U.S. at a resolution of 1 sq km. Now applied both across space and through time, the clustering technique yields temporally-varying climate regimes predicted by transient runs of the Parallel Climate Model (PCM). Using a business-as-usual (BAU) scenario and clustering four fields of significance to the global water cycle (surface temperature, precipitation, soil moisture, and snow depth) from 1871 through 2098, the authors' analysis shows an increase in spatial area occupied by the cluster or climate regime which typifies desert regions (i.e., an increase in desertification) and a decrease in the spatial area occupied by the climate regime typifying winter-time high latitude perma-frost regions. The patterns of cluster changes have been analyzed to understand the predicted variability in the water cycle on global and continental scales. In addition, representative climate regimes were determined by taking three 10-year averages of the fields 100 years apart for northern hemisphere winter (December, January, and February) and summer (June, July, and August). The result is global maps of typical seasonal climate regimes for 100 years in the past, for the present, and for 100 years into the future. Using three-dimensional data or phase space representations of these climate regimes (i.e., the cluster centroids), the authors demonstrate the portion of this phase space occupied by the land surface at all points in space and time. Any single spot on the globe will exist in one of these climate regimes at any single point in time. By incrementing time, that same spot will trace out a trajectory or orbit between and among these climate regimes (or atmospheric states) in phase (or state) space. When a geographic region enters a state it never previously visited, a climatic change is said to have occurred. Tracing out the entire trajectory of a single spot on the globe yields a 'manifold' in state space representing the shape of its predicted climate occupancy. This sort of analysis enables a researcher to more easily grasp the multivariate behavior of the climate system.

Oglesby, Robert↗

Clonal Selection Based Artificial Immune System for Generalized Pattern Recognition

The last two decades has seen a rapid increase in the application of AIS (Artificial Immune Systems) modeled after the human immune system to a wide range of areas including network intrusion detection, job shop scheduling, classification, pattern recognition, and robot control. JPL (Jet Propulsion Laboratory) has developed an integrated pattern recognition/classification system called AISLE (Artificial Immune System for Learning and Exploration) based on biologically inspired models of B-cell dynamics in the immune system. When used for unsupervised or supervised classification, the method scales linearly with the number of dimensions, has performance that is relatively independent of the total size of the dataset, and has been shown to perform as well as traditional clustering methods. When used for pattern recognition, the method efficiently isolates the appropriate matches in the data set. The paper presents the underlying structure of AISLE and the results from a number of experimental studies.

pattern recognition↗

Applications of Some Artificial Intelligence Methods to Satellite Soundings

Hard clustering of temperature profiles and regression temperature retrievals were used to refine the method using the probabilities of membership of each pattern vector in each of the clusters derived with discriminant analysis. In hard clustering the maximum probability is taken and the corresponding cluster as the correct cluster are considered discarding the rest of the probabilities. In fuzzy partitioned clustering these probabilities are kept and the final regression retrieval is a weighted regression retrieval of several clusters. This method was used in the clustering of brightness temperatures where the purpose was to predict tropopause height. A further refinement is the division of temperature profiles into three major regions for classification purposes. The results are summarized in the tables total r.m.s. errors are displayed. An approach based on fuzzy logic which is intimately related to artificial intelligence methods is recommended.

Munteanu, M. J.↗

Top-Down Fabrication of Atomic Patterns in Twisted Bilayer Graphene

Atomic-scale engineering typically involves bottom-up approaches, leveraging parameters such as temperature, partial pressures, and chemical affinity to promote spontaneous arrangement of atoms. These parameters are applied globally, resulting in atomic-scale features scattered probabilistically throughout the material. In a top-down approach, different regions of the material are exposed to different parameters, resulting in structural changes varying on the scale of the resolution. In this work, the application of global and local parameters is combined in an aberration-corrected scanning transmission electron microscope (STEM) to demonstrate atomic-scale precision patterning of atoms in twisted bilayer graphene. The focused electron beam is used to define attachment points for foreign atoms through the controlled ejection of carbon atoms from the graphene lattice. The sample environment is staged with nearby source materials such that the sample temperature can induce migration of the source atoms across the sample surface. Under these conditions, the electron-beam (top-down) enables carbon atoms in the graphene to be replaced spontaneously by diffusing adatoms (bottom-up). Using image-based feedback control, arbitrary patterns of atoms and atom clusters are attached to the twisted bilayer graphene with limited human interaction. Finally, the role of substrate temperature on adatom and vacancy diffusion is explored by first-principles simulations.

36 MATERIALS SCIENCE↗

Spatiotemporal pattern detection, generation, and computation with circuits

Abstract Implementations of neurons, delays, and synapse circuits are presented with simulations. These neural elements are used to create two small spiking neural networks, the Rate-Window and Order-Biased clusters, which are capable of detecting simple two-spike spatiotemporal patterns. A simple pattern detecting network (SPDN) is created by combining the Rate-Window and Order-Biased clusters, where clusters are small spiking neural networks, and its simple pattern detection ability is demonstrated in simulation. The SPDN is used to implement a complex pattern detecting network (CPDN) and its complex pattern detection ability is demonstrated in simulation. Methods for generating arbitrary spatiotemporal patterns are presented. The CPDN and spatiotemporal pattern generation methods are then used to implement a novel spatiotemporal computing paradigm based on detecting and responding to spatiotemporal symbols. A simulation of a spatiotemporal half adder is presented to demonstrate the computing paradigm.

97 - MATHEMATICS AND COMPUTING↗

Climate Teleconnections and Recent Patterns of Human and Animal Disease Outbreaks

Recent clusters of outbreaks of mosquito-borne diseases (Rift Valley fever and chikungunya) in Africa and parts of the Indian Ocean islands illustrate how interannual climate variability influences the changing risk patterns of disease outbreaks. Extremes in rainfall (drought and flood) during the period 2004 - 2009 have privileged different disease vectors. Chikungunya outbreaks occurred during the severe drought from late 2004 to 2006 over coastal East Africa and the western Indian Ocean islands and in the later years India and Southeast Asia. The chikungunya pandemic was caused by a Central/East African genotype that appears to have been precipitated and then enhanced by global-scale and regional climate conditions in these regions. Outbreaks of Rift Valley fever occurred following excessive rainfall period from late 2006 to late 2007 in East Africa and Sudan, and then in 2008 - 2009 in Southern Africa. The shift in the outbreak patterns of Rift Valley fever from East Africa to Southern Africa followed a transition of the El Nino/Southern Oscillation (ENSO) phenomena from the warm El Nino phase (2006-2007) to the cold La Nina phase (2007-2009) and associated patterns of variability in the greater Indian Ocean basin that result in the displacement of the centres of above normal rainfall from Eastern to Southern Africa. Understanding the background patterns of climate variability both at global and regional scale and their impacts on ecological drivers of vector borne-diseases is critical in long-range planning of appropriate response and mitigation measures.

Anyamba, Assaf↗

Statistical relationships across epigenomes using large-scale hierarchical clustering

Recent advances in genomics and sequencing platforms have revolutionized our ability to create immense data sets, particularly for studying epigenetic regulation of gene expression. However, the avalanche of epigenomic data is difficult to parse for biological interpretation given nonlinear complex patterns and relationships. This attractive challenge in epigenomic data lends itself to machine learning for discerning infectivity and susceptibility. In this study, we explore over 3000 epigenomes of uninfected individuals and provide a framework to characterize the relationships among epigenetic modifiers, their modifiers, genetic loci, and specific immune cell types across all chromosomes using hierarchical clustering. Hierarchical clustering of epigenomic data revealed consistent epigenetic patterns across chromosomes, demonstrating that variation due to epigenetic modifiers is greater than variation between cell types. Gene Ontology and KEGG pathway analyses indicated significant enrichment of genes involved in chromatin remodeling, mRNA splicing, immune responses, and the regulation of microRNAs and snoRNAs. Epigenetic modifiers frequently formed biologically relevant clusters, including the cohesin complex, RNA Polymerase II transcription factors, and PRC2 complex members. These clustering behaviors remained consistent across all chromosomes, supported by entropy analysis and high Adjusted Rand Index scores, indicating robust cross-chromosomal similarity. Co-occurrence analysis further revealed specific sets of modifiers that consistently appeared together within clusters, reflecting shared biological functions and interactions. Validation using another dataset confirmed the reproducibility of these clustering patterns and modifier co-occurrence relationships, underscoring the reliability and generalizability of the methodology.

97 MATHEMATICS AND COMPUTING↗

Some approaches to optimal cluster labeling of aerospace imagery

Some approaches are presented to the problem of labeling clusters using information from a given set of labeled and unlabeled aerospace imagery patterns. The assignment of class labels to the clusters is formulated as the determination of the best assignment over all possible ones with respect to some criterion. Cluster labeling is also viewed as the probability of correct labeling with a maximization of likelihood function. Results of the application of these techniques in the processing of remotely sensed multispectral scanner imagery data are presented.

Chittineni, C. B.↗

Regional and Model-Specific Response Types in A Global Gridded Crop Model Ensemble

Crop models are often employed to project crop yields under changing conditions such as global warming and associated management change for adaptation. Multi-model ensembles are promoted to enhance the robustness of projections, but questions remain on what causes often large differences between projections of individual models. Global Gridded Crop Models (GGCMs) are especially exposed to this question when applied for assessing climate change impacts, adaptation, environmental impacts of agricultural production, because their results are used in downstream analyses, such as in integrated assessment or economic modeling for projecting future land-use change. Even though global gridded crop models are often based on detailed field-scale models or have implemented similar modeling principles in other ecosystem models, global-scale models are subject to substantial uncertainties from both model structure and parametrization as well as from calibration and input data quality. AgMIP’s Global Gridded Crop Model Intercomparison (GGCMI) has thus set out to intercompare GGCMs in order to evaluate model performance, describe model uncertainties, identify inconsistencies within the ensemble and underlying reasons, and to ultimately improve models and modeling capacities. In phase 2 of the GGCMI activities, 12 modeling groups followed a modeling protocol that asked for up to 1404 31-year global simulations at 0.5 arc-degree spatial resolution to assess models’ sensitivities to changes in carbon dioxide (C; 4 different levels) temperature (T; 7 different offset levels), water supply (W; 9 levels), and nitrogen (N; 3 levels), the so-called CTWN experiment (Franke et al. 2020; http://dx.doi.org/10.5194/gmd-13-2315-2020). We here present analyses of model response types using impact response surfaces along the C, T, W, and N dimensions, respectively and collectively. Doing so, we can understand differences in simulated responses per driver rather than aggregated changes in yields. We find that models’ sensitivities to the individual driver dimensions are substantially different and often more different across models than across regions. A cluster analysis finds regional and model-specific patterns. There is some agreement across models with respect to the spatial patterns of response types but strong differences in the distribution of response type clusters across models suggests that models need to undergo further scrutiny. We suggest establishing standards in model process evaluation not only against historical dynamics but also against dedicated experiments across the CTWN dimensions.

crop models↗

Analysis of multispectral data using an unsupervised classification technique: Application to VAS

A statistical classification method based on clustering of multidimensional histograms was applied to several channels of the VAS multispectral imagery. The method automatically discriminates and classifies atmospheric ground features such as cloud types, atmospheric moisture patterns, ocean, or ground. Such a clustering method has the advantage of forming natural data groupings, without a priori classification. Clusters are not limited by straight lines or plane surfaces as is the case in threshold methods. The method was applied to simultaneous full resolution images from channels 8 (11.2 micron), 10 (6.7 micron), and 12 (3.9 micron). Twenty image segments of 64 by 64, 12 image segments of 128 by 128, and 4 image segments of 254 by 254 picture elements were analyzed. In addition, normal VISSR mode images at 1800, 1830, and 2000 GMT were used to identify the classes. The gray levels measured along a scan line and the result of the classification scheme (dashed curves) for the three channels investigated are shown. Each point of the image is affected to a class. Each class is identified by a center of gravity that is represented by a vector in the three dimensional space of gray levels.

Szejwach, G.↗

Data-driven analysis and prediction of wastewater treatment plant performance: Insights and forecasting for sustainable operations

Here this study presents a comprehensive performance and forecasting analysis of the As-Samra wastewater treatment plant (WWTP) in Jordan, with two main objectives. Firstly, a thorough evaluation of the plant's performance is conducted. The analysis involves independently assessing historical operational conditions, plant production, and their statistical correlations using various statistical techniques. The second objective focuses on developing a data-driven forecasting approach to predict the plant's production one month in advance, using multiple machine learning models. The results highlight the effectiveness of principal component analysis (PCA) in simplifying operational data, revealing distinct operational clusters, and identifying seasonal production patterns while showing correlations between operational conditions and overall power production. The support vector machine (SVM) forecasting model emerged as the top performer, showcasing the potential of a hybrid forecasting approach. The findings offer valuable perspectives for enhancing operational efficiency, refining production planning, and ultimately improving the environmental impact of the plant.

42 ENGINEERING↗

Effects of warming on bacterial growth rates in a peat soil under ambient and elevated CO 2

Boreal peatlands are important global carbon reservoirs that are particularly vulnerable to predicted climate changes such as increasing CO 2 and temperature. Since microbial activities regulate the balance of carbon sequestered into soil organic matter or remineralized to CO 2 , characterizing their response to these environmental factors is critical to predicting how peatland ecosystems will affect climate-carbon cycle feedbacks. Here we examined in-situ taxon-specific variation in microbial growth under long-term elevated CO 2 and across a gradient of warming treatments in a northern Minnesota peat bog using quantitative stable isotope probing with 18 O-water. Across temperatures, bacterial taxa were grouped according to the excess atom fraction 18 O (EAF) of their genomes, a proxy for DNA replication and hence, growth. Taxon-specific growth across CO 2 and temperature treatments clustered into relatively few response patterns. While a large portion of taxa showed little to no growth under ambient CO 2 , many of the same taxa grew rapidly under elevated CO 2 . We found support for phylogenetic conservation of response patterns among Acidobacteria and Proteobacteria, the two most abundant phyla in our data. Our results suggest certain taxa may be primed for new climate conditions and have a greater influence on carbon cycling with implications for future climate mitigation strategies.

16S amplicon sequencing, Carbon Dioxide (CO2), pea↗

SchedInspector: A Batch Job Scheduling Inspector Using Reinforcement Learning

Improving the performance of job executions is an important goal of HPC batch job schedulers, such as minimizing job waiting time, slowdown, or completion time. Such a goal is often accomplished using carefully designed heuristics based on job features, such as job size and job duration. However, these heuristics overlook important runtime factors (e.g., cluster availability and waiting job patterns), which may vary across time and make a previously sound scheduling decision not hold any longer. In this study, we propose a new approach to incorporate runtime factors into batch job scheduling for better job execution performance. The key idea is to add a scheduling inspector on top of the base job scheduler to scrutinize its scheduling decisions. The inspector will take the runtime factors into consideration and accordingly determine the fitness of the scheduled job. It then either accepts the scheduled job or rejects it and asks the base schedulers to try again later. We realize such an inspector, namely SchedInspector, by leveraging the intelligence of reinforcement learning. Through extensive experiments, we show SchedInspector can intelligently integrate the runtime factors into various batch job scheduling policies, including the state-of-the-art one, to gain better job execution performance, such as smaller average bounded job slowdown (up to 69% better) or average job waiting time (up to 52% better), across various real-world workloads. We also show that although rejecting scheduling decisions may leave the resources idle hence affect the system utilization, SchedInspector is able to achieve the job execution performance improvement with marginal impact on the system utilization (typically less than 1%). We consider one key advantage of SchedInspector is it automatically learns to work with and improve existing job scheduling policies without changing them, which makes it promising to serve as a generic enhancer for various batch job scheduling policies.

Zhang, Di↗

Watershed zonation through hillslope clustering for tractably quantifying above- and below-ground watershed heterogeneity and functions

Abstract. In this study, we develop a watershed zonation approach for characterizing watershed organization and functions in a tractable manner by integrating multiple spatial data layers. We hypothesize that (1) a hillslope is an appropriate unit for capturing the watershed-scale heterogeneity of key bedrock-through-canopy properties and for quantifying the co-variability of these properties representing coupled ecohydrological and biogeochemical interactions, (2) remote sensing data layers and clustering methods can be used to identify watershed hillslope zones having the unique distributions of these properties relative to neighboring parcels, and (3) property suites associated with the identified zones can be used to understand zone-based functions, such as response to early snowmelt or drought and solute exports to the river. We demonstrate this concept using unsupervised clustering methods that synthesize airborne remote sensing data (lidar, hyperspectral, and electromagnetic surveys) along with satellite and streamflow data collected in the East River Watershed, Crested Butte, Colorado, USA. Results show that (1) we can define the scale of hillslopes at which the hillslope-averaged metrics can capture the majority of the overall variability in key properties (such as elevation, net potential annual radiation, and peak snow-water equivalent – SWE), (2) elevation and aspect are independent controls on plant and snow signatures, (3) near-surface bedrock electrical resistivity (top 20 m) and geological structures are significantly correlated with surface topography and plan species distribution, and (4) K-means, hierarchical clustering, and Gaussian mixture clustering methods generate similar zonation patterns across the watershed. Using independently collected data, we show that the identified zones provide information about zone-based watershed functions, including foresummer drought sensitivity and river nitrogen exports. The approach is expected to be applicable to other sites and generally useful for guiding the selection of hillslope-experiment locations and informing model parameterization.

58 GEOSCIENCES↗

Application of an automatic cloud tracking technique to Meteosat water vapor and infrared observations

The automatic cloud tracking system was applied to METEOSAT 6.7 micrometers water vapor measurements to learn whether the system can track the motions of water vapor patterns. Data for the midlatitudes, subtropics, and tropics were selected from a sequence of METEOSAT pictures for 25 April 1978. Trackable features in the water vapor patterns were identified using a clustering technique and the features were tracked by two different methods. In flat (low contrast) water vapor fields, the automatic motion computations were not reliable, but in areas where the water vapor fields contained small scale structure (such as in the vicinity of active weather phenomena) the computations were successful. Cloud motions were computed using METEOSAT infrared observations (including tropical convective systems and midlatitude jet stream cirrus).

Endlich, R. M.↗

Unsupervised Segmentation and Clustering Workflow for Efficient Processing of 4D-STEM and 5D-STEM Data

Four-dimensional scanning transmission electron microscopy (4D-STEM) enables mapping of diffraction information with nanometer-scale spatial resolution, offering detailed insight into local structure, orientation, and strain. However, as data dimensionality and sampling density increase, particularly for in situ scanning diffraction experiments (5D-STEM), robust segmentation of structurally consistent behavior across sequential measurements becomes essential for efficient and physically meaningful analysis. Here, we introduce a clustering framework that identifies crystallographically distinct domains from 4D-STEM datasets. By using local diffraction-pattern similarity as a metric, the method extracts closed contours delineating spatially contiguous regions. This approach produces cluster-averaged diffraction patterns that improve signal quality while reducing data volume by orders of magnitude, enabling rapid and accurate orientation, phase, and strain mapping. We demonstrate the applicability of this approach to in situ liquid-cell 4D-STEM data of gold nanoparticle growth. Our method provides a scalable and generalizable route for spatially coherent segmentation, data compression, and quantitative structure–strain mapping across diverse 4D-STEM modalities. The full analysis code and example workflows are publicly available to support reproducibility and reuse.

4D-STEM↗

Three dimensional cluster analysis for atom probe tomography using Ripley’s K-function and machine learning

The size and structure of spatial molecular and atomic clustering can significantly impact material properties and is therefore important to accurately quantify. Ripley’s K-function (K(r)), a measure of spatial correlation, can be used to perform such quantification when the material system of interest can be represented as a marked point pattern. This work demonstrates how machine learning models based on K (r)-derived metrics can accurately estimate cluster size and intra-cluster density in simulated three dimensional (3D) point patterns containing spherical clusters of varying size; over 90% of model estimates for cluster size and intra-cluster density fall within 11% and 18% error of the true values, respectively. These K (r)-based size and density estimates are then applied to an experimental APT reconstruction to characterize MgZn clusters in a 7000 series aluminum alloy. Here we find that the estimates are more accurate, consistent, and robust to user interaction than estimates from the popular maximum separation algorithm. Using K (r) and machine learning to measure clustering is an accurate and repeatable way to quantify this important material attribute.

36 MATERIALS SCIENCE↗

Timing based clustering in the Northern Finland Birth Cohorts 1966 and 1986 suggests two new patterns for childhood BMI curve [Poster]

Childhood body mass index (BMI) is a widely used measure of adiposity in children (<18 years of age). Children grow with individual tempo and individuals of the same age, or of the same BMI, might be in different phases in their individual growth curves. Variability between different childhood BMI curves can be separated in two components: phase variability (x-axis; time) and amplitude variability (y-axis; BMI). Phase variability can be thought of arising from differences in maturational age between individuals. This is related to the timing of peaks and valleys in a child’s BMI curve.

59 BASIC BIOLOGICAL SCIENCES↗