Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “pattern clustering”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Different feeding patterns affect meat quality of Tibetan pigs associated with intestinal microbiota alterations

This study aimed to investigate the effects of different feeding patterns on meat quality, gut microbiota and its metabolites of Tibetan pigs. Tibetan pigs with similar body weight were fed the high energy diets ( HEP , 20 pigs) and the regular diets ( RFP , 20 pigs), and free-ranging Tibetan pigs ( FRP , 20 pigs) were selected as the reference. After 6 weeks of experiment, meat quality indexes of semitendinosus muscle ( SM ) and cecal microbiota were measured. The results of meat quality demonstrated that the shear force of pig SM in FRP group was higher than that in HEP and RFP groups ( p < 0.001); the pH-value of SM in HEP pigs was higher at 45 min ( p < 0.05) and lower at 24 h ( p < 0.01) after slaughter than that in FRP and RFP groups; the SM lightness ( L* value) of FRP pigs increased compared with RFP and HEP groups ( p < 0.001), while the SM redness ( a* value) of FRP pigs was higher than that of RFP group ( p < 0.05). The free fatty acid ( FA ) profile exhibited that the total FAs and unsaturated FAs of pig SM in HEP and RFP groups were higher than those in FRP group ( p < 0.05); the RFP pigs had more reasonable FA composition with higher n-3 polyunsaturated FAs ( PUFAs ) and lower n-6/n-3 PUFA ratio than HEP pigs ( p < 0.05). Based on that, we observed that Tibetan pigs fed high energy diets (HEP) had lower microbial α-diversity in cecum ( p < 0.05), and distinct feeding patterns exhibited a different microbial cluster. Simultaneously, the short-chain FA levels in cecum of FRP and RFP pigs were higher compared with HEP pigs ( p < 0.05). A total of 11 genera related to muscle lipid metabolism or meat quality, including Alistipes , Anaerovibrio , Acetitomaculun , etc., were identified under different feeding patterns ( p < 0.05). Spearman correlation analysis demonstrated that alterations of free FAs in SM were affected by the genera Prevotellaceae_NK3B31_group , Prevotellaceae UCG-003 and Christensenellaceae_R-7_group ( p < 0.05). Taken together, distinct feeding patterns affected meat quality of Tibetan pigs related to gut microbiota alterations.

Zhu, Yanbin↗

Correlating and Simulating Socio-Demographically Driven Residential End-Use Activity Schedules

Incorporating socio-demographic and behavioral considerations into decision-support tools is crucial for identifying gaps and addressing consumer needs to ensure reliable and affordable energy solutions. In energy simulation models, the correlation between socio-demographics and time-use behavior is not well-captured. Thus, we developed a large-scale simulation workflow to generate schedules for 10 residential activities across 24 population segments defined by age, income, and employment status. Using pre-pandemic 2015-2019 American Time Use Survey (ATUS) data, we used ANOVA to confirm the correlation between demographic factors and time use. We explored three k-modes clustering methods-backward, forward, and a new hybrid approach-to delineate the occupancy patterns based on demographics. Using the probability of cluster membership for each population segment and a time inhomogeneous Markov chain to generate activity transition probabilities for each cluster, we simulated 50,000 schedules per segment and validated them against the ATUS data. The hybrid method produced the most socio-demographically differentiated clusters while demonstrating comparable performance to other approaches, with an overall root mean square error of 0.12 for both weekday and weekend schedules. Thus, the hybrid method, where each cluster is dominated by certain demographic segments and occupancy patterns, offers more modeling versatility in terms of scenario analysis. The new workflow improves the socio demographic differentiation of energy consumption by considering differences in time use. This approach enables future research on demographically segmented time of use (TOU) energy consumption, including impacts of TOU utility bills and rate analysis, long-run marginal emissions, and energy retrofits.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Top-Down Fabrication of Atomic Patterns in Twisted Bilayer Graphene

Atomic-scale engineering typically involves bottom-up approaches, leveraging parameters such as temperature, partial pressures, and chemical affinity to promote spontaneous arrangement of atoms. These parameters are applied globally, resulting in atomic-scale features scattered probabilistically throughout the material. In a top-down approach, different regions of the material are exposed to different parameters, resulting in structural changes varying on the scale of the resolution. In this work, the application of global and local parameters is combined in an aberration-corrected scanning transmission electron microscope (STEM) to demonstrate atomic-scale precision patterning of atoms in twisted bilayer graphene. The focused electron beam is used to define attachment points for foreign atoms through the controlled ejection of carbon atoms from the graphene lattice. The sample environment is staged with nearby source materials such that the sample temperature can induce migration of the source atoms across the sample surface. Under these conditions, the electron-beam (top-down) enables carbon atoms in the graphene to be replaced spontaneously by diffusing adatoms (bottom-up). Using image-based feedback control, arbitrary patterns of atoms and atom clusters are attached to the twisted bilayer graphene with limited human interaction. Finally, the role of substrate temperature on adatom and vacancy diffusion is explored by first-principles simulations.

36 MATERIALS SCIENCE↗

Spatiotemporal pattern detection, generation, and computation with circuits

Abstract Implementations of neurons, delays, and synapse circuits are presented with simulations. These neural elements are used to create two small spiking neural networks, the Rate-Window and Order-Biased clusters, which are capable of detecting simple two-spike spatiotemporal patterns. A simple pattern detecting network (SPDN) is created by combining the Rate-Window and Order-Biased clusters, where clusters are small spiking neural networks, and its simple pattern detection ability is demonstrated in simulation. The SPDN is used to implement a complex pattern detecting network (CPDN) and its complex pattern detection ability is demonstrated in simulation. Methods for generating arbitrary spatiotemporal patterns are presented. The CPDN and spatiotemporal pattern generation methods are then used to implement a novel spatiotemporal computing paradigm based on detecting and responding to spatiotemporal symbols. A simulation of a spatiotemporal half adder is presented to demonstrate the computing paradigm.

97 - MATHEMATICS AND COMPUTING↗

Statistical relationships across epigenomes using large-scale hierarchical clustering

Recent advances in genomics and sequencing platforms have revolutionized our ability to create immense data sets, particularly for studying epigenetic regulation of gene expression. However, the avalanche of epigenomic data is difficult to parse for biological interpretation given nonlinear complex patterns and relationships. This attractive challenge in epigenomic data lends itself to machine learning for discerning infectivity and susceptibility. In this study, we explore over 3000 epigenomes of uninfected individuals and provide a framework to characterize the relationships among epigenetic modifiers, their modifiers, genetic loci, and specific immune cell types across all chromosomes using hierarchical clustering. Hierarchical clustering of epigenomic data revealed consistent epigenetic patterns across chromosomes, demonstrating that variation due to epigenetic modifiers is greater than variation between cell types. Gene Ontology and KEGG pathway analyses indicated significant enrichment of genes involved in chromatin remodeling, mRNA splicing, immune responses, and the regulation of microRNAs and snoRNAs. Epigenetic modifiers frequently formed biologically relevant clusters, including the cohesin complex, RNA Polymerase II transcription factors, and PRC2 complex members. These clustering behaviors remained consistent across all chromosomes, supported by entropy analysis and high Adjusted Rand Index scores, indicating robust cross-chromosomal similarity. Co-occurrence analysis further revealed specific sets of modifiers that consistently appeared together within clusters, reflecting shared biological functions and interactions. Validation using another dataset confirmed the reproducibility of these clustering patterns and modifier co-occurrence relationships, underscoring the reliability and generalizability of the methodology.

97 MATHEMATICS AND COMPUTING↗

Data-driven analysis and prediction of wastewater treatment plant performance: Insights and forecasting for sustainable operations

Here this study presents a comprehensive performance and forecasting analysis of the As-Samra wastewater treatment plant (WWTP) in Jordan, with two main objectives. Firstly, a thorough evaluation of the plant's performance is conducted. The analysis involves independently assessing historical operational conditions, plant production, and their statistical correlations using various statistical techniques. The second objective focuses on developing a data-driven forecasting approach to predict the plant's production one month in advance, using multiple machine learning models. The results highlight the effectiveness of principal component analysis (PCA) in simplifying operational data, revealing distinct operational clusters, and identifying seasonal production patterns while showing correlations between operational conditions and overall power production. The support vector machine (SVM) forecasting model emerged as the top performer, showcasing the potential of a hybrid forecasting approach. The findings offer valuable perspectives for enhancing operational efficiency, refining production planning, and ultimately improving the environmental impact of the plant.

42 ENGINEERING↗

Effects of warming on bacterial growth rates in a peat soil under ambient and elevated CO 2

Boreal peatlands are important global carbon reservoirs that are particularly vulnerable to predicted climate changes such as increasing CO 2 and temperature. Since microbial activities regulate the balance of carbon sequestered into soil organic matter or remineralized to CO 2 , characterizing their response to these environmental factors is critical to predicting how peatland ecosystems will affect climate-carbon cycle feedbacks. Here we examined in-situ taxon-specific variation in microbial growth under long-term elevated CO 2 and across a gradient of warming treatments in a northern Minnesota peat bog using quantitative stable isotope probing with 18 O-water. Across temperatures, bacterial taxa were grouped according to the excess atom fraction 18 O (EAF) of their genomes, a proxy for DNA replication and hence, growth. Taxon-specific growth across CO 2 and temperature treatments clustered into relatively few response patterns. While a large portion of taxa showed little to no growth under ambient CO 2 , many of the same taxa grew rapidly under elevated CO 2 . We found support for phylogenetic conservation of response patterns among Acidobacteria and Proteobacteria, the two most abundant phyla in our data. Our results suggest certain taxa may be primed for new climate conditions and have a greater influence on carbon cycling with implications for future climate mitigation strategies.

16S amplicon sequencing, Carbon Dioxide (CO2), pea↗

SchedInspector: A Batch Job Scheduling Inspector Using Reinforcement Learning

Improving the performance of job executions is an important goal of HPC batch job schedulers, such as minimizing job waiting time, slowdown, or completion time. Such a goal is often accomplished using carefully designed heuristics based on job features, such as job size and job duration. However, these heuristics overlook important runtime factors (e.g., cluster availability and waiting job patterns), which may vary across time and make a previously sound scheduling decision not hold any longer. In this study, we propose a new approach to incorporate runtime factors into batch job scheduling for better job execution performance. The key idea is to add a scheduling inspector on top of the base job scheduler to scrutinize its scheduling decisions. The inspector will take the runtime factors into consideration and accordingly determine the fitness of the scheduled job. It then either accepts the scheduled job or rejects it and asks the base schedulers to try again later. We realize such an inspector, namely SchedInspector, by leveraging the intelligence of reinforcement learning. Through extensive experiments, we show SchedInspector can intelligently integrate the runtime factors into various batch job scheduling policies, including the state-of-the-art one, to gain better job execution performance, such as smaller average bounded job slowdown (up to 69% better) or average job waiting time (up to 52% better), across various real-world workloads. We also show that although rejecting scheduling decisions may leave the resources idle hence affect the system utilization, SchedInspector is able to achieve the job execution performance improvement with marginal impact on the system utilization (typically less than 1%). We consider one key advantage of SchedInspector is it automatically learns to work with and improve existing job scheduling policies without changing them, which makes it promising to serve as a generic enhancer for various batch job scheduling policies.

Zhang, Di↗

Watershed zonation through hillslope clustering for tractably quantifying above- and below-ground watershed heterogeneity and functions

Abstract. In this study, we develop a watershed zonation approach for characterizing watershed organization and functions in a tractable manner by integrating multiple spatial data layers. We hypothesize that (1) a hillslope is an appropriate unit for capturing the watershed-scale heterogeneity of key bedrock-through-canopy properties and for quantifying the co-variability of these properties representing coupled ecohydrological and biogeochemical interactions, (2) remote sensing data layers and clustering methods can be used to identify watershed hillslope zones having the unique distributions of these properties relative to neighboring parcels, and (3) property suites associated with the identified zones can be used to understand zone-based functions, such as response to early snowmelt or drought and solute exports to the river. We demonstrate this concept using unsupervised clustering methods that synthesize airborne remote sensing data (lidar, hyperspectral, and electromagnetic surveys) along with satellite and streamflow data collected in the East River Watershed, Crested Butte, Colorado, USA. Results show that (1) we can define the scale of hillslopes at which the hillslope-averaged metrics can capture the majority of the overall variability in key properties (such as elevation, net potential annual radiation, and peak snow-water equivalent – SWE), (2) elevation and aspect are independent controls on plant and snow signatures, (3) near-surface bedrock electrical resistivity (top 20 m) and geological structures are significantly correlated with surface topography and plan species distribution, and (4) K-means, hierarchical clustering, and Gaussian mixture clustering methods generate similar zonation patterns across the watershed. Using independently collected data, we show that the identified zones provide information about zone-based watershed functions, including foresummer drought sensitivity and river nitrogen exports. The approach is expected to be applicable to other sites and generally useful for guiding the selection of hillslope-experiment locations and informing model parameterization.

58 GEOSCIENCES↗

Unsupervised Segmentation and Clustering Workflow for Efficient Processing of 4D-STEM and 5D-STEM Data

Four-dimensional scanning transmission electron microscopy (4D-STEM) enables mapping of diffraction information with nanometer-scale spatial resolution, offering detailed insight into local structure, orientation, and strain. However, as data dimensionality and sampling density increase, particularly for in situ scanning diffraction experiments (5D-STEM), robust segmentation of structurally consistent behavior across sequential measurements becomes essential for efficient and physically meaningful analysis. Here, we introduce a clustering framework that identifies crystallographically distinct domains from 4D-STEM datasets. By using local diffraction-pattern similarity as a metric, the method extracts closed contours delineating spatially contiguous regions. This approach produces cluster-averaged diffraction patterns that improve signal quality while reducing data volume by orders of magnitude, enabling rapid and accurate orientation, phase, and strain mapping. We demonstrate the applicability of this approach to in situ liquid-cell 4D-STEM data of gold nanoparticle growth. Our method provides a scalable and generalizable route for spatially coherent segmentation, data compression, and quantitative structure–strain mapping across diverse 4D-STEM modalities. The full analysis code and example workflows are publicly available to support reproducibility and reuse.

4D-STEM↗

Three dimensional cluster analysis for atom probe tomography using Ripley’s K-function and machine learning

The size and structure of spatial molecular and atomic clustering can significantly impact material properties and is therefore important to accurately quantify. Ripley’s K-function (K(r)), a measure of spatial correlation, can be used to perform such quantification when the material system of interest can be represented as a marked point pattern. This work demonstrates how machine learning models based on K (r)-derived metrics can accurately estimate cluster size and intra-cluster density in simulated three dimensional (3D) point patterns containing spherical clusters of varying size; over 90% of model estimates for cluster size and intra-cluster density fall within 11% and 18% error of the true values, respectively. These K (r)-based size and density estimates are then applied to an experimental APT reconstruction to characterize MgZn clusters in a 7000 series aluminum alloy. Here we find that the estimates are more accurate, consistent, and robust to user interaction than estimates from the popular maximum separation algorithm. Using K (r) and machine learning to measure clustering is an accurate and repeatable way to quantify this important material attribute.

36 MATERIALS SCIENCE↗

Timing based clustering in the Northern Finland Birth Cohorts 1966 and 1986 suggests two new patterns for childhood BMI curve [Poster]

Childhood body mass index (BMI) is a widely used measure of adiposity in children (<18 years of age). Children grow with individual tempo and individuals of the same age, or of the same BMI, might be in different phases in their individual growth curves. Variability between different childhood BMI curves can be separated in two components: phase variability (x-axis; time) and amplitude variability (y-axis; BMI). Phase variability can be thought of arising from differences in maturational age between individuals. This is related to the timing of peaks and valleys in a child’s BMI curve.

59 BASIC BIOLOGICAL SCIENCES↗

Subseasonal Clustering of Atmospheric Rivers Over the Western United States

Abstract The serial occurrence of atmospheric rivers (ARs) along the US West Coast can lead to prolonged and exacerbated hydrologic impacts, threatening flood‐control and water‐supply infrastructure due to soil saturation and diminished recovery time between storms. Here a statistical approach for quantifying subseasonal temporal clustering among extreme events is applied to a 41‐year (1979–2019) wintertime AR catalog across the western United States (US). Observed AR occurrence, compared against a randomly distributed AR timeseries with the same average event density, reveals temporal clustering at a greater‐than‐random rate across the western US with a distinct geographical pattern. Compared to the Pacific Northwest, significant AR clusters over the northern Coastal Range of California and Sierra Nevada are more frequent and occur over longer time periods. Clusters along the California Coastal Range typically persist for 2 weeks, are composed of 4–5 ARs per cluster, and account for over 85% of total AR occurrence. Across the northwest Coast‐Cascade Ranges, clusters account for ∼50% of total AR occurrence, typically last 8–10 days, and contain 3–4 individual AR events. Based on precipitation data from a high‐resolution dynamical downscaling of reanalysis, the fractions of total and extreme hourly precipitation attributable to AR clusters are largest along the northern California coast and in the Sierra Nevada. Interannual variability among clusters highlights their importance for determining whether a particular water year is anomalously wet or dry. The mechanisms behind this unusual clustering are unclear and require further research.

Meteorology & Atmospheric Sciences↗

Conserved unique peptide patterns (CUPP) online platform 2.0: implementation of +1000 JGI fungal genomes

Carbohydrate-processing enzymes, CAZymes, are classified into families based on sequence and three-dimensional fold. Because many CAZyme families contain members of diverse molecular function (different EC-numbers), sophisticated tools are required to further delineate these enzymes. Such delineation is provided by the peptide-based clustering method CUPP, Conserved Unique Peptide Patterns. CUPP operates synergistically with the CAZy family/subfamily categorizations to allow systematic exploration of CAZymes by defining small protein groups with shared sequence motifs. The updated CUPP library contains 21,930 of such motif groups including 3,842,628 proteins. The new implementation of the CUPP-webserver, https://cupp.info/, now includes all published fungal and algal genomes from the Joint Genome Institute (JGI), genome resources MycoCosm and PhycoCosm, dynamically subdivided into motif groups of CAZymes. This allows users to browse the JGI portals for specific predicted functions or specific protein families from genome sequences. Thus, a genome can be searched for proteins having specific characteristics. All JGI proteins have a hyperlink to a summary page which links to the predicted gene splicing including which regions have RNA support. The new CUPP implementation also includes an update of the annotation algorithm that uses only a fourth of the RAM while enabling multi-threading, providing an annotation speed below 1 ms/protein.

59 BASIC BIOLOGICAL SCIENCES↗

Comparing Synoptic Pattern Evolution for Flash‐Flood‐Producing and Non‐Flash‐Flood‐Producing Mesoscale Convective Systems in the United States

Understanding how the short-term evolution of synoptic weather patterns influence Mesoscale Convective Systems (MCSs) is essential, as these systems are responsible for over half of central U.S. flash floods, leading to substantial socioeconomic and water resource management impacts. This study analyzes long-term MCS data, flash flood reports, and atmospheric reanalyses from 2007 to 2017 using a machine learning clustering algorithm to examine how the synoptic weather patterns evolve prior to MCS initiation. While the clusters reflect seasonal and regional differences in MCS occurrence, they do not consistently distinguish between MCSs that do and do not produce flash floods. Systems in the southern Great Plains are more flood-prone when a synoptic-scale forcing, located near the system, drives strong water vapor transport from the nearby moisture source. More generally under different synoptic weather patterns, a broader precipitating area is the most dominant factor governing MCS flash flood potential.

atmospheric dynamics↗

Observations of the Bright Star in the Globular Cluster 47 Tucanae (NGC 104)

The Bright Star in the globular cluster 47 Tucanae (NGC 104) is a post-asymptotic giant branch (post-AGB) star of spectral type B8 III. The ultraviolet spectra of late-B stars exhibit myriad absorption features, many due to species unobservable from the ground. The Bright Star thus represents a unique window into the chemistry of 47 Tuc. We have analyzed observations obtained with the Far Ultraviolet Spectroscopic Explorer, the Cosmic Origins Spectrograph aboard the Hubble Space Telescope, and the Magellan Inamori Kyocera Echelle Spectrograph on the Magellan Telescope. By fitting these data with synthetic spectra, we determine various stellar parameters (T {sub eff} = 10,850 ± 250 K, logg=2.20±0.13) and the photospheric abundances of 26 elements, including Ne, P, Cl, Ga, Pd, In, Sn, Hg, and Pb, which have not previously been published for this cluster. Abundances of intermediate-mass elements (Mg through Ga) generally scale with Fe, while the heaviest elements (Pd through Pb) have roughly solar abundances. Its low C/O ratio indicates that the star did not undergo third dredge-up and suggests that its heavy elements were made by a previous generation of stars. If so, this pattern should be present throughout the cluster, not just in this star. Stellar-evolution models suggest that the Bright Star is powered by a He-burning shell, having left the AGB during or immediately after a thermal pulse. Its mass (0.54 ± 0.16M {sub ⊙}) implies that single stars in 47 Tuc lose 0.1–0.2 M {sub ⊙} on the AGB, only slightly less than they lose on the red giant branch.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Large-scale protein level comparison of Deltaproteobacteria reveals cohesive metabolic groups

Abstract Deltaproteobacteria, now proposed to be the phyla Desulfobacterota, Myxococcota, and SAR324, are ubiquitous in marine environments and play essential roles in global carbon, sulfur, and nutrient cycling. Despite their importance, our understanding of these bacteria is biased towards cultured organisms. Here we address this gap by compiling a genomic catalog of 1 792 genomes, including 402 newly reconstructed and characterized metagenome-assembled genomes (MAGs) from coastal and deep-sea sediments. Phylogenomic analyses reveal that many of these novel MAGs are uncultured representatives of Myxococcota and Desulfobacterota that are understudied. To better characterize Deltaproteobacteria diversity, metabolism, and ecology, we clustered ~1 500 genomes based on the presence/absence patterns of their protein families. Protein content analysis coupled with large-scale metabolic reconstructions separates eight genomic clusters of Deltaproteobacteria with unique metabolic profiles. While these eight clusters largely correspond to phylogeny, there are exceptions where more distantly related organisms appear to have similar ecological roles and closely related organisms have distinct protein content. Our analyses have identified previously unrecognized roles in the cycling of methylamines and denitrification among uncultured Deltaproteobacteria. This new view of Deltaproteobacteria diversity expands our understanding of these dominant bacteria and highlights metabolic abilities across diverse taxa.

Langwig, Marguerite V. (ORCID:0000000202472816)↗

Exploratory analysis of machine learning techniques in the Nevada geothermal play fairway analysis

Play fairway analysis (PFA) is commonly used to generate geothermal potential maps and guide exploration studies, with a particular focus on locating and characterizing blind geothermal systems. This study evaluates the application of machine learning techniques to PFA in the Great Basin region of Nevada. Following the evaluation of various techniques, we identified two approaches to PFA that produced promising results, 1) supervised Bayesian probabilistic neural networks to generate geothermal potential maps with confidence intervals, and 2) unsupervised principal component analysis paired with k-means clustering to generate both cluster maps to help identify spatial patterns, as well as new combined feature inputs. We applied these techniques to perform a comparative analysis between two principal sets of geological and geophysical features related to permeability and heat and a set of positive (known geothermal resources) and negative training sites (known drill sites with unsuitable geothermal conditions). We found that these methods constrain previously unrecognized feature controls on geothermal favorability, many of which are spatially organized within the extent of cluster groups and the major structural-hydrologic domains of the study area. Furthermore, we utilized exploratory unsupervised modeling to highlight spatial relationships between input data and predictive output results of our supervised modeling. As a result, we demonstrate how our models compare to the previous Nevada PFA and how the rapid insights these machine learning techniques offer may support future assessments of both known and undiscovered blind geothermal systems in the Great Basin region of Nevada and beyond.

15 GEOTHERMAL ENERGY↗