Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “pattern clustering”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Subseasonal Clustering of Atmospheric Rivers Over the Western United States

Abstract The serial occurrence of atmospheric rivers (ARs) along the US West Coast can lead to prolonged and exacerbated hydrologic impacts, threatening flood‐control and water‐supply infrastructure due to soil saturation and diminished recovery time between storms. Here a statistical approach for quantifying subseasonal temporal clustering among extreme events is applied to a 41‐year (1979–2019) wintertime AR catalog across the western United States (US). Observed AR occurrence, compared against a randomly distributed AR timeseries with the same average event density, reveals temporal clustering at a greater‐than‐random rate across the western US with a distinct geographical pattern. Compared to the Pacific Northwest, significant AR clusters over the northern Coastal Range of California and Sierra Nevada are more frequent and occur over longer time periods. Clusters along the California Coastal Range typically persist for 2 weeks, are composed of 4–5 ARs per cluster, and account for over 85% of total AR occurrence. Across the northwest Coast‐Cascade Ranges, clusters account for ∼50% of total AR occurrence, typically last 8–10 days, and contain 3–4 individual AR events. Based on precipitation data from a high‐resolution dynamical downscaling of reanalysis, the fractions of total and extreme hourly precipitation attributable to AR clusters are largest along the northern California coast and in the Sierra Nevada. Interannual variability among clusters highlights their importance for determining whether a particular water year is anomalously wet or dry. The mechanisms behind this unusual clustering are unclear and require further research.

Meteorology & Atmospheric Sciences↗

Collected Notes on the Workshop for Pattern Discovery in Large Databases

These collected notes are a record of material presented at the Workshop. The core data analysis is addressed that have traditionally required statistical or pattern recognition techniques. Some of the core tasks include classification, discrimination, clustering, supervised and unsupervised learning, discovery and diagnosis, i.e., general pattern discovery.

Buntine, Wray↗

Conserved unique peptide patterns (CUPP) online platform 2.0: implementation of +1000 JGI fungal genomes

Carbohydrate-processing enzymes, CAZymes, are classified into families based on sequence and three-dimensional fold. Because many CAZyme families contain members of diverse molecular function (different EC-numbers), sophisticated tools are required to further delineate these enzymes. Such delineation is provided by the peptide-based clustering method CUPP, Conserved Unique Peptide Patterns. CUPP operates synergistically with the CAZy family/subfamily categorizations to allow systematic exploration of CAZymes by defining small protein groups with shared sequence motifs. The updated CUPP library contains 21,930 of such motif groups including 3,842,628 proteins. The new implementation of the CUPP-webserver, https://cupp.info/, now includes all published fungal and algal genomes from the Joint Genome Institute (JGI), genome resources MycoCosm and PhycoCosm, dynamically subdivided into motif groups of CAZymes. This allows users to browse the JGI portals for specific predicted functions or specific protein families from genome sequences. Thus, a genome can be searched for proteins having specific characteristics. All JGI proteins have a hyperlink to a summary page which links to the predicted gene splicing including which regions have RNA support. The new CUPP implementation also includes an update of the annotation algorithm that uses only a fourth of the RAM while enabling multi-threading, providing an annotation speed below 1 ms/protein.

59 BASIC BIOLOGICAL SCIENCES↗

Comparing Synoptic Pattern Evolution for Flash‐Flood‐Producing and Non‐Flash‐Flood‐Producing Mesoscale Convective Systems in the United States

Understanding how the short-term evolution of synoptic weather patterns influence Mesoscale Convective Systems (MCSs) is essential, as these systems are responsible for over half of central U.S. flash floods, leading to substantial socioeconomic and water resource management impacts. This study analyzes long-term MCS data, flash flood reports, and atmospheric reanalyses from 2007 to 2017 using a machine learning clustering algorithm to examine how the synoptic weather patterns evolve prior to MCS initiation. While the clusters reflect seasonal and regional differences in MCS occurrence, they do not consistently distinguish between MCSs that do and do not produce flash floods. Systems in the southern Great Plains are more flood-prone when a synoptic-scale forcing, located near the system, drives strong water vapor transport from the nearby moisture source. More generally under different synoptic weather patterns, a broader precipitating area is the most dominant factor governing MCS flash flood potential.

atmospheric dynamics↗

Observations of the Bright Star in the Globular Cluster 47 Tucanae (NGC 104)

The Bright Star in the globular cluster 47 Tucanae (NGC 104) is a post-asymptotic giant branch (post-AGB) star of spectral type B8 III. The ultraviolet spectra of late-B stars exhibit myriad absorption features, many due to species unobservable from the ground. The Bright Star thus represents a unique window into the chemistry of 47 Tuc. We have analyzed observations obtained with the Far Ultraviolet Spectroscopic Explorer, the Cosmic Origins Spectrograph aboard the Hubble Space Telescope, and the Magellan Inamori Kyocera Echelle Spectrograph on the Magellan Telescope. By fitting these data with synthetic spectra, we determine various stellar parameters (T {sub eff} = 10,850 ± 250 K, logg=2.20±0.13) and the photospheric abundances of 26 elements, including Ne, P, Cl, Ga, Pd, In, Sn, Hg, and Pb, which have not previously been published for this cluster. Abundances of intermediate-mass elements (Mg through Ga) generally scale with Fe, while the heaviest elements (Pd through Pb) have roughly solar abundances. Its low C/O ratio indicates that the star did not undergo third dredge-up and suggests that its heavy elements were made by a previous generation of stars. If so, this pattern should be present throughout the cluster, not just in this star. Stellar-evolution models suggest that the Bright Star is powered by a He-burning shell, having left the AGB during or immediately after a thermal pulse. Its mass (0.54 ± 0.16M {sub ⊙}) implies that single stars in 47 Tuc lose 0.1–0.2 M {sub ⊙} on the AGB, only slightly less than they lose on the red giant branch.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Large-scale protein level comparison of Deltaproteobacteria reveals cohesive metabolic groups

Abstract Deltaproteobacteria, now proposed to be the phyla Desulfobacterota, Myxococcota, and SAR324, are ubiquitous in marine environments and play essential roles in global carbon, sulfur, and nutrient cycling. Despite their importance, our understanding of these bacteria is biased towards cultured organisms. Here we address this gap by compiling a genomic catalog of 1 792 genomes, including 402 newly reconstructed and characterized metagenome-assembled genomes (MAGs) from coastal and deep-sea sediments. Phylogenomic analyses reveal that many of these novel MAGs are uncultured representatives of Myxococcota and Desulfobacterota that are understudied. To better characterize Deltaproteobacteria diversity, metabolism, and ecology, we clustered ~1 500 genomes based on the presence/absence patterns of their protein families. Protein content analysis coupled with large-scale metabolic reconstructions separates eight genomic clusters of Deltaproteobacteria with unique metabolic profiles. While these eight clusters largely correspond to phylogeny, there are exceptions where more distantly related organisms appear to have similar ecological roles and closely related organisms have distinct protein content. Our analyses have identified previously unrecognized roles in the cycling of methylamines and denitrification among uncultured Deltaproteobacteria. This new view of Deltaproteobacteria diversity expands our understanding of these dominant bacteria and highlights metabolic abilities across diverse taxa.

Langwig, Marguerite V. (ORCID:0000000202472816)↗

Exploratory analysis of machine learning techniques in the Nevada geothermal play fairway analysis

Play fairway analysis (PFA) is commonly used to generate geothermal potential maps and guide exploration studies, with a particular focus on locating and characterizing blind geothermal systems. This study evaluates the application of machine learning techniques to PFA in the Great Basin region of Nevada. Following the evaluation of various techniques, we identified two approaches to PFA that produced promising results, 1) supervised Bayesian probabilistic neural networks to generate geothermal potential maps with confidence intervals, and 2) unsupervised principal component analysis paired with k-means clustering to generate both cluster maps to help identify spatial patterns, as well as new combined feature inputs. We applied these techniques to perform a comparative analysis between two principal sets of geological and geophysical features related to permeability and heat and a set of positive (known geothermal resources) and negative training sites (known drill sites with unsuitable geothermal conditions). We found that these methods constrain previously unrecognized feature controls on geothermal favorability, many of which are spatially organized within the extent of cluster groups and the major structural-hydrologic domains of the study area. Furthermore, we utilized exploratory unsupervised modeling to highlight spatial relationships between input data and predictive output results of our supervised modeling. As a result, we demonstrate how our models compare to the previous Nevada PFA and how the rapid insights these machine learning techniques offer may support future assessments of both known and undiscovered blind geothermal systems in the Great Basin region of Nevada and beyond.

15 GEOTHERMAL ENERGY↗

Understanding the Influence of Urban Form on the Spatial Pattern of Precipitation

Urban areas are known to modify the spatial pattern of precipitation climatology. Existing observational evidence suggests that precipitation can be enhanced downwind of a city. Among the proposed mechanisms, the thermodynamic and aerodynamic processes in the urban lower atmosphere interact with the meteorological conditions and can play a key role in determining the resulting precipitation patterns. In addition, these processes are influenced by urban form, such as the impervious surface extent. This study aims to unravel how different urban forms impact the spatial patterns of precipitation climatology under different meteorological conditions. We use the Multi-Radar Multi-Sensor quantitative precipitation estimation data products and analyze the hourly precipitation maps for 27 selected cities across the continental United States from the years 2015–2021 summer months. Results show that about 80% of the studied cities exhibit a statistically significant downwind enhancement of precipitation. Additionally, we find that the precipitation pattern tends to be more spatially clustered in intensity under higher wind speed; the location of radial precipitation maxima is located closer to the city center under low background winds but shifts downwind under high wind conditions. The magnitude of downwind precipitation enhancement is highly dependent on wind directions and is positively correlated with the city size for the south, southwest, and west directions. This study presents observational evidence through a cross-city analysis that the urban precipitation pattern can be influenced by the urban modification of atmospheric processes, providing insight into the mechanistic link between future urban land-use change and hydroclimates.

54 ENVIRONMENTAL SCIENCES↗

Isolating the Contributions of Specific Network Sites to the Diffuse Vibrational Spectrum of Interfacial Water with Isotopomer-Selective Spectroscopy of Cold Clusters

Decoding the structural information contained in the interfacial vibrational spectrum of water requires understanding how the spectral signatures of individual water molecules respond to their local hydrogen bonding environments. In this study, we isolated the contributions for the five classes of sites that differ according to the number of donor (D) and acceptor (A) hydrogen bonds that characterize each site. These patterns were measured by exploiting the unique properties of the water cluster cage structures formed in the gas phase upon hydration of a series of cations M+·(H2O)n (M = Li, Na, Cs, NH4, CH3NH3, H3O, and n = 5, 20-22). This selection of ions was chosen to systematically express the A, AD, AAD, ADD, and AADD hydrogen bonding motifs. The spectral signatures of each site were measured using two-color, IR-IR isotopomer-selective photofragmentation vibrational spectroscopy of the cryogenically cooled, mass selected cluster ions in which a single intact H2O is introduced without isotopic scrambling, an important advantage afforded by the cluster regime. The resulting patterns provide an unprecedented picture of the intrinsic line shapes and spectral complexities associated with excitation of the individual OH groups, as well as the correlation between the frequencies of the two OH groups on the same water molecule, as a function of network site. The properties of the surrounding water network that govern this frequency map are evaluated by dissecting electronic structure calculations that explore how changes in the nearby network structures, both within and beyond the first hydration shell, affect the local frequency of an OH oscillator. The qualitative trends are recovered with a simple model that correlates the OH frequency with the network-modulated local electron density in the center of the OH bond.

Yang, Nan↗

A Kinematic Perspective on the Formation Process of the Stellar Groups in the Rosette Nebula

Stellar kinematics is a powerful tool for understanding the formation process of stellar associations. Here, we present a kinematic study of the young stellar population in the Rosette nebula using recent Gaia data and high-resolution spectra. We first isolate member candidates using the published mid-infrared photometric data and the list of X-ray sources. A total of 403 stars with similar parallaxes and proper motions are finally selected as members. The spatial distribution of the members shows that this star-forming region is highly substructured. The young open cluster NGC 2244 in the center of the nebula has a pattern of radial expansion and rotation. We discuss its implication on the cluster formation, e.g., monolithic cold collapse or hierarchical assembly. On the other hand, we also investigate three groups located around the border of the H ii bubble. The western group seems to be spatially correlated with the adjacent gas structure, but their kinematics is not associated with that of the gas. The southern group does not show any systematic motion relative to NGC 2244. These two groups might be spontaneously formed in filaments of a turbulent cloud. The eastern group is spatially and kinematically associated with the gas pillar receding away from NGC 2244. This group might be formed by feedback from massive stars in NGC 2244. Our results suggest that the stellar population in the Rosette Nebula may form through three different processes: the expansion of stellar clusters, hierarchical star formation in turbulent clouds, and feedback-driven star formation.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

A comparison of land-use determinations using data from ERTS-1 and high altitude aircraft

A manual interpretation of ERTS-1 MSS system corrected imagery has been performed on a study area within the Houston Area Test Site to classify land use using the Level 1 categories proposed by the Department of the Interior. The two types of imagery used included: (1) black and white transparencies of each band enlarged to a scale of approximately 1:250,000 and (2) color transparencies composited from the computer compatible tapes using the film recorder on a multispectral data analysis station. The results of this interpretation have been compared with the 1970 land use inventory of HATS which was compiled using color ektachrome imagery from high altitude aircraft (scale 1:120,000). Urban data from the same scene was also analyzed using a computer-aided (clustering) technique. The resulting clusters, representing areas of similar content, were compared with existing land use patterns in Houston. A technique was developed to correlate the spectral clusters to specific urban features on aircraft imagery by the location of specific, high contrast objects in particular resolution elements. It was concluded that ERTS-1 data could be used to develop Level 1 and many Level 2 land use categories for regional inventories and perhaps to some degree on a local level.

Lundelius, M. A.↗

Galactic ArchaeoLogIcaL ExcavatiOns (GALILEO): I. An updated census of APOGEE N-rich giants across the Milky Way

We use the 17th data release of the second phase of the Apache Point Observatory Galactic Evolution Experiment (APOGEE-2) to provide a homogenous census of N-rich red giant stars across the Milky Way (MW). We report a total of 149 newly identified N-rich field giants toward the bulge, metal-poor disk, and halo of our Galaxy. They exhibit significant enrichment in their nitrogen abundance ratios ([N/Fe] ≳ +0.5), along with simultaneous depletions in their [C/Fe] abundance ratios ([C/Fe] < +0.15), and they cover a wide range of metallicities (-1.8 < [Fe/H] < -0.7). The final sample of candidate N-rich red giant stars with globular-cluster-like (GC-like) abundance patterns from the APOGEE survey includes a grand total of ~412 unique objects. These strongly N-enhanced stars are speculated to have been stripped from GCs based on their chemical similarities with these systems. Even though we have not found any strong evidence for binary companions or signatures of pulsating variability yet, we cannot rule out the possibility that some of these objects were members of binary systems in the past and/or are currently part of a variable system. In particular, the fact that we identify such stars among the field stars in our Galaxy provides strong evidence that the nucleosynthetic process(es) producing the anomalous [N/Fe] abundance ratios occurs over a wide range of metallicities. This may provide evidence either for or against the uniqueness of the progenitor stars to GCs and/or the existence of chemical anomalies associated with likely tidally shredded clusters in massive dwarf galaxies such as “Kraken/Koala”, Gaia-Enceladus-Sausage, among others, before or during their accretion by the MW. In conclusion, a dynamical analysis reveals that the newly identified N-rich stars exhibit a wide range of dynamical characteristics throughout the MW, indicating that they were produced in a variety of Galactic environments.

59 BASIC BIOLOGICAL SCIENCES↗

An analysis of pilot error-related aircraft accidents

A multidisciplinary team approach to pilot error-related U.S. air carrier jet aircraft accident investigation records successfully reclaimed hidden human error information not shown in statistical studies. New analytic techniques were developed and applied to the data to discover and identify multiple elements of commonality and shared characteristics within this group of accidents. Three techniques of analysis were used: Critical element analysis, which demonstrated the importance of a subjective qualitative approach to raw accident data and surfaced information heretofore unavailable. Cluster analysis, which was an exploratory research tool that will lead to increased understanding and improved organization of facts, the discovery of new meaning in large data sets, and the generation of explanatory hypotheses. Pattern recognition, by which accidents can be categorized by pattern conformity after critical element identification by cluster analysis.

Kowalsky, N. B.↗

Investigating Covalent Bonding in f-elements using Gas-phase Ion Chemistry

Introduction In the reprocessing of f-elements present in used nuclear fuels, a variety of diglycolamides (DGA’s) are used as extractants for actinide partitioning. In particular, the Actinide-Lanthanide Separation (ALSEP) process typically utilizes either the N,N,N’,N’-tetraoctyl diglycolamide (TODGA) or N,N,N',N'-tetra-2-ethylhexyl diglycolamide (T2EHDGA) extractant ligands following the partitioning of uranium and plutonium from used nuclear fuel. To better understand fundamental interactions in these processes, covalent bonding of several f-elements with diglycolamides, primarily TODGA, is investigated in the gas phase using nanospray ionization and a quadrupole time-of-flight mass spectrometer. Further, analysis of the identity and relative strength of the cluster is enabled by MS2 isolation and collision induced dissociation. Methods Metal ion cluster analysis was completed using a Bruker mircOTOF-Q II quadrupole time-of-flight mass spectrometer equipped with a CaptiveSpray nanospray ion source. Metal:ligand solutions were prepared as 30 µM europium nitrate, samarium nitrate, cerium nitrate, or holmium nitrate and 3 µM DGA in acetonitrile or a 50:50 mixture of acetonitrile: isopropanol. Cluster mass spectra and collision-induced dissociation experiments were conducted in positive mode. Preliminary data To examine the patterns and relative strength of lanthanide: DGA interactions, MS2 experiments were completed with each lanthanide species listed above. Preliminary analyses of samarium and europium TODGA clusters suggest several combinations of TODGA and nitrate forming. The samarium cluster experiments yielded Sm(TODGA)x clusters with a samarium:TODGA ratio of up to 1:7 able to be isolated and evidence of greater ratios present in the mass spectrum. This is surprising, as metal clusters are not expected to have a coordination space able to accommodate this many ligands as large as TODGA. MS2 experiments show that, at higher ratios and with sufficient collision energy, entire TODGA ligands are removed instead of being fragmented. These experiments show that a lower collision energy is required to remove ligands as the number of bound TODGA’s increases, suggesting that in larger clusters, ligands are more delicately complexed to the metal. In addition to Sm(TODGA)x, several clusters were observed with nitrate ions bound to the metal in addition to TODGA. With a single nitrate ion, clusters with up to six TODGA’s were able to be isolated. In a similar pattern to the samarium clusters with only TODGA, less collision energy is required to eliminate one or more TODGA’s with increasing size. MS2 experiments suggest clusters with one nitrate appear to be of an equivalent or greater stability to clusters which replace the nitrate with a TODGA, as more collision energy is required to remove a TODGA ligand. These species with one nitrate are also in a higher abundance than the equivalent TODGA only cluster. With two nitrate ions, only clusters with a single TODGA were able to be isolated. Analogous europium experiments resulted in very similar clusters. Ratios of up to 1:7 Eu:TODGA were able to be isolated, and clusters with one nitrate and up to five TODGAs were isolated. In clusters with two nitrate ions, one or two TODGA’s could also be bound to the metal. MS2 experiments suggested, similarly to samarium, that larger clusters required less collision energy to eliminate TODGA. Europium clusters with one nitrate are in greater abundance and are stronger than the equivalent cluster which replaces the nitrate with TODGA. Similar analysis with cerium and holmium is ongoing, as well as analysis with other DGA ligands to compare relative strengths of the lanthanide metals with various extractant ligands.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Improving the Accuracy of Clustering Electric Utility Net Load Data using Dynamic Time Warping

Identifying patterns in electric utility net load data in a time-series format is very useful in preparing the operation for next day. Machine learning algorithms have been used in other domains and those concepts are applied in this paper on real-world net load measurement data. Clustering is the practice of grouping data with similar characteristics as determined by the distance measure. The K-means clustering algorithm is utilized here with actual electric utility data. The paper uses the standard distance measure, Euclidean distance (ED), and compares its performance against the dynamic time warping (DTW) measure. An actual case study with real data is presented, and DTW distance measure-based method observed to result better accuracy compared to the ED based method for substation net load measurements predominantly with residential customers.

clustering↗

JGI Plant Gene Atlas: an updateable transcriptome resource to improve functional gene descriptions across the plant kingdom

Abstract Gene functional descriptions offer a crucial line of evidence for candidate genes underlying trait variation. Conversely, plant responses to environmental cues represent important resources to decipher gene function and subsequently provide molecular targets for plant improvement through gene editing. However, biological roles of large proportions of genes across the plant phylogeny are poorly annotated. Here we describe the Joint Genome Institute (JGI) Plant Gene Atlas, an updateable data resource consisting of transcript abundance assays spanning 18 diverse species. To integrate across these diverse genotypes, we analyzed expression profiles, built gene clusters that exhibited tissue/condition specific expression, and tested for transcriptional response to environmental queues. We discovered extensive phylogenetically constrained and condition-specific expression profiles for genes without any previously documented functional annotation. Such conserved expression patterns and tightly co-expressed gene clusters let us assign expression derived additional biological information to 64 495 genes with otherwise unknown functions. The ever-expanding Gene Atlas resource is available at JGI Plant Gene Atlas (https://plantgeneatlas.jgi.doe.gov) and Phytozome (https://phytozome.jgi.doe.gov/), providing bulk access to data and user-specified queries of gene sets. Combined, these web interfaces let users access differentially expressed genes, track orthologs across the Gene Atlas plants, graphically represent co-expressed genes, and visualize gene ontology and pathway enrichments.

59 BASIC BIOLOGICAL SCIENCES↗

Simulating water dynamics related to pedogenesis across space and time: Implications for four-dimensional digital soil mapping

Digital soil mapping (DSM) relies on machine-learning and geostatistics to represent soil property observations across space. DSM techniques are powerful but often empirical, being limited to the quality and density of point samples. Water dynamics are closely related to soil variability, and the physics that govern water movement are well known. Hydrological properties can hence be simulated by physical models through space and time, unveiling key characteristics about soils. We propose the use of hydrologic models to map soils across the surface (2D), depth (1D), and time (1D)–which provides a 4D approach to digital soil mapping (4DSM). The Distributed Hydrology Soil Vegetation Model (DHSVM) was applied to a watershed currently under pasture. Moisture sensors and wells were installed at different depths in the watershed on summit, sideslope and toeslope positions to validate the model. DHSVM simulations of soil moisture distribution and depth to saturation were performed during the hydrological year (October 2008-September 2009). Clusters of similar pixels based on soil moisture values were determined using Dynamic Time Warping (DTW) to align temporal data and K-means. Clustering was performed both seasonally and for the entire year. Temporal patterns simulated by DHSVM matched measurements given by moisture sensors and wells. Seasonal clusters differed from the annual cluster. Distinct clusters were observed for each season and with depth, showing that spatiotemporal soil variability is lost when statically assessing soils. Spatiotemporal clusters corroborated field observations of fragipan occurrence not explicitly spatially mapped by Soil Survey Geographic Database (SSURGO). If a connection can be made between water and soils, static and dynamic soil variability can be predicted using physically based hydrologic models. Hydrologic models can benefit soil mapping by enabling reliable 4D simulation of water dynamics, which are fundamental to soil variability and soil classification and directly relate to biological, physical and chemical soil processes not captured by typical soil sampling protocols.

54 ENVIRONMENTAL SCIENCES↗