Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Bayesian clustering”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

AutoClass: A Bayesian Approach to Classification

We describe a Bayesian approach to the untutored discovery of classes in a set of cases, sometimes called finite mixture separation or clustering. The main difference between clustering and our approach is that we search for the "best" set of class descriptions rather than grouping the cases themselves. We describe our classes in terms of a probability distribution or density function, and the locally maximal posterior probability valued function parameters. We rate our classifications with an approximate joint probability of the data and functional form, marginalizing over the parameters. Approximation is necessitated by the computational complexity of the joint probability. Thus, we marginalize w.r.t. local maxima in the parameter space. We discuss the rationale behind our approach to classification. We give the mathematical development for the basic mixture model and describe the approximations needed for computational tractability. We instantiate the basic model with the discrete Dirichlet distribution and multivariant Gaussian density likelihoods. Then we show some results for both constructed and actual data.

Stutz, John↗

Fossil Signatures Using Elemental Abundance Distributions and Bayesian Probabilistic Classification

Elemental abundances (C6, N7, O8, Na11, Mg12, Al3, P15, S16, Cl17, K19, Ca20, Ti22, Mn25, Fe26, and Ni28) were obtained for a set of terrestrial fossils and the rock matrix surrounding them. Principal Component Analysis extracted five factors accounting for the 92.5% of the data variance, i.e. information content, of the elemental abundance data. Hierarchical Cluster Analysis provided unsupervised sample classification distinguishing fossil from matrix samples on the basis of either raw abundances or PCA input that agreed strongly with visual classification. A stochastic, non-linear Artificial Neural Network produced a Bayesian probability of correct sample classification. The results provide a quantitative probabilistic methodology for discriminating terrestrial fossils from the surrounding rock matrix using chemical information. To demonstrate the applicability of these techniques to the assessment of meteoritic samples or in situ extraterrestrial exploration, we present preliminary data on samples of the Orgueil meteorite. In both systems an elemental signature produces target classification decisions remarkably consistent with morphological classification by a human expert using only structural (visual) information. We discuss the possibility of implementing a complexity analysis metric capable of automating certain image analysis and pattern recognition abilities of the human eye using low magnification optical microscopy images and discuss the extension of this technique across multiple scales.

Hoover, Richard B.↗

Simultaneous global and local clustering in multiplex networks with covariate information

Understanding both global and layer-specific group structures is useful for uncovering complex patterns in networks with multiple interaction types. In this work, we introduce a new model, the hierarchical multiplex stochastic blockmodel, which simultaneously detects communities within individual layers of a multiplex network while inferring a global node clustering across the layers. A stochastic blockmodel is assumed in each layer, with probabilities of layer-level group memberships determined by a node’s global group assignment. Our model uses a Bayesian framework, employing a probit stick-breaking process to construct node-specific mixing proportions over a set of shared Griffiths–Engen–McCloseky distributions. These proportions determine layer-level community assignment, allowing for an unknown and varying number of groups across layers, while incorporating nodal covariate information to inform the global clustering. We propose a scalable variational inference procedure with parallelisable updates for application to large networks. Extensive simulation studies demonstrate our model’s ability to accurately recover both global and layer-level clusters in complicated settings, and applications to real data showcase the model’s effectiveness in uncovering interesting latent network structure.

community detection↗

Probabilistic Classification Using Elemental Abundance Distributions and Lossless Image Compression in Apollo 17 Lunar Dust Samples from Mare Serenitatis

We have previously outlined a strategy for the detection of fossils [Storrie-Lombardi and Hoover, 2004] and extant microbial life [Storrie-Lombaudi and Hoover, 20051 during robotic missions to Mars using co-registered structural and chemical signatures. Data inputs included image lossless compression indices to estimate relative textural complexity and elemental abundance distributions. Two exploratory classification algorithms (principal component analysis and hierarchical cluster analysis) provide an initial tentative classification of all targets. Nonlinear stochastic neural networks are then trained to produce a Bayesian estimate of algorithm classification accuracy. The strategy previously has been successful in distinguishing regions of biotic and abiotic alteration of basalt glass from unaltered samples. [Storrie-Lombardi and Fisk, 2004; Storrie-Lombardi and Fisk, 2004] Such investigations of abiotic versus biotic alteration of terrestrial mineralogy on Earth are compromised by .the difficulty finding mineralogy completely unaffected by the ubiquitous presence of microbial life on the planet. The renewed interest in lunar exploration offers an opportunity to investigate geological materials that may exhibit signs of aqueous alteration, but are highly unlikely to contain contaminating biological weathering signatures. We here present an extension of our earlier data set to include lunar dust samples obtained during the Apollo 17 mission. Apollo 17 landed in the Taurus-Littrow Valley in Mare Serenitatis. Most of the rock samples from this region of the lunar highlands are basalts comprised primarily of plagioclase and pyroxene and selected examples of orange and black volcanic glass. SEM images and elemental abundances (C6, N7, O8, Na11, Mg12, Al13, Si14, P15, S16, Cll7, K19, Ca20, Fe26) for a series of targets in the lunar dust samples are compared to the extant cyanobacteria, fossil trilobites, Orgueil meteorite, and terrestrial basalt targets previously discussed. The data set provides a first step in producing a quantitative probabilistic methodology for geobiological analysis of returned lunar samples or in situ exploration.

Storrie-Lombardi, Michael C.↗

Deep Galex Observations of the Coma Cluster: Source Catalog and Galaxy Counts

We present a source catalog from deep 26 ks GALEX observations of the Coma cluster in the far-UV (FUV; 1530 Angstroms) and near-UV (NUV; 2310 Angstroms) wavebands. The observed field is centered 0.9 deg. (1.6 Mpc) south-west of the Coma core, and has full optical photometric coverage by SDSS and spectroscopic coverage to r-21. The catalog consists of 9700 galaxies with GALEX and SDSS photometry, including 242 spectroscopically-confirmed Coma member galaxies that range from giant spirals and elliptical galaxies to dwarf irregular and early-type galaxies. The full multi-wavelength catalog (cluster plus background galaxies) is 80% complete to NUV=23 and FUV=23.5, and has a limiting depth at NUV=24.5 and FUV=25.0 which corresponds to a star formation rate of 10(exp -3) solar mass yr(sup -1) at the distance of Coma. The GALEX images presented here are very deep and include detections of many resolved cluster members superposed on a dense field of unresolved background galaxies. This required a two-fold approach to generating a source catalog: we used a Bayesian deblending algorithm to measure faint and compact sources (using SDSS coordinates as a position prior), and used the GALEX pipeline catalog for bright and/or extended objects. We performed simulations to assess the importance of systematic effects (e.g. object blends, source confusion, Eddington Bias) that influence source detection and photometry when using both methods. The Bayesian deblending method roughly doubles the number of source detections and provides reliable photometry to a few magnitudes deeper than the GALEX pipeline catalog. This method is also free from source confusion over the UV magnitude range studied here: conversely, we estimate that the GALEX pipeline catalogs are confusion limited at NUV approximately 23 and FUV approximately 24. We have measured the total UV galaxy counts using our catalog and report a 50% excess of counts across FUV=22-23.5 and NUV=21.5-23 relative to previous GALEX measurements, which is not attributed to cluster member galaxies. Our galaxy counts are a better match to deeper UV counts measured with HST.

Hammer, D.↗

Cluster characterization in atom probe tomography: Machine learning using multiple summary functions

In this work, we develop a machine learning-based method to characterize intracluster concentration (ρ c ), background concentration (ρ b ), clustering radius (r̄), and radius dispersity (δ r ) in simulated atom probe tomography data using multiple spatial statistics summary functions to train a Bayesian regularized neural network. Here, we build upon previous work that utilized Ripley’s K-function by incorporating additional features from nearest-neighbor spatial statistics summary functions to better characterize concentration-based metrics. The addition of nearest-neighbor based features allows for highly accurate estimates of ρ c and ρ b , both with 90% of the predictions within 4.0% of the real value; the root-mean-square errors are reduced by 81.5% and 92.8% from predictions using only K-function based features, respectively. Additionally, including these nearest-neighbor based features improves the ability to differentiate between r̄ and δ r .

36 MATERIALS SCIENCE↗

O 16 O 16 collisions at energies available at the BNL Relativistic Heavy Ion Collider and at the CERN Large Hadron Collider comparing α clustering versus substructure

Collisions of . light and heavy nuclei in relativistic heavy-ion collisions have been shown to be sensitive to nuclear structure. With a proposed 16 O 16 O run at the CERN Large Hadron Collider (LHC) and at the BNL Relativistic Heavy Ion Collider (RHIC) we study the potential for finding α clustering in 16 O. Here we use the state-of-the-art iEBE-VISHNU package with 16 O nucleonic configurations from ab initio nuclear lattice simulations. This setup was tuned using a Bayesian analysis on p Pb and PbPb systems. We find that the 16 O 16 O system always begins far from equilibrium and that at LHC and RHIC it approaches the regime of hydrodynamic applicability only at very late times. Finally, by taking ratios of flow harmonics we are able to find measurable differences between α-clustering, nucleonic, and subnucleonic degrees of freedom in the initial state.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Permutation-adapted complete and independent basis for atomic cluster expansion descriptors

Atomic cluster expansion (ACE) methods provide a systematic way to describe particle local environments of arbitrary body order. For practical applications it is often required that the basis of cluster functions be symmetrized with respect to rotations and permutations. Existing methodologies yield sets of symmetrized functions that are over-complete. These methodologies thus require an additional numerical procedure, such as singular value decomposition (SVD), to eliminate redundant functions. In this work, it is shown that analytical linear relationships for subsets of cluster functions may be derived using recursion and permutation properties of generalized Wigner symbols. From these relationships, subsets (blocks) of cluster functions can be selected such that, within each block, functions are guaranteed to be linearly independent. It is conjectured that this block-wise independent set of permutation-adapted rotation and permutation invariant (PA-RPI) functions forms a complete, independent basis for ACE. Along with the first analytical proofs of block-wise linear dependence of ACE cluster functions and other theoretical arguments, numerical results are offered to demonstrate this. The utility of the method is demonstrated in the development of an ACE interatomic potential for tantalum. Using the new basis functions in combination with Bayesian compressive sensing sparse regression, some high degree descriptors are observed to persist and help achieve high-accuracy models.

Angular momentum↗

As-built design specification for proportion estimate software subsystem

The Proportion Estimate Processor evaluates four estimation techniques in order to get an improved estimate of the proportion of a scene that is planted in a selected crop. The four techniques to be evaluated were provided by the techniques development section and are: (1) random sampling; (2) proportional allocation, relative count estimate; (3) proportional allocation, Bayesian estimate; and (4) sequential Bayesian allocation. The user is given two options for computation of the estimated mean square error. These are referred to as the cluster calculation option and the segment calculation option. The software for the Proportion Estimate Processor is operational on the IBM 3031 computer.

Obrien, S.↗

Portfolios in Stochastic Local Search: Efficiently Computing Most Probable Explanations in Bayesian Networks

Portfolio methods support the combination of different algorithms and heuristics, including stochastic local search (SLS) heuristics, and have been identified as a promising approach to solve computationally hard problems. While successful in experiments, theoretical foundations and analytical results for portfolio-based SLS heuristics are less developed. This article aims to improve the understanding of the role of portfolios of heuristics in SLS. We emphasize the problem of computing most probable explanations (MPEs) in Bayesian networks (BNs). Algorithmically, we discuss a portfolio-based SLS algorithm for MPE computation, Stochastic Greedy Search (SGS). SGS supports the integration of different initialization operators (or initialization heuristics) and different search operators (greedy and noisy heuristics), thereby enabling new analytical and experimental results. Analytically, we introduce a novel Markov chain model tailored to portfolio-based SLS algorithms including SGS, thereby enabling us to analytically form expected hitting time results that explain empirical run time results. For a specific BN, we show the benefit of using a homogenous initialization portfolio. To further illustrate the portfolio approach, we consider novel additive search heuristics for handling determinism in the form of zero entries in conditional probability tables in BNs. Our additive approach adds rather than multiplies probabilities when computing the utility of an explanation. We motivate the additive measure by studying the dramatic impact of zero entries in conditional probability tables on the number of zero-probability explanations, which again complicates the search process. We consider the relationship between MAXSAT and MPE, and show that additive utility (or gain) is a generalization, to the probabilistic setting, of MAXSAT utility (or gain) used in the celebrated GSAT and WalkSAT algorithms and their descendants. Utilizing our Markov chain framework, we show that expected hitting time is a rational function - i.e. a ratio of two polynomials - of the probability of applying an additive search operator. Experimentally, we report on synthetically generated BNs as well as BNs from applications, and compare SGSs performance to that of Hugin, which performs BN inference by compilation to and propagation in clique trees. On synthetic networks, SGS speeds up computation by approximately two orders of magnitude compared to Hugin. In application networks, our approach is highly competitive in Bayesian networks with a high degree of determinism. In addition to showing that stochastic local search can be competitive with clique tree clustering, our empirical results provide an improved understanding of the circumstances under which portfolio-based SLS outperforms clique tree clustering and vice versa.

Mengshoel, Ole J.↗

Advanced Statistical Methods in Spacecraft Flight Software Cost Estimation: Bayesian Regression and Nonlinear Principal Components Analysis to Support System Engineering in the Early Project Lifecycle

This paper provides an overview of the new features and model updates in the upcoming release of the NASA Analogy Software Cost Tool (ASCoT). ASCoT, hosted within the Online NASA Space Estimation Tools (ONSET) on the One NASA Cost Engineering (ONCE) Database, is a web-based tool that provides a suite of estimation tools to support early lifecycle NASA flight software cost analysis. In addition to the traditional parametric flight software costing method COCOMO II, ASCoT contains a Bayesian linear regression to predict total flight software development cost as a function of total spacecraft cost, as well as four analogic methods: k-Nearest Neighbors (kNN) and Clustering models to predict Effort (in work-months) and total source lines of code (SLOC). These methods are designed to work primarily with system-level inputs such as mission type (orbiter, lander, etc.), mission destination (Earth, Inner Planetary, etc.), and the number of instruments and deployables. Nonlinear principal components analysis (NLPCA) is performed to find the principal features of the data composed of both categorical and numerical variables and is necessary prior to defining our analogic methods. Sensitivity analyses and in- and out-of-sample model performance results are presented for the Bayesian CER and the analogic models.

Johnson, James K.↗

Exploratory analysis of machine learning techniques in the Nevada geothermal play fairway analysis

Play fairway analysis (PFA) is commonly used to generate geothermal potential maps and guide exploration studies, with a particular focus on locating and characterizing blind geothermal systems. This study evaluates the application of machine learning techniques to PFA in the Great Basin region of Nevada. Following the evaluation of various techniques, we identified two approaches to PFA that produced promising results, 1) supervised Bayesian probabilistic neural networks to generate geothermal potential maps with confidence intervals, and 2) unsupervised principal component analysis paired with k-means clustering to generate both cluster maps to help identify spatial patterns, as well as new combined feature inputs. We applied these techniques to perform a comparative analysis between two principal sets of geological and geophysical features related to permeability and heat and a set of positive (known geothermal resources) and negative training sites (known drill sites with unsuitable geothermal conditions). We found that these methods constrain previously unrecognized feature controls on geothermal favorability, many of which are spatially organized within the extent of cluster groups and the major structural-hydrologic domains of the study area. Furthermore, we utilized exploratory unsupervised modeling to highlight spatial relationships between input data and predictive output results of our supervised modeling. As a result, we demonstrate how our models compare to the previous Nevada PFA and how the rapid insights these machine learning techniques offer may support future assessments of both known and undiscovered blind geothermal systems in the Great Basin region of Nevada and beyond.

15 GEOTHERMAL ENERGY↗

The emergence and transmission dynamics of HIV-1 CRF07_BC in Mainland China

A total of 1155 partial pol gene sequences of human immunodeficiency virus (HIV)-1 CRF07_BC were sampled between 1997 and 2015, spanning 13 provinces in Mainland China and risk groups [heterosexual, injecting drug users (IDU), and men who have sex with men (MSM)] to investigate the evolution, adaptation, spatiotemporal and risk group dynamics, migration patterns, and protein structure of HIV-1 CRF07_BC. Due to the unequal distribution of sequences across time, location, and risk group in the complete dataset (‘full1155’), subsampling methods were used. Maximum-likelihood and Bayesian phylogenetic analysis as well as discrete trait analysis of geographical location and risk group were carried out. To study mutations of a cluster of HIV-1 CRF07_BC (CRF07-1), we performed a comparative analysis of this cluster to the other CRF07_BC sequences (‘backbone_295’) and mapped the mutations observed in the respective protein structure. Our findings showed that HIV-1 CRF07_BC most likely originated among IDU in Yunnan Province between October 1992 to July 1993 [95 per cent hightest posterior density (HPD): May 1989–August 1995] and that IDU in Yunnan Province and MSM in Guangdong Province likely served as the viral sources during the early and more recent spread in Mainland China. We also revealed that HIV-1 CRF07-1 has been spreading for roughly 20 years and continues to cause local transmission in Mainland China and worldwide. Overall, our study sheds light on the dynamics of HIV-1 CRF07_BC distribution patterns in Mainland China. Our research may also be useful in formulating public health policies aimed at controlling acquired immune deficiency syndrome in Mainland China and globally.

60 APPLIED LIFE SCIENCES↗

Selection function of clusters in Dark Energy Survey year 3 data from cross-matching with South Pole Telescope detections

Context. Galaxy clusters selected based on overdensities of galaxies in photometric surveys provide the largest cluster samples. However, modeling the selection function of such samples is complicated by noncluster members projected along the line of sight (projection effects) and the potential detection of unvirialized objects (contamination). Aims. We empirically constrained the magnitude of these effects by cross-matching galaxy clusters selected in the Dark Energy Survey data with the redMaPPer algorithm with significant detections in three South Pole Telescope surveys (SZ, pol-ECS, pol-500d). Methods. For matched clusters, we augmented the redMaPPer catalog with the SPT detection significance. For unmatched objects we used the SPT detection threshold as an upper limit on the SZe signature. Using a Bayesian population model applied to the collected multiwavelength data, we explored various physically motivated models to describe the relationship between observed richness and halo mass. Results. Our analysis reveals a clear preference for models with an additional skewed scatter component associated with projection effects over a purely log-normal scatter model. We rule out significant contamination by unvirialized objects at the high-richness end of the sample. While dedicated simulations offer a well-fitting calibration of projection effects, our findings suggest the presence of redshift-dependent trends that these simulations may not have captured. Our findings highlight that modeling the selection function of optically detected clusters remains a complicated challenge that requires a combination of simulation and data-driven approaches.

79 ASTRONOMY AND ASTROPHYSICS↗

New Determination of Fundamental Properties of Palomar 5 Using Deep DESI Imaging Data

The legacy imaging surveys for the Dark Energy Spectroscopic Instrument project provides multiple-color photometric data, which are about 2 mag deeper than those from the SDSS. In this study, we redetermine the fundamental properties for an old halo globular cluster of Palomar 5 based on these new imaging data, including structure parameters, stellar population parameters, and luminosity and mass functions. These characteristics, together with its tidal tails, are key for dynamical studies of the cluster and constraining the mass model of the Milky Way. By fitting the King model to the radial surface density profile of Palomar 5, we derive the core radius of [[Formula]], tidal radius of [[Formula]], and concentration parameter of c = 0.78 ± 0.04. We apply a Bayesian analysis method to derive the stellar population properties and get an age of 11.508 ± 0.027 Gyr, metallicity of [Fe/H] = −1.798 ± 0.014, reddening of E(B − V) = 0.0552 ± 0.0005, and distance modulus of (m−M){sub 0} = 16.835±0.006. The main-sequence luminosity and mass functions for both the cluster center and tidal tails are investigated. The luminosity and mass functions at different distances from the cluster center suggest that there is obvious spatial mass segregation. Many faint low-mass stars have been evaporated at the cluster center, and the tidal tails are enhanced by low-mass stars. Both the concentration and relaxation times suggest that Palomar 5 is a totally relaxed system.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Deducing subnanometer cluster size and shape distributions of heterogeneous supported catalysts

Abstract Infrared (IR) spectra of adsorbate vibrational modes are sensitive to adsorbate/metal interactions, accurate, and easily obtainable in-situ or operando. While they are the gold standards for characterizing single-crystals and large nanoparticles, analogous spectra for highly dispersed heterogeneous catalysts consisting of single-atoms and ultra-small clusters are lacking. Here, we combine data-based approaches with physics-driven surrogate models to generate synthetic IR spectra from first-principles. We bypass the vast combinatorial space of clusters by determining viable, low-energy structures using machine-learned Hamiltonians, genetic algorithm optimization, and grand canonical Monte Carlo calculations. We obtain first-principles vibrations on this tractable ensemble and generate single-cluster primary spectra analogous to pure component gas-phase IR spectra. With such spectra as standards, we predict cluster size distributions from computational and experimental data, demonstrated in the case of CO adsorption on Pd/CeO 2 (111) catalysts, and quantify uncertainty using Bayesian Inference. We discuss extensions for characterizing complex materials towards closing the materials gap.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Measurement of $\nu_\mu$ CC Interactions With Two-Proton Final State in MINERvA

This dissertation presents a measurement of charged–current (CC) muon–neutrino interactions with exactly two protons and no pions in the final state (CC~$2p\,0\pi$), using data collected by the MINERvA detector in the NuMI medium–energy beam at Fermilab. Such two–proton topologies are a sensitive probe of nuclear dynamics in the few–GeV regime, including multi–nucleon correlations (npnh, notably $2p2h$) and intranuclear final–state interactions (FSI) such as pion absorption and nucleon rescattering. A precise experimental characterization of these processes is essential both for neutrino–interaction theory and for reducing systematic uncertainties in oscillation experiments that rely on accurate modeling of neutrino–nucleus interactions. Events are selected by requiring a $\nu_\mu$ CC interaction with a reconstructed $\mu^-$ and two proton tracks originating from a common vertex in MINERvA’s finely segmented scintillator tracker, with no reconstructed mesons. Muon charge and momentum are constrained by matching to the MINOS Near Detector, while proton identification exploits energy–loss profiles and stopping–proton features. Backgrounds from pion–producing channels that enter the signal region through FSI or reconstruction effects are constrained with data–driven sidebands (Michel–electron and isolated–cluster “blob” samples) and tuned via a simultaneous fit across signal and sideband regions. To correct detector resolution and acceptance effects, the analysis employs iterative Bayesian unfolding with extensive validation: statistical pseudo–experiments, and robustness checks against generator systematic “universes” and additional strong shape warps. Single–differential cross sections are reported for three observables tailored to the two–proton final state: the opening–angle cosine $\cos\!\left(\theta_{pp}\right)$, the leading–proton momentum, and the subleading–proton momentum. Systematic uncertainties include contributions from neutrino flux, interaction modeling (e.g., npnh and resonance parameters, pion FSI), and detector response (calibration, reconstruction efficiencies). The resulting distributions provide targeted constraints on the interplay of multi–nucleon dynamics and FSI that shape CC~$2p\,0\pi$ final states on hydrocarbon. Comparisons to modern GENIE–based simulations highlight kinematic regions where model components require refinement. These measurements thus inform generator tuning and improve the reliability of neutrino–energy reconstruction strategies for current and future long–baseline oscillation programs.

Syrotenko, Vladyslav S. [Tufts U.]↗

Description of Pegethrix niliensis sp. nov., a Novel Cyanobacterium from the Nile River Basin, Egypt: A Polyphasic Analysis and Comparative Study of Related Genera in the Oculatellales Order

In this paper, we examine the filamentous cyanobacterial strain NILCB16 and describe it as a new species within the genus Pegethrix. The original population was sampled from a mat growing in an irrigation canal in the Nile River, Egypt. Initially classified under Plectonema or Planktolyngbya, the strain is a potential producer of the toxins microcystin and β-N-Methylamino-L-Alanine (BMAA). Additionally, we reviewed the taxonomic relationships between the Oculatellales genera. To describe the new species, we conducted a polyphasic study, encompassing 16S rRNA gene phylogenetic analyses performed using both Maximum Likelihood and Bayesian methods, sequence identity (p-distance) analysis, 16S-23S ITS secondary structures, and morphological and habitat comparisons. The phylogenetic analysis revealed that strain NILCB16 clustered within the Pegethrix clade with strong phylogenetic support, but in a distinct position from other species in the genus. The strain shared a maximum 16S rRNA gene identity of 97.3% with P. qiandaoensis and 96.1% with the type species, P. bostrychoides. Morphologically, NILCB16 can be differentiated from other species in the genus by its lack of false branching. Our phylogenetic analyses also show that Pegethrix, Cartusia, Elainella, and Maricoleus are clustered with strong phylogenetic support. They exhibit high 16S rRNA gene identity and are morphologically indistinguishable, suggesting they could potentially be merged into a single genus in the future.

Hentschke, Guilherme Scotta (ORCID:000000034396024↗