Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Bayesian clustering”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Probabilisitc Geobiological Classification Using Elemental Abundance Distributions and Lossless Image Compression in Recent and Modern Organisms

Last year we presented techniques for the detection of fossils during robotic missions to Mars using both structural and chemical signatures[Storrie-Lombardi and Hoover, 2004]. Analyses included lossless compression of photographic images to estimate the relative complexity of a putative fossil compared to the rock matrix [Corsetti and Storrie-Lombardi, 2003] and elemental abundance distributions to provide mineralogical classification of the rock matrix [Storrie-Lombardi and Fisk, 2004]. We presented a classification strategy employing two exploratory classification algorithms (Principal Component Analysis and Hierarchical Cluster Analysis) and non-linear stochastic neural network to produce a Bayesian estimate of classification accuracy. We now present an extension of our previous experiments exploring putative fossil forms morphologically resembling cyanobacteria discovered in the Orgueil meteorite. Elemental abundances (C6, N7, O8, Na11, Mg12, Ai13, Si14, P15, S16, Cl17, K19, Ca20, Fe26) obtained for both extant cyanobacteria and fossil trilobites produce signatures readily distinguishing them from meteorite targets. When compared to elemental abundance signatures for extant cyanobacteria Orgueil structures exhibit decreased abundances for C6, N7, Na11, All3, P15, Cl17, K19, Ca20 and increases in Mg12, S16, Fe26. Diatoms and silicified portions of cyanobacterial sheaths exhibiting high levels of silicon and correspondingly low levels of carbon cluster more closely with terrestrial fossils than with extant cyanobacteria. Compression indices verify that variations in random and redundant textural patterns between perceived forms and the background matrix contribute significantly to morphological visual identification. The results provide a quantitative probabilistic methodology for discriminating putatitive fossils from the surrounding rock matrix and &om extant organisms using both structural and chemical information. The techniques described appear applicable to the geobiological analysis of meteoritic samples or in situ exploration of the Mars regolith. Keywords: cyanobacteria, microfossils, Mars, elemental abundances, complexity analysis, multifactor analysis, principal component analysis, hierarchical cluster analysis, artificial neural networks, paleo-biosignatures

Storrie-Lombardi, Michael C.↗

Evaluation of a pair-based, joint-likelihood association approach for regional infrasound event identification

SUMMARY A Bayesian framework for the association of infrasonic detections is presented and evaluated for analysis at regional propagation scales. A pair-based, joint-likelihood association approach is developed that identifies events by computing the probability that individual detection pairs are attributable to a hypothetical common source and applying hierarchical clustering to identify events from the pair-based analysis. The framework is based on a Bayesian formulation introduced for infrasonic source localization and utilizes the propagation models developed for that application with modifications to improve the numerical efficiency of the analysis. Clustering analysis is completed using hierarchical analysis via weighted linkage for a non-Euclidean distance matrix defined by the negative log-joint-likelihood values. The method is evaluated using regional synthetic data with propagation distances of hundreds of kilometres in order to study the sensitivity of the method to uncertainties and errors in backazimuth and time of arrival. The method is found to be robust and stable for typical uncertainties, able to effectively distinguish noise detections within the data set from those in events, and can be made numerically efficient due to its ease of parallelization.

58 GEOSCIENCES↗

Understanding the Scalability of Bayesian Network Inference using Clique Tree Growth Curves

Bayesian networks (BNs) are used to represent and efficiently compute with multi-variate probability distributions in a wide range of disciplines. One of the main approaches to perform computation in BNs is clique tree clustering and propagation. In this approach, BN computation consists of propagation in a clique tree compiled from a Bayesian network. There is a lack of understanding of how clique tree computation time, and BN computation time in more general, depends on variations in BN size and structure. On the one hand, complexity results tell us that many interesting BN queries are NP-hard or worse to answer, and it is not hard to find application BNs where the clique tree approach in practice cannot be used. On the other hand, it is well-known that tree-structured BNs can be used to answer probabilistic queries in polynomial time. In this article, we develop an approach to characterizing clique tree growth as a function of parameters that can be computed in polynomial time from BNs, specifically: (i) the ratio of the number of a BN's non-root nodes to the number of root nodes, or (ii) the expected number of moral edges in their moral graphs. Our approach is based on combining analytical and experimental results. Analytically, we partition the set of cliques in a clique tree into different sets, and introduce a growth curve for each set. For the special case of bipartite BNs, we consequently have two growth curves, a mixed clique growth curve and a root clique growth curve. In experiments, we systematically increase the degree of the root nodes in bipartite Bayesian networks, and find that root clique growth is well-approximated by Gompertz growth curves. It is believed that this research improves the understanding of the scaling behavior of clique tree clustering, provides a foundation for benchmarking and developing improved BN inference and machine learning algorithms, and presents an aid for analytical trade-off studies of clique tree clustering using growth curves.

Mengshoel, Ole Jakob↗

Implementation of a practical Markov chain Monte Carlo sampling algorithm in PyBioNetFit

Abstract Summary Bayesian inference in biological modeling commonly relies on Markov chain Monte Carlo (MCMC) sampling of a multidimensional and non-Gaussian posterior distribution that is not analytically tractable. Here, we present the implementation of a practical MCMC method in the open-source software package PyBioNetFit (PyBNF), which is designed to support parameterization of mathematical models for biological systems. The new MCMC method, am, incorporates an adaptive move proposal distribution. For warm starts, sampling can be initiated at a specified location in parameter space and with a multivariate Gaussian proposal distribution defined initially by a specified covariance matrix. Multiple chains can be generated in parallel using a computer cluster. We demonstrate that am can be used to successfully solve real-world Bayesian inference problems, including forecasting of new Coronavirus Disease 2019 case detection with Bayesian quantification of forecast uncertainty. Availability and implementation PyBNF version 1.1.9, the first stable release with am, is available at PyPI and can be installed using the pip package-management system on platforms that have a working installation of Python 3. PyBNF relies on libRoadRunner and BioNetGen for simulations (e.g. numerical integration of ordinary differential equations defined in SBML or BNGL files) and Dask.Distributed for task scheduling on Linux computer clusters. The Python source code can be freely downloaded/cloned from GitHub and used and modified under terms of the BSD-3 license (https://github.com/lanl/pybnf). Online documentation covering installation/usage is available (https://pybnf.readthedocs.io/en/latest/). A tutorial video is available on YouTube (https://www.youtube.com/watch?v=2aRqpqFOiS4&t=63s). Supplementary information Supplementary data are available at Bioinformatics online.

59 BASIC BIOLOGICAL SCIENCES↗

Noise-aware optimization in nominally identical manufacturing and measuring systems for high-throughput parallel workflows

Device-to-device variability in experimental noise critically impacts reproducibility, especially in automated, high-throughput systems like additive manufacturing farms. While manageable in small labs, such variability can escalate into serious risks at larger scales, such as architectural 3D printing, where noise may cause structural or economic failures. This contribution presents a noise-aware decision-making algorithm that quantifies and models device-specific noise profiles to manage variability adaptively. It uses distributional analysis and pairwise divergence metrics with clustering to choose between single-device and robust multi-device Bayesian optimization strategies. Unlike conventional methods that assume homogeneous devices or enforce generic robustness, the proposed framework explicitly determines whether shared optimization across devices is appropriate based on the degree of inter-device noise heterogeneity. This enables improved performance, reproducibility, and efficiency. An experimental case study involving three nominally identical 3D printers (same brand, model, and close serial numbers) demonstrates reduced redundancy, lower resource usage, and improved reliability, along with improved convergence stability and solution quality through the selection of the appropriate optimization strategy based on the degree of inter-device noise heterogeneity. Overall, this framework establishes a general approach for precision- and resource-aware optimization in scalable, automated experimental platforms, demonstrated here on a representative multi-device 3D printing case study.

Schenk, Christina↗

Thermodynamically informed priors for uncertainty propagation in first-principles statistical mechanics

Here, this work demonstrates how first-principles statistical mechanics approaches within a Bayesian framework can quantify and propagate uncertainties to downstream thermodynamic calculations. To address the issue of Bayesian prior selection, knowledge of 0 K ground states in the material system of interest is incorporated into the prior. The effectiveness of this framework is shown by creating a phase diagram for the fcc zirconium nitride system, including confidence intervals on order-disorder transition temperatures.

Bayesian methods↗

Deep-learning-enabled Bayesian inference of fuel magnetization in magnetized liner inertial fusion

Fuel magnetization in magneto-inertial fusion (MIF) experiments improves charged burn product confinement, reducing requirements on fuel areal density and pressure to achieve self-heating. By elongating the path length of 1.01 MeV tritons produced in a pure deuterium fusion plasma, magnetization enhances the probability for deuterium–tritium reactions producing 11.8-17.1 MeV neutrons. Nuclear diagnostics thus enable a sensitive probe of magnetization. Characterization of magnetization, including uncertainty quantification, is crucial for understanding the physics governing target performance in MIF platforms, such as magnetized liner inertial fusion (MagLIF) experiments conducted at Sandia National Laboratories, Z-facility. We demonstrate a deep-learned surrogate of a physics-based model of nuclear measurements. A single model evaluation is reduced from O ( 10 – 100 ) (10–100) CPU hours on a high-performance computing cluster down to O ( 10 – 100 ) (10) ms on a laptop. This enables a Bayesian inference of magnetization, rigorously accounting for uncertainties from surrogate modeling and noisy nuclear measurements. The approach is validated by testing on synthetic data and comparing with a previous study. We analyze a series of MagLIF experiments systematically varying preheat, resulting in the first ever systematic experimental study of magnetic confinement properties of the fuel plasma as a function of fundamental inputs on any neutron-producing MIF platform. We demonstrate that magnetization decreases from B R ∼ 0.2 ~0.5 to B R ∼ 0.2 ~0.2 MG cm as laser preheat energy deposited increases from E preheat ∼ 460 ~460 J to E preheat ∼ 460 ~1.4 kJ. This trend is consistent with 2D LASNEX simulations showing Nernst advection of the magnetic field out of the hot fuel and diffusion into the target liner.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

ELG×LRG Distribution through Dark Matter Halo Dynamics

We investigate the clustering and halo occupation distribution (HOD) of DESI Y1 emission-line (ELGs) and luminous red (LRGs) galaxies at 0.8 < z < 1.1, including their cross-correlation (ELG×LRG), using the A BACUS S UMMIT suite and a new Halo Occupation Model (H OME ) for galaxy multitracers. This integrates intrahalo dynamics, halo exclusion, and quenching, bridging insights from hydrodynamical, HOD, abundance-matching, and semianalytic studies. Leveraging full phase-space information from the Uchuu N-body simulation, and sampling satellites from dark-matter particle positions via physically motivated prescriptions, Home reproduces the anisotropic clustering down to s = 200 h −1 kpc with unprecedented accuracy. Model parameters are inferred solely from two-point statistics using a two-level Bayesian framework, yielding high-fidelity ELG, LRG, and cross-reference catalogs. We find that satellite ELGs behave as incoherent flows within their parent halos, dominating the clustering below 4 h −1 Mpc. The HOD from the best-fit Home has the following properties: (i) 90.50% (85.91%) of ELGs (LRGs) are central galaxies without satellites, residing in halos of M vir ∼ 6.6 × 10 11 (1.2 × 10 13 ) h −1 M ⊙ ; (ii) the ELG×LRG cross-correlation is governed by central-central pairs and shaped by halo exclusion on 2–5 h −1 Mpc scales; (iii) 9.50% (14.09%) of ELGs (LRGs) are satellites, of which 1.09% (3.52%) inhabit halos with a central galaxy of the same species in a maximally conformal configuration, 7.02% (0.005%) orbit complementary hosts in a minimally conformal state, and 0.58% (10.57%) are orphans. The high sensitivity of Home precisely captures the dynamics of satellites in different host environments, opening a promising avenue for understanding systematics and the dynamical nature of dark matter, potentially distinguishing gravity models.

Favole, Ginevra [Universidad de La Laguna (Spain);↗

Uncertainty-aware molecular dynamics from Bayesian active learning for phase transformations and thermal transport in SiC

Abstract Machine learning interatomic force fields are promising for combining high computational efficiency and accuracy in modeling quantum interactions and simulating atomistic dynamics. Active learning methods have been recently developed to train force fields efficiently and automatically. Among them, Bayesian active learning utilizes principled uncertainty quantification to make data acquisition decisions. In this work, we present a general Bayesian active learning workflow, where the force field is constructed from a sparse Gaussian process regression model based on atomic cluster expansion descriptors. To circumvent the high computational cost of the sparse Gaussian process uncertainty calculation, we formulate a high-performance approximate mapping of the uncertainty and demonstrate a speedup of several orders of magnitude. We demonstrate the autonomous active learning workflow by training a Bayesian force field model for silicon carbide (SiC) polymorphs in only a few days of computer time and show that pressure-induced phase transformations are accurately captured. The resulting model exhibits close agreement with both ab initio calculations and experimental measurements, and outperforms existing empirical models on vibrational and thermal properties. The active learning workflow readily generalizes to a wide range of material systems and accelerates their computational understanding.

36 MATERIALS SCIENCE↗

Enhancing DESI DR1 full-shape analyses using HOD-informed priors

We present an analysis of DESI Data Release 1 (DR1) that incorporates Halo Occupation Distribution (HOD)-informed priors into Full-Shape (FS) modeling of the power spectrum based on cosmological perturbation theory (PT). By leveraging physical insights from the galaxy-halo connection, these HOD-informed priors on nuisance parameters substantially mitigate projection effects in extended cosmological models that allow for dynamical dark energy. The resulting credible intervals now encompass the posterior maximum from the baseline analysis using gaussian priors, eliminating a significant posterior shift observed in baseline studies. In the ΛCDM framework, a combined DESI DR1 FS information and constraints from the DESI DR1 baryon acoustic oscillations (BAO) — including Big Bang Nucleosynthesis (BBN) constraints and a weak prior on the scalar spectral index — yields Ω m = 0.2994 ± 0.0090 and σ 8 = 0.836$^{+0.024}_{-0.027}$, representing improvements of approximately 4% and 23% over the baseline analysis, respectively. For the w 0 w a CDM model, our results from various data combinations are highly consistent, with all configurations converging to a region with w 0 > -1 and w a < 0. This convergence not only suggests intriguing hints of dynamical dark energy but also underscores the robustness of our HOD-informed prior approach in delivering reliable cosmological constraints.

59 BASIC BIOLOGICAL SCIENCES↗

HSC-XXL: Baryon budget of the 136 XXL groups and clusters

Abstract We present our determination of the baryon budget for an X-ray-selected XXL sample of 136 galaxy groups and clusters spanning nearly two orders of magnitude in mass (M500 ∼ 1013–1015 M⊙) and the redshift range 0 ≲ z ≲ 1. Our joint analysis is based on the combination of Hyper Suprime-Cam Subaru Strategic Program (HSC-SSP) weak-lensing mass measurements, XXL X-ray gas mass measurements, and HSC and Sloan Digital Sky Survey multiband photometry. We carry out a Bayesian analysis of multivariate mass-scaling relations of gas mass, galaxy stellar mass, stellar mass of brightest cluster galaxies (BCGs), and soft-band X-ray luminosity, by taking into account the intrinsic covariance between cluster properties, selection effect, weak-lensing mass calibration, and observational error covariance matrix. The mass-dependent slope of the gas mass–total mass (M500) relation is found to be $1.29_{-0.10}^{+0.16}$, which is steeper than the self-similar prediction of unity, whereas the slope of the stellar mass–total mass relation is shallower than unity; $0.85_{-0.09}^{+0.12}$. The BCG stellar mass weakly depends on cluster mass with a slope of $0.49_{-0.10}^{+0.11}$. The baryon, gas mass, and stellar mass fractions as a function of M500 agree with the results from numerical simulations and previous observations. We successfully constrain the full intrinsic covariance of the baryonic contents. The BCG stellar mass shows the larger intrinsic scatter at a given halo total mass, followed in order by stellar mass and gas mass. We find a significant positive intrinsic correlation coefficient between total (and satellite) stellar mass and BCG stellar mass and no evidence for intrinsic correlation between gas mass and stellar mass. All the baryonic components show no redshift evolution.

Akino, Daichi↗

Automated and efficient local adaptive regression for principal component-based reduced-order modeling of turbulent reacting flows

Principal Component Analysis can be used to reduce the cost of Computational Fluid Dynamics simulations of turbulent reacting flows by reducing the dimensionality of the transported variables through projection of the thermochemical state onto a lower-dimensional manifold. However, because of the nonlinearity of the principal component source terms, nonlinear regression techniques must be utilized for the source terms in terms of the principal components. Unfortunately, widely available and utilized nonlinear regression techniques can have prohibitive computational requirements and/or accuracy that is highly dependent on user experience in ad hoc tuning of model architecture and hyperparameters. Here, in this work, a new nonlinear regression approach is proposed that is both computationally efficient and automated so does not require any user input. The approach is evaluated through a priori prediction of principal component source terms using data from a Direct Numerical Simulation of a turbulent nonpremixed n-heptane/air jet flame. In particular, the proposed framework consists of local regressions whose complexity is adapted according to the local nonlinearity of the data: local linear regression when accurate enough and local Artificial Neural Networks when nonlinear regression is required. The number of local clusters for local regression is determined automatically using the Davies-Bouldin index. In addition, Bayesian optimization is utilized for model training (i.e., to select the best architectures and hyperparameters of the nonlinear regressions in an unsupervised fashion), eliminating ad hoc hand-tuning and/or expensive grid searches. Overall, compared to a single, global neural network, the new local adaptive regression approach is shown to have comparable accuracy but 69% less training time due to the utilization of local linear regression and faster training of local neural networks.

42 ENGINEERING↗

aphBO-2GP-3B: a budgeted asynchronous parallel multi-acquisition functions for constrained Bayesian optimization on high-performing computing architecture

High-fidelity complex engineering simulations are often predictive, but also computationally expensive and often require substantial computational efforts. The mitigation of computational burden is usually enabled through parallelism in high-performance cluster (HPC) architecture. Optimization problems associated with these applications is a challenging problem due to the high computational cost of the high-fidelity simulations. In this paper, an asynchronous parallel constrained Bayesian optimization method is proposed to efficiently solve the computationally expensive simulation-based optimization problems on the HPC platform, with a budgeted computational resource, where the maximum number of simulations is a constant. The advantage of this method are three-fold. Firstly, the efficiency of the Bayesian optimization is improved, where multiple input locations are evaluated parallel in an asynchronous manner to accelerate the optimization convergence with respect to physical runtime. This efficiency feature is further improved so that when each of the inputs is finished, another input is queried without waiting for the whole batch to complete. Second, the proposed method can handle both known and unknown constraints. Third, the proposed method samples several acquisition functions based on their rewards using a modified GP-Hedge scheme. The proposed framework is termed aphBO-2GP-3B, which means asynchronous parallel hedge Bayesian optimization with two Gaussian processes and three batches. The numerical performance of the proposed framework aphBO-2GP-3B is comprehensively benchmarked using 16 numerical examples, compared against other 6 parallel Bayesian optimization variants and 1 parallel Monte Carlo as a baseline, and demonstrated using two real-world high-fidelity expensive industrial applications. The first engineering application is based on finite element analysis (FEA) and the second one is based on computational fluid dynamics (CFD) simulations.

97 MATHEMATICS AND COMPUTING↗

Dark Energy Survey Year 3 results: optimized $w$CDM simulation-based inference with weak lensing map-level hybrid statistics

We present cosmological constraints from the Dark Energy Survey Year 3 (DES Y3) weak lensing data using hierarchical hybrid statistics within a Bayesian simulation-based inference framework that is based on the Gower Street simulations. To maximize the precision of the inference, we have developed a new, information-theory based, data compression of the weak lensing maps to just seven highly informative summary statistics. The hybrid scheme exploits the high information content of the power spectrum, compressing both the power spectrum and neural-based summaries that are designed to extract further information. Our simulation-based approach enables principled forward modelling of all major sources of systematic uncertainty and survey properties into realistic mock observations, including the survey mask, photometric redshift uncertainties, intrinsic galaxy alignments, multiplicative shear calibration bias, source galaxy clustering, non-Gaussian shape noise, and non-linear structure formation. The summary statistics are then used in a Bayesian simulation-based inference pipeline. The inference is validated through coverage tests and checks for robustness against baryonic feedback. Assuming a $w$CDM cosmology, our analysis yields $S_8 = 0.808 \pm 0.017$, $Ω_{\rm m} = 0.325 \pm 0.024$, and $w < -0.766$ (marginalized posterior 68 per cent credible intervals). This rigorous combination of information theory, physics- and neural network-based extreme data compression, and principled Bayesian analysis improves the figure of merit for $(Ω_{\rm m}, S_8, w)$ by 60 per cent over the previous state-of-the-art, and by almost a factor of 3 over two-point analyses of the same data. They are the most precise joint constraints on $(Ω_{\rm m}, S_8, w)$ from weak gravitational lensing data alone of any survey to date. We intend to apply this analysis to the more recent DES Y6 data.

Williamson, J. [University Coll. London]↗

Cluster expansion by transfer learning for phase stability predictions

Recent progress towards universal machine-learned interatomic potentials holds considerable promise for materials discovery. Yet the accuracy of these potentials for predicting phase stability may still be limited. In contrast, cluster expansions provide accurate phase stability predictions but are computationally demanding to parameterize from first principles, especially for structures of low dimension or with a large number of components, such as interfaces or multimetal catalysts. We overcome this trade-off via transfer learning. Using Bayesian inference, we incorporate prior statistical knowledge from machine-learned and physics-based potentials, enabling us to sample the most informative configurations and to efficiently fit first-principles cluster expansions. Furthermore, this algorithm is tested on Pt:Ni, showing robust convergence of the mixing energies as a function of sample size with reduced statistical fluctuations.

36 MATERIALS SCIENCE↗

The Attraction Indian Buffet Distribution

We propose the attraction Indian buffet distribution (AIBD), a distribution for binary feature matrices influenced by pairwise similarity information. Binary feature matrices are used in Bayesian models to uncover latent variables (i.e., features) that explain observed data. The Indian buffet process (IBP) is a popular exchangeable prior distribution for latent feature matrices. In the presence of additional information, however, the exchangeability assumption is not reasonable or desirable. The AIBD can incorporate pairwise similarity information, yet it preserves many properties of the IBP, including the distribution of the total number of features. Thus, much of the interpretation and intuition that one has for the IBP directly carries over to the AIBD. A temperature parameter controls the degree to which the similarity information affects feature-sharing between observations. Unlike other nonexchangeable distributions for feature allocations, the probability mass function of the AIBD has a tractable normalizing constant, making posterior inference on hyperparameters straight-forward using standard MCMC methods. A novel posterior sampling algorithm is proposed for the IBP and the AIBD. We demonstrate the feasibility of the AIBD as a prior distribution in feature allocation models and compare the performance of competing methods in simulations and an application.

97 MATHEMATICS AND COMPUTING↗

The Dark Energy Survey Year 3 high-redshift sample: selection, characterization, and analysis of galaxy clustering

ABSTRACT The fiducial cosmological analyses of imaging surveys like DES typically probe the Universe at redshifts z < 1. We present the selection and characterization of high-redshift galaxy samples using DES Year 3 data, and the analysis of their galaxy clustering measurements. In particular, we use galaxies that are fainter than those used in the previous DES Year 3 analyses and a Bayesian redshift scheme to define three tomographic bins with mean redshifts around z ∼ 0.9, 1.2, and 1.5, which extend the redshift coverage of the fiducial DES Year 3 analysis. These samples contain a total of about 9 million galaxies, and their galaxy density is more than 2 times higher than those in the DES Year 3 fiducial case. We characterize the redshift uncertainties of the samples, including the usage of various spectroscopic and high-quality redshift samples, and we develop a machine-learning method to correct for correlations between galaxy density and survey observing conditions. The analysis of galaxy clustering measurements, with a total signal to noise S/N ∼ 70 after scale cuts, yields robust cosmological constraints on a combination of the fraction of matter in the Universe Ωm and the Hubble parameter h, $\Omega _m h = 0.195^{+0.023}_{-0.018}$, and 2–3 per cent measurements of the amplitude of the galaxy clustering signals, probing galaxy bias and the amplitude of matter fluctuations, bσ8. A companion paper (in preparation) will present the cross-correlations of these high-z samples with cosmic microwave background lensing from Planck and South Pole Telescope, and the cosmological analysis of those measurements in combination with the galaxy clustering presented in this work.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

AutoClass: A Bayesian Approach to Classification

We describe a Bayesian approach to the untutored discovery of classes in a set of cases, sometimes called finite mixture separation or clustering. The main difference between clustering and our approach is that we search for the "best" set of class descriptions rather than grouping the cases themselves. We describe our classes in terms of a probability distribution or density function, and the locally maximal posterior probability valued function parameters. We rate our classifications with an approximate joint probability of the data and functional form, marginalizing over the parameters. Approximation is necessitated by the computational complexity of the joint probability. Thus, we marginalize w.r.t. local maxima in the parameter space. We discuss the rationale behind our approach to classification. We give the mathematical development for the basic mixture model and describe the approximations needed for computational tractability. We instantiate the basic model with the discrete Dirichlet distribution and multivariant Gaussian density likelihoods. Then we show some results for both constructed and actual data.

Stutz, John↗