Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Bayesian clustering”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Quantifying the basic reproduction number and underestimated fraction of Mpox cases worldwide at the onset of the outbreak

In 2022, there was a global resurgence of mpox, with different clinical-epidemiological features compared with previous outbreaks. Sexual contact was hypothesized as the primary transmission route, and the community of men having sex with men (MSM) was disproportionately affected. Because of the stigma associated with sexually transmitted infections, the real burden of mpox could be masked. We quantified the basic reproduction number (R 0 ) and the underestimated fraction of mpox cases in 16 countries, from the onset of the outbreak until early September 2022, using Bayesian inference and a compartmentalized, risk-structured (high-/low-risk populations) and two-route (sexual/non-sexual transmission) mathematical model. Machine learning (ML) was harnessed to identify underestimation determinants. Estimated R 0 ranged between 1.37 (Canada) and 3.68 (Germany). The underestimation rates for the high- and low-risk populations varied between 25–93% and 65–85%, respectively. The estimated total number of mpox cases, relative to the reported cases, is highest in Colombia (3.60) and lowest in Canada (1.08). In the ML analysis, two clusters of countries could be identified, differing in terms of attitudes towards the 2SLGBTQIAP+ community and the importance of religion. Given the substantial mpox underestimation, surveillance should be enhanced, and country-specific campaigns against the stigmatization of MSM should be organized, leveraging community-based interventions.

60 APPLIED LIFE SCIENCES↗

Discovery of Self-Assembling π-Conjugated Peptides by Active Learning-Directed Coarse-Grained Molecular Simulation

Electronically active organic molecules have demonstrated great promise as novel soft materials for energy harvesting and transport. Self-assembled nanoaggregates formed from π-conjugated oligopeptides composed of an aromatic core flanked by oligopeptide wings offer emergent optoelectronic properties within a water-soluble and biocompatible substrate. Nanoaggregate properties can be controlled by tuning core chemistry and peptide composition, but the sequence–structure–function relations remain poorly characterized. Here, we employ coarse-grained molecular dynamics simulations within an active learning protocol employing deep representational learning and Bayesian optimization to efficiently identify molecules capable of assembling pseudo-1D nanoaggregates with good stacking of the electronically active π-cores. We consider the DXXX-OPV3-XXXD oligopeptide family, where D is an Asp residue and OPV3 is an oligophenylenevinylene oligomer (1,4-distyrylbenzene), to identify the top performing XXX tripeptides within all 20 3 = 8000 possible sequences. By direct simulation of only 2.3% of this space, we identify molecules predicted to exhibit superior assembly relative to those reported in prior work. Spectral clustering of the top candidates reveals new design rules governing assembly. This work establishes new understanding of DXXX-OPV3-XXXD assembly, identifies promising new candidates for experimental testing, and presents a computational design platform that can be generically extended to other peptide-based and peptide-like systems.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

The cosmic waltz of Coma Berenices and Latyshev 2 (Group X)

Context. Open clusters (OCs) are fundamental benchmarks where theories of star formation and stellar evolution can be tested and validated. Coma Berenices (Coma Ber) and Latyshev 2 (Group X) are the second and third OCs closest to the Sun, making them excellent targets to search for low-mass stars and ultra-cool dwarfs. In addition, this pair will experience a flyby in 10–16 Myr, making it a benchmark to test pair interactions of OCs. Aims. We aim to analyse the membership, luminosity, mass, phase-space (i.e. positions and velocities), and energy distributions for Coma Ber and Latyshev 2 and test the hypothesis of the mixing of their populations at the encounter time. Methods. We developed a new phase-space membership methodology and applied it to Gaia data. With the recovered members, we inferred the phase-space, luminosity, and mass distributions using publicly available Bayesian inference codes. Then, with a publicly available orbit integration code and members’ positions and velocities, we integrated their orbits 20 Myr into the future. Results. In Coma Ber, we identified 302 candidate members distributed in the core and tidal tails. The tails are dynamically cold and asymmetrically populated. The stellar system called Group X is made of two structures: the disrupted OC Latyshev 2 (186 candidate members) and a loose stellar association called Mecayotl 1 (146 candidate members), and both of them will fly by Coma Ber in 11.3 ± 0.5 Myr and 14.0 ± 0.6 Myr, respectively, and each other in 8.1 ± 1.3 Myr. Conclusions. We study the dynamical properties of the core and tails of Coma Ber and also confirm the existence of the OC Latyshev 2 and its neighbour stellar association Mecayotl 1. Although these three systems will experience encounters, we find no evidence supporting the mixing of their populations.

79 ASTRONOMY AND ASTROPHYSICS↗

A scalable variational method for estimating the latent infection-rate field of an outbreak

In this paper, we explore whether the infection-rate of a disease can serve as a robust monitoring variable in epidemiological surveillance algorithms. The infection-rate is dependent on population mixing patterns that do not vary erratically day-to-day; in contrast, daily case-counts used in contemporary surveillance algorithms are corrupted by reporting errors. The technical challenge lies in estimating the latent infection-rate from case-counts. Here we devise a Bayesian method to estimate the infection-rate across multiple adjoining areal units, and then use it, via an anomaly detector, to discern a change in epidemiological dynamics. We extend an existing model for estimating the infection-rate in an areal unit by incorporating a Markov random field model, so that we may estimate infection-rates across multiple areal units, while preserving spatial correlations observed in the epidemiological dynamics. To carry out the high-dimensional Bayesian inverse problem, we develop an implementation of mean-field variational inference specific to the infection model and integrate it with the random field model to incorporate correlations across counties. The method is tested on estimating the COVID-19 infection-rates across all 33 counties in New Mexico using data from the summer of 2020, and then employing them to detect the arrival of the Fall 2020 COVID-19 wave. We perform the detection using a temporal algorithm that is applied county-by-county. We also show how the infection-rate field can be used to cluster counties with similar epidemiological dynamics.

60 APPLIED LIFE SCIENCES↗

Galaxy Clustering with LSST: Effects of Number Count Bias from Blending

The Vera C. Rubin Observatory Legacy Survey of Space and Time (LSST) will survey the southern sky to create the largest galaxy catalog to date, and its statistical power demands an improved understanding of systematic effects such as source overlaps, also known as blending. In this work we study how blending introduces a bias in the number counts of galaxies (instead of the flux and colors), and how it propagates into galaxy clustering statistics. We use the 300 deg 2 DC2 image simulation and its resulting galaxy catalog (LSST Dark Energy Science Collaboration et al. 2021) to carry out this study. We find that, for a LSST Year 1 (Y1)-like cosmological analyses, the number count bias due to blending leads to small but statistically significant differences in mean redshift measurements when comparing an observed sample to an unblended calibration sample. In the two-point correlation function, blending causes differences greater than 3σ on scales below approximately 10', but large scales are unaffected. We fit Ω m and linear galaxy bias in a Bayesian cosmological analysis and find that the recovered parameters from this limited area sample, with the LSST Y1 scale cuts, are largely unaffected by blending. Our main results hold when considering photometric redshift and a LSST Year 5 (Y5)-like sample.

79 ASTRONOMY AND ASTROPHYSICS↗

Data augmentation for disruption prediction via robust surrogate models

The goal of this work is to generate large statistically representative data sets to train machine learning models for disruption prediction provided by data from few existing discharges. Such a comprehensive training database is important to achieve satisfying and reliable prediction results in artificial neural network classifiers. Here, we aim for a robust augmentation of the training database for multivariate time series data using Student t process regression. We apply Student t process regression in a state space formulation via Bayesian filtering to tackle challenges imposed by outliers and noise in the training data set and to reduce the computational complexity. Thus, the method can also be used if the time resolution is high. We use an uncorrelated model for each dimension and impose correlations afterwards via colouring transformations. We demonstrate the efficacy of our approach on plasma diagnostics data of three different disruption classes from the DIII-D tokamak. To evaluate if the distribution of the generated data is similar to the training data, we additionally perform statistical analyses using methods from time series analysis, descriptive statistics and classic machine learning clustering algorithms.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Data augmentation for disruption prediction via robust surrogate models

The goal of this work is to generate large statistically representative datasets to train machine learning models for disruption prediction provided by data from few existing discharges. Such a comprehensive training database is important to achieve satisfying and reliable prediction results in artificial neural network classifiers. Here, we aim for a robust augmentation of the training database for multivariate time series data using Student-t process regression. We apply Student-t process regression in a state space formulation via Bayesian filtering to tackle challenges imposed by outliers and noise in the training data set and to reduce the computational complexity. Thus, the method can also be used if the time resolution is high. We use an uncorrelated model for each dimension and impose correlations afterwards via coloring transformations. We demonstrate the efficacy of our approach on plasma diagnostics data of three different disruption classes from the DIII-D tokamak. To evaluate if the distribution of the generated data is similar to the training data, we additionally perform statistical analyses using methods from time series analysis, descriptive statistics, and classic machine learning clustering algorithms.

97 MATHEMATICS AND COMPUTING↗

On evaluating the accuracy of SAR sea-ice classification using multifrequency polarimetric AIRSAR data

We investigate how multifrequency and polarimetric synthetic aperture radar (SAR) imagery enhances present capability to discriminate different ice conditions in single-frequency, single-polarization satellite SAR data. Frequencies considered are C- (lambda = 5.6cm), L- (lambda = 24cm) and P- (lambda = 68cm) band. Radar backscatter characteristics of six radiometrically and polarimetrically distinct ice types are selected from a cluster analysis of the multifrequency polarimetric SAR data and used to classify SAR images. Validation of these ice conditions is based on information provided by aerial photos, weather and ice surface measurements acquired at an ice camp, together with airborne passive microwave imagery, and visual analysis of the SAR data. The six identified sea-ice types are: (1) multiyear sea-ice; (2) compressed first year ice; (3) first year rubble and ridges; (4) first year rough ice; (5) first year smooth ice; and (6) thin ice. Open water is absent in all analyzed data. Classification of the SAR imagery into those six ice types is performed using a Bayesian Maximum A Posteriori classifier. Two complete scenes acquired at different dates in different locations are classified. The scenes were chosen such that they are representative of typical ice conditions in the Beaufort Sea in March 1988 and because ancillary information is available for validating the segmentation of various ice surface conditions.

Drinkwater, Mark R.↗

The cosmic DANCe of Perseus

Context. Star-forming regions are excellent benchmarks for testing and validating theories of star formation and stellar evolution. The Perseus star-forming region, being one of the youngest (< 10 Myr), closest (280-320 pc), and most studied in the literature, is a fundamental benchmark. Aims. We aim to study the membership, phase-space structure, mass, and energy (kinetic plus potential) distribution of the Perseus star-forming region using public catalogues (Gaia, APOGEE, 2MASS, and Pan-STARRS). Methods. We used Bayesian methodologies that account for extinction to identify the Perseus physical groups in the phase-space, retrieve their candidate members, derive their properties (age, mass, 3D positions, 3D velocities, and energy), and attempt to reconstruct their origin. Results. We identify 1052 candidate members in seven physical groups (one of them new) with ages between 3 and 10 Myr, dynamical super-virial states, and large fractions of energetically unbounded stars. Their mass distributions are broadly compatible with that of Chabrier for masses ≳0.1 M ⊙ and do not show hints of over-abundance of low-mass stars in NGC 1333 with respect to IC 348. These groups’ ages, spatial structure, and kinematics are compatible with at least three generations of stars. Future work is still needed to clarify if the formation of the youngest was triggered by the oldest. Conclusions. The exquisite Gaia data complemented with public archives and mined with comprehensive Bayesian methodologies allow us to identify 31% more members than previous studies, discover a new physical group (Gorgophone: 7 Myr, 191 members, and 145 M ⊙ ), and confirm that the spatial, kinematic, and energy distributions of these groups support the hierarchical star formation scenario.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Improving Subsurface Stress Characterization for Carbon Dioxide Storage Projects by Incorporating Machine Learning Techniques

The overall objective of this project is to develop a framework for reliable characterization and prediction of the state of stress in the overburden and underburden (including the basement) in CO 2 storage reservoirs using machine learning and integrated geomechanics and geophysical methods. Specifically, we propose to develop workflow encompassing of technologies and/or methods to predict stress and pressure changes due to CO 2 injection in an active tertiary recovery site and their impacts on subtle fault activation, fractures and occurrence of microseismic events and compare responses to field observations. In this project, we anticipate using dataset from the Farnsworth field Unit (FWU) which is operated by Purdure Petroleum. A novel elastic-waveform VSP inversion technique will be used to estimate high-resolution spatial and temporal changes of elastic moduli in CO 2 storage reservoirs, which will be combined with velocity-stress relationship derived from laboratory tests to obtain subsurface pressure and stress. Clustered microseismic data will be jointly inverted for improved focal mechanisms. Least-squares reverse-time migration of microseismic waveform data will be performed to directly image fracture/fault zones. Additionally, a deep neural network machine learning technique with convolutional and recurrent layers will be used for learning the spectro-temporal structures in microseismic waveforms. The results of this geotechnical data analysis will be integrated to develop a high-resolution 3D mechanical earth model extending from the overburden sealing formations to the underburden including the basement. Mechanical properties will be derived through integration of mechanical logs, tests, available results from chemo-mechanical laboratory tests, and elastic inversion of seismic data using a combination of Bayesian and stochastic methods as well as machine learning technique. Failure features (faults/fractures) will be represented and/or modeled based on seismic and core data analysis. A transient hydrodynamic-geomechanical model will be developed through coupling with the calibrated FWU reservoir simulation model. The full physics coupled model will be used to train a reduced order proxy model using machine learning algorithm for estimating stress which will then be used with appropriate constitutive relationships and forward seismological models to simulate pressure changes and induced microseismicity. An advanced optimization framework will be developed to perform a history match to minimize error between field observations and simulated. The history matched proxy model will be verified against the full-physics equivalent. The field observations that will be used in the coupled model calibration process include pressure/stress inverted from VSP, moment magnitude from microseismic analysis, real time downhole pressure measurements, production and injection data. Parameter sensitivity and uncertainty analysis will be performed to characterize the impact of model parameter uncertainty on stress estimates. The proposed project will have significant impact on future field implementation of the proposed technology. Because the project field site is an ongoing CO 2 EOR development, the value of the new technology will be demonstrated in an operational context and evaluated as a viable risk mitigation strategy. Cost/benefit will be evaluated together with the various commercial incentives for CO 2 sequestration available to oil and gas operators. The extensive available dataset and ongoing data acquisition under the SWP Phase III work plan provides flexibility for investigation of multiple approaches and reduces technical risk.

58 GEOSCIENCES↗

Emulating ab initio computations of infinite nucleonic matter

We construct efficient emulators for the computation of the infinite nuclear matter equation of state. These emulators are based on the subspace-projected coupled-cluster method for which we here develop a new algorithm called small-batch voting to eliminate spurious states that might appear when emulating quantum many-body methods based on a non-Hermitian Hamiltonian. The efficiency and accuracy of these emulators facilitate a rigorous statistical analysis within which we explore nuclear matter predictions for > 10 6 different parametrizations of a chiral interaction model with explicit Δ -isobars at next-to-next-to leading order. Constrained by nucleon-nucleon scattering phase shifts and bound-state observables of light nuclei up to He 4 , we use history matching to identify nonimplausible domains for the low-energy coupling constants of the chiral interaction. Within these domains we perform a Bayesian analysis using sampling and importance resampling with different likelihood calibrations and study correlations between interaction parameters, calibration observables in light nuclei, and nuclear matter saturation properties. Published by the American Physical Society 2024

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Using Regionalized Air Quality Model Performance and Bayesian Maximum Entropy Data Fusion to Map Global Surface Ozone Concentration and Associated Uncertainty

Estimates of ground-level ozone concentrations have been improved through data fusion of observations and atmospheric chemistry models. Our previous global ozone estimates for the Global Burden of Disease study corrected for bias uniformly across continents and then corrected near monitoring stations using the Bayesian Maximum Entropy (BME) framework for data fusion. Here, we use the Regionalized Air Quality Model Performance (RAMP) framework to correct model bias over a much larger spatial range than BME can, accounting for the spatial inhomogeneity of bias and nonlinearity as a function of modeled ozone. RAMP bias correction is applied to a composite of 9 global chemistry-climate models, based on the nearest set of monitors. These estimates are then fused with observations using BME, which matches observations at measurement stations, with the influence of observations declining with distance in space and time. We create global ozone maps for each year from 1990 to 2017 at fine spatial resolution. RAMP is shown to create unrealistic discontinuities due to the spatial clustering of ozone monitors, which we overcome by applying a weighting for RAMP based on the number of monitors nearby. Incorporating RAMP before BME has little effect on model performance near stations, but strongly increases R 2 by 0.15 at locations farther from stations, shown through a checkerboard cross-validation. Corrections to estimates differ based on location in space and time, confirming heterogeneity. We quantify the likelihood of exceeding selected ozone levels, finding that parts of the Middle East, India, and China are most likely to exceed 55 parts per billion (ppb) in 2017. About 96% of the global population was exposed to ozone levels above the World Health Organization guideline of 60 µg m −3 (30 ppb) in 2017. Our annual fine-resolution ozone estimates may be useful for several applications including epidemiology and assessments of impacts on health, agriculture, and ecosystems.

Ozone↗

Initialization and Restart in Stochastic Local Search: Computing a Most Probable Explanation in Bayesian Networks

For hard computational problems, stochastic local search has proven to be a competitive approach to finding optimal or approximately optimal problem solutions. Two key research questions for stochastic local search algorithms are: Which algorithms are effective for initialization? When should the search process be restarted? In the present work we investigate these research questions in the context of approximate computation of most probable explanations (MPEs) in Bayesian networks (BNs). We introduce a novel approach, based on the Viterbi algorithm, to explanation initialization in BNs. While the Viterbi algorithm works on sequences and trees, our approach works on BNs with arbitrary topologies. We also give a novel formalization of stochastic local search, with focus on initialization and restart, using probability theory and mixture models. Experimentally, we apply our methods to the problem of MPE computation, using a stochastic local search algorithm known as Stochastic Greedy Search. By carefully optimizing both initialization and restart, we reduce the MPE search time for application BNs by several orders of magnitude compared to using uniform at random initialization without restart. On several BNs from applications, the performance of Stochastic Greedy Search is competitive with clique tree clustering, a state-of-the-art exact algorithm used for MPE computation in BNs.

Mengshoel, Ole J.↗

Status of experimental knowledge on the unbound nucleus 13 Be

The structure of the unbound nucleus 13 Be is important for understanding the Borromean, two-neutron halo nucleus 14 Be. The experimental studies conducted over the last four decades are reviewed in the context of the beryllium chain of isotopes and some significant theoretical studies. The focus of this paper is the comparison of new data from a 12 Be(d,p) reaction in inverse kinematics, which was analyzed using Geant4 simulations and a Bayesian fitting procedure, with previous measurements. Two possible scenarios to explain the strength below 1 MeV above the neutron separation energy were proposed in that study: a single p-wave resonance or a mixture of an s-wave virtual state with a weaker p- or d-wave resonance. Comparisons of recent invariant mass and the (d,p) experiments show good agreement between the transfer measurement and the two most recent high-energy nucleon removal measurements.

12Be↗

Multiscale Physics of Atomic Nuclei from First Principles

Atomic nuclei exhibit multiple energy scales ranging from hundreds of MeV in binding energies to fractions of an MeV for low-lying collective excitations. As the limits of nuclear binding are approached near the neutron and proton drip lines, traditional shell structure starts to melt with an onset of deformation and an emergence of coexisting shapes. It is a long-standing challenge to describe this multiscale physics starting from nuclear forces with roots in quantum chromodynamics. Here, we achieve this within a unified and nonperturbative quantum many-body framework that captures both short- and long-range correlations starting from modern nucleon-nucleon and three-nucleon forces from chiral effective field theory. The short-range (dynamic) correlations which account for the bulk of the binding energy are included within a symmetry-breaking framework, while long-range (static) correlations (and fine details about the collective structure) are included by employing symmetry projection techniques. Our calculations accurately reproduce—within theoretical error bars—available experimental data for low-lying collective states and the electromagnetic quadrupole transitions in 20−30 Ne. In addition, we reveal coexisting spherical and deformed shapes in 30 Ne, which indicates the breakdown of the magic neutron number 𝑁 = 20 as the key nucleus 28 O is approached, and we predict that the drip line nuclei 32,34 Ne are strongly deformed and collective. By developing reduced-order models for symmetry-projected states, we perform a global sensitivity analysis and find that the subleading singlet 𝑆-wave contact and a pion-nucleon coupling strongly impact nuclear deformation in chiral effective field theory. The techniques developed in this work clarify how microscopic nuclear forces generate the multiscale physics of nuclei spanning collective phenomena as well as short-range correlations and allow one to capture emergent and dynamical phenomena in finite fermion systems such as atom clusters, molecules, and atomic nuclei.

74 ATOMIC AND MOLECULAR PHYSICS↗

Almost medium-free measurement of the Hoyle state direct-decay component with a TPC

The structure of the Hoyle state, a highly α -clustered state at 7.65 MeV in 12 C, has long been the subject of debate. Understanding if the system comprises of three weakly interacting α particles in the 0s orbital, known as an α-condensate state, is possible by studying the decay branches of the Hoyle state. The direct decay of the Hoyle state into three α particles, rather than through the 8 Be ground state, can be identified by studying the energy partition of the three α particles arising from the decay. This paper provides details on the breakup mechanism of the Hoyle stating using a new experimental technique. Method: By using β-delayed charged-particle spectroscopy of 12 N using the Texas active target time-projection chamber, a high-sensitivity measurement of the direct 3α decay ratio can be performed without contributions from pileup events. A Bayesian approach to understanding the contribution of the direct components via a likelihood function shows that the direct component is <0.043% at the 95% confidence level. This value is in agreement with several other studies, and, here, we can demonstrate that a small nonsequential component with a decay fraction of about 10 –4 is most likely. Here, the measurement of the non-sequential component of the Hoyle state decay is performed in an almost medium-free reaction for the first time. The derived upper limit is in agreement with previous studies and demonstrates sensitivity to the absolute branching ratio. Further experimental studies would need to be combined with robust microscopic theoretical understanding of the decay dynamics to provide additional insight into the idea of the Hoyle state as an α condensate.

6 ≤ A ≤ 19↗

Physics-based reward driven image analysis in microscopy

The rise of electron microscopy has expanded our ability to acquire nanometer and atomically resolved images of complex materials. The resulting vast datasets are typically analyzed by human operators, an intrinsically challenging process due to the multiple possible analysis steps and the corresponding need to build and optimize complex analysis workflows. We present a methodology based on the concept of a Reward Function coupled with Bayesian Optimization, to optimize image analysis workflows dynamically. The Reward Function is engineered to closely align with the experimental objectives and broader context and is quantifiable upon completion of the analysis. Here, cross-section, high-angle annular dark field (HAADF) images of ion-irradiated (Y, Dy)Ba 2 Cu 3 O 7–δ thin-films were used as a model system. The reward functions were formed based on the expected materials density and atomic spacings and used to drive multi-objective optimization of the classical Laplacian-of-Gaussian (LoG) method. These results can be benchmarked against the DCNN segmentation. This optimized LoG* compares favorably against DCNN in the presence of the additional noise. We further extend the reward function approach towards the identification of partially-disordered regions, creating a physics-driven reward function and action space of high-dimensional clustering. We pose that with correct definition, the reward function approach allows real-time optimization of complex analysis workflows at much higher speeds and lower computational costs than classical DCNN-based inference, ensuring the attainment of results that are both precise and aligned with the human-defined objectives.

47 OTHER INSTRUMENTATION↗

Inductive Approaches to Improving Diagnosis and Design for Diagnosability

The first research area under this grant addresses the problem of classifying time series according to their morphological features in the time domain. A supervised learning system called CALCHAS, which induces a classification procedure for signatures from preclassified examples, was developed. For each of several signature classes, the system infers a model that captures the class's morphological features using Bayesian model induction and the minimum message length approach to assign priors. After induction, a time series (signature) is classified in one of the classes when there is enough evidence to support that decision. Time series with sufficiently novel features, belonging to classes not present in the training set, are recognized as such. A second area of research assumes two sources of information about a system: a model or domain theory that encodes aspects of the system under study and data from actual system operations over time. A model, when it exists, represents strong prior expectations about how a system will perform. Our work with a diagnostic model of the RCS (Reaction Control System) of the Space Shuttle motivated the development of SIG, a system which combines information from a model (or domain theory) and data. As it tracks RCS behavior, the model computes quantitative and qualitative values. Induction is then performed over the data represented by both the 'raw' features and the model-computed high-level features. Finally, work on clustering for operating mode discovery motivated some important extensions to the clustering strategy we had used. One modification appends an iterative optimization technique onto the clustering system; this optimization strategy appears to be novel in the clustering literature. A second modification improves the noise tolerance of the clustering system. In particular, we adapt resampling-based pruning strategies used by supervised learning systems to the task of simplifying hierarchical clusterings, thus making post-clustering analysis easier.

Fisher, Douglas H.↗