Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Bayesian model selection”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

The multicategory case of the sequential Bayesian pixel selection and estimation procedure

A Bayesian technique for stratified proportion estimation and a sampling based on minimizing the mean squared error of this estimator were developed and tested on LANDSAT multispectral scanner data using the beta density function to model the prior distribution in the two-class case. An extention of this procedure to the k-class case is considered. A generalization of the beta function is shown to be a density function for the general case which allows the procedure to be extended.

Pore, M. D.↗

Probabilistic system identification in the time domain

The objective of system identification is to determine reliable dynamical models of a structure by systematically using its measured excitation and response. It brings together in an integrated fashion, experimental, analytical, and computational techniques in structural dynamics. Areas of application for system identification include the following: (1) Model Evaluation--assessing assumptions (linearity and equivalent viscous damping) and techniques (finite-element modeling) used to construct theoretical models of a structure; (2) Model Improvement--updating of a theoretical model to enable more accurate response predictions for possible future loads on the structure, or for control of the structure; (3) Empirical Modelling--developing empirical relationships (nonlinear models) or empirical parameter values (modal damping) because the present state of the art does not provide theoretical results; and (4) Damage Detection and Assessment--continual or episodic updating of a structural model through vibration monitoring to detect and locate any structural damage. It can be argued that since the construction or modification of models using test data is subject to inherent uncertainties, the above problems should be properly treated within a Bayesian probabilistic framework. Such a methodology is presented which allows the precision of the estimates of the model parameters to be computed. It also leads to a guiding principle in applications. Namely, when selecting a single model from a given class of models, one should take the most probable model in the class based on the experimental data. Practical applications of this principle are given which are based on the utilization of measured seismic motions in large civil structures. Examples include the application of a computer program MODE-ID to identify modal properties directly from seismic excitation and response time histories from a nine-story steel-frame building at JPL and from a freeway overpass bridge.

Beck, James L.↗

A Bayesian Framework for Reliability Analysis of Spacecraft Deployments

Deployable subsystems are essential to mission success of most spacecraft. These subsystems enable critical functions including power, communications and thermal control. The loss of any of these functions will generally result in loss of the mission. These subsystems and their components often consist of unique designs and applications for which various standardized data sources are not applicable for estimating reliability and for assessing risks. In this study, a two stage sequential Bayesian framework for reliability estimation of spacecraft deployment was developed for this purpose. This process was then applied to the James Webb Space Telescope (JWST) Sunshield subsystem, a unique design intended for thermal control of the Optical Telescope Element. Initially, detailed studies of NASA deployment history, "heritage information", were conducted, extending over 45 years of spacecraft launches. This information was then coupled to a non-informative prior and a binomial likelihood function to create a posterior distribution for deployments of various subsystems uSing Monte Carlo Markov Chain sampling. Select distributions were then coupled to a subsequent analysis, using test data and anomaly occurrences on successive ground test deployments of scale model test articles of JWST hardware, to update the NASA heritage data. This allowed for a realistic prediction for the reliability of the complex Sunshield deployment, with credibility limits, within this two stage Bayesian framework.

Evans, John W.↗

Towards a Rigorous Basis for Specific Operations Risk Assessment of UAS

The Specific Operations Risk Assessment (SORA) guidance represents the consensus of various national aviation authorities on a common process to identify, qualitatively assess, and manage the safety risk posed by unmanned aircraft systems (UAS), when preparing the safety case required for regulatory approval to conduct certain types of operations. As such, it can be considered a de facto standard, being increasingly adopted by various relevant stakeholders. This paper first gives an overview of the SORA process and associated methods, identifying a number of inconsistencies in risk identification and assessment, also discussing plausible strategies to close the associated gaps. Then, we give a well-founded basis for the applicable concepts, such as barrier integrity, assurance, and robustness, following which we present a preliminary and simple probabilistic formalization of the underpinning barrier-based safety model. We illustrate our overall approach through a worked example, also discussing how a Bayesian framework can facilitate extending and enhancing our initial formalization. We conclude with a discussion of the opportunities afforded by our approach, such as a well-founded basis for barrier selection, whilst addressing the associated challenges. The main objective of this work is to complement the current SORA guidance through a principled, mathematicallybased approach to risk assessment, particularly when it is applied to higher-risk operational concepts that warrant greater rigor in safety assessment and assurance.

Safety Cases↗

Perceptual learning through optimization of attentional weighting: human versus optimal Bayesian learner

Human performance in visual detection, discrimination, identification, and search tasks typically improves with practice. Psychophysical studies suggest that perceptual learning is mediated by an enhancement in the coding of the signal, and physiological studies suggest that it might be related to the plasticity in the weighting or selection of sensory units coding task relevant information (learning through attention optimization). We propose an experimental paradigm (optimal perceptual learning paradigm) to systematically study the dynamics of perceptual learning in humans by allowing comparisons to that of an optimal Bayesian algorithm and a number of suboptimal learning models. We measured improvement in human localization (eight-alternative forced-choice with feedback) performance of a target randomly sampled from four elongated Gaussian targets with different orientations and polarities and kept as a target for a block of four trials. The results suggest that the human perceptual learning can occur within a lapse of four trials (<1 min) but that human learning is slower and incomplete with respect to the optimal algorithm (23.3% reduction in human efficiency from the 1st-to-4th learning trials). The greatest improvement in human performance, occurring from the 1st-to-2nd learning trial, was also present in the optimal observer, and, thus reflects a property inherent to the visual task and not a property particular to the human perceptual learning mechanism. One notable source of human inefficiency is that, unlike the ideal observer, human learning relies more heavily on previous decisions than on the provided feedback, resulting in no human learning on trials following a previous incorrect localization decision. Finally, the proposed theory and paradigm provide a flexible framework for future studies to evaluate the optimality of human learning of other visual cues and/or sensory modalities.

Non-NASA Center↗

Building a Database for a Quantitative Model

A database can greatly benefit a quantitative analysis. The defining characteristic of a quantitative risk, or reliability, model is the use of failure estimate data. Models can easily contain a thousand Basic Events, relying on hundreds of individual data sources. Obviously, entering so much data by hand will eventually lead to errors. Not so obviously entering data this way does not aid linking the Basic Events to the data sources. The best way to organize large amounts of data on a computer is with a database. But a model does not require a large, enterprise-level database with dedicated developers and administrators. A database built in Excel can be quite sufficient. A simple spreadsheet database can link every Basic Event to the individual data source selected for them. This database can also contain the manipulations appropriate for how the data is used in the model. These manipulations include stressing factors based on use and maintenance cycles, dormancy, unique failure modes, the modeling of multiple items as a single "Super component" Basic Event, and Bayesian Updating based on flight and testing experience. A simple, unique metadata field in both the model and database provides a link from any Basic Event in the model to its data source and all relevant calculations. The credibility for the entire model often rests on the credibility and traceability of the data.

Kahn, C. Joseph↗

The Growth of Mass and Reliability Source Database and Me

Mass and Reliability Source (MaRS) is a database which draws together information from multiple other sources and databases (MADS, VMDB, ISS PART, Bayesian PowerPoint, etc.). Information such as mass, operating hours, and failure data are compiled to create a tool for the building of probabilistic Risk Assessment (PRA) models. MaRS also has great potential for deep space mission planning. Understanding the lowest point of failure within an ORU helps to inform the efficient selection of spare parts necessary for a journey where mass and volume are at a premium and resupply is not possible.

George, Cory A.↗

Probabilistic Calibration of Expensive Models using Efficiently Trained Surrogates

Calibration of computational models in the presence of uncertainty is often cast as a Bayesian inference problem and solved via sampling methods, e.g., Markov chain Monte Carlo. When the computational model is expensive, this task becomes intractable due to the large number of samples required to accurately estimate the posterior distribution of the calibration parameters. A popular solution to this problem is to use machine learning to develop a faster-to-evaluate, lower-fidelity substitute for the original model to serve as a surrogate while solving the inference problem. Although considered an offline cost, generating training data to construct this surrogate model can still be an expensive task in practice. An active learning algorithm is presented that focuses training on improving surrogate accuracy specifically in and around the bulk of the posterior distribution, as this is where the model is exercised during calibration. Candidate samples are drawn from families of distributions related to an approximation of the posterior. The sample maximizing predictive variance is then selected for evaluation by the original computational model, yielding a label for the training point. Iterating this approach increases efficiency relative to space filling designs (e.g., Latin hypercube sampling) by avoiding low probability points. Practical considerations are discussed, including the benefits of using a sequential Monte Carlo sampling approach, convergence heuristics, and the importance of both exploration and exploitation given that the true posterior is unknown a priori.

uncertainty quantification↗

Application of Machine Learning to Rotorcraft Health Monitoring

Machine learning is a powerful tool for data exploration and model building with large data sets. This project aimed to use machine learning techniques to explore the inherent structure of data from rotorcraft gear tests, relationships between features and damage states, and to build a system for predicting gear health for future rotorcraft transmission applications. Classical machine learning techniques are difficult, if not irresponsible to apply to time series data because many make the assumption of independence between samples. To overcome this, Hidden Markov Models were used to create a binary classifier for identifying scuffing transitions and Recurrent Neural Networks were used to leverage long distance relationships in predicting discrete damage states. When combined in a workflow, where the binary classifier acted as a filter for the fatigue monitor, the system was able to demonstrate accuracy in damage state prediction and scuffing identification. The time dependent nature of the data restricted data exploration to collecting and analyzing data from the model selection process. The limited amount of available data was unable to give useful information, and the division of training and testing sets tended to heavily influence the scores of the models across combinations of features and hyper-parameters. This work built a framework for tracking scuffing and fatigue on streaming data and demonstrates that machine learning has much to offer rotorcraft health monitoring by using Bayesian learning and deep learning methods to capture the time dependent nature of the data. Suggested future work is to implement the framework developed in this project using a larger variety of data sets to test the generalization capabilities of the models and allow for data exploration.

machine learning↗

Autoclass: An automatic classification system

The task of inferring a set of classes and class descriptions most likely to explain a given data set can be placed on a firm theoretical foundation using Bayesian statistics. Within this framework, and using various mathematical and algorithmic approximations, the AutoClass System searches for the most probable classifications, automatically choosing the number of classes and complexity of class descriptions. A simpler version of AutoClass has been applied to many large real data sets, has discovered new independently-verified phenomena, and has been released as a robust software package. Recent extensions allow attributes to be selectively correlated within particular classes, and allow classes to inherit, or share, model parameters through a class hierarchy. The mathematical foundations of AutoClass are summarized.

Stutz, John↗

Bayesian classification theory

The task of inferring a set of classes and class descriptions most likely to explain a given data set can be placed on a firm theoretical foundation using Bayesian statistics. Within this framework and using various mathematical and algorithmic approximations, the AutoClass system searches for the most probable classifications, automatically choosing the number of classes and complexity of class descriptions. A simpler version of AutoClass has been applied to many large real data sets, has discovered new independently-verified phenomena, and has been released as a robust software package. Recent extensions allow attributes to be selectively correlated within particular classes, and allow classes to inherit or share model parameters though a class hierarchy. We summarize the mathematical foundations of AutoClass.

Hanson, Robin↗

Assimilation of satellite microwave observations over the rainbands of tropical cyclones

A novel Bayesian Monte Carlo integration (BMCI) technique was developed to retrieve geophysical variables from satellite microwave radiometer data in the presence of tropical cyclones. The BMCI technique includes three steps: generating a stochastic database, simulating satellite brightness temperatures using a radiative transfer model, and retrieving geophysical variables such as profiles of temperature, relative humidity, and cloud liquid and ice water content from real observations. The technique also provides uncertainty estimates for each retrieval and can output the error covariance matrix of selected parameters. The measurements from the Advanced Technology Microwave Sounder (ATMS) on board Suomi National Polar-Orbiting Partnership (Suomi NPP) and the Global Precipitation Measurement (GPM) Microwave Imager (GMI) were used as input. A new technique was developed to correct the ATMS and GMI observations for the beam-filling effect, which is due to small-scale variability of precipitation and clouds when compared with the instrument footprint and also the nonlinear relation between the brightness temperature and precipitation. In addition, the assimilation of the BMCI retrievals into the NASA GEOS model is discussed for Hurricane Maria. The results show that assimilating the BMCI retrievals can influence the dynamical features of the cyclone, including a stronger warm core, a symmetric eye, and vertically aligned wind columns. Two possible factors that may limit the impact of the BMCI retrievals include 1) the resolution of the model (about 25 km), which was too coarse to show the potential of the BMCI data in improving the representation of tropical storms in the model forecast, and 2) the data assimilation system not being able to consider vertically correlated observation errors.

Isaac Moradi↗

Atacama Cosmology Telescope measurements of a large sample of candidates from the Massive and Distant Clusters of WISE Survey: Sunyaev-Zeldovich effect confirmation of MaDCoWS candidates using ACT

Context. Galaxy clusters are an important tool for cosmology, and their detection and characterization are key goals for current and future surveys. Using data from the Wide-field Infrared Survey Explorer (WISE), the Massive and Distant Clusters of WISE Survey (MaDCoWS) located 2839 significant galaxy overdensities at redshifts 0.7 . z . 1.5, which included extensive follow-up imaging from the Spitzer Space Telescope to determine cluster richnesses. Concurrently, the Atacama Cosmology Telescope (ACT) has produced large area millimeter-wave maps in three frequency bands along with a large catalog of Sunyaev-Zeldovich (SZ)-selected clusters as part of its Data Release 5 (DR5). Aims. We aim to verify and characterize MaDCoWS clusters using measurements of, or limits on, their thermal SZ effect signatures. We also use these detections to establish the scaling relation between SZ mass and the MaDCoWS-defined richness. Methods. Using the maps and cluster catalog from DR5, we explore the scaling between SZ mass and cluster richness. We do this by comparing cataloged detections and extracting individual and stacked SZ signals from the MaDCoWS cluster locations. We use complementary radio survey data from the Very Large Array, submillimeter data from Herschel, and ACT 224 GHz data to assess the impact of contaminating sources on the SZ signals from both ACT and MaDCoWS clusters. We use a hierarchical Bayesian model to fit the mass-richness scaling relation, allowing for clusters to be drawn from two populations: one, a Gaussian centered on the mass-richness relation, and the other, a Gaussian centered on zero SZ signal. Results. We find that MaDCoWS clusters have submillimeter contamination that is consistent with a gray-body spectrum, while the ACT clusters are consistent with no submillimeter emission on average. Additionally, the intrinsic radio intensities of ACT clusters are lower than those of MaDCoWS clusters, even when the ACT clusters are restricted to the same redshift range as the MaDCoWS clusters. We find the best-fit ACT SZ mass versus MaDCoWS richness scaling relation has a slope of p1 = 1.84+0.15 −0.14, where the slope is defined as M ∝ λ p1 15 and λ15 is the richness. We also find that the ACT SZ signals for a significant fraction (∼57%) of the MaDCoWS sample can statistically be described as being drawn from a noise-like distribution, indicating that the candidates are possibly dominated by low-mass and unvirialized systems that are below the mass limit of the ACT sample. Further, we note that a large portion of the optically confirmed ACT clusters located in the same volume of the sky as MaDCoWS are not selected by MaDCoWS, indicating that the MaDCoWS sample is not complete with respect to SZ selection. Finally, we find that the radio loud fraction of MaDCoWS clusters increases with richness, while we find no evidence that the submillimeter emission of the MaDCoWS clusters evolves with richness. Conclusions. We conclude that the original MaDCoWS selection function is not well defined and, as such, reiterate the MaDCoWS collaboration’s recommendation that the sample is suited for probing cluster and galaxy evolution, but not cosmological analyses. We find a best-fit mass-richness relation slope that agrees with the published MaDCoWS preliminary results. Additionally, we find that while the approximate level of infill of the ACT and MaDCoWS cluster SZ signals (1–2%) is subdominant to other sources of uncertainty for current generation experiments, characterizing and removing this bias will be critical for next-generation experiments hoping to constrain cluster masses at the sub-percent level.

large↗

Bayesian image reconstruction - The pixon and optimal image modeling

In this paper we describe the optimal image model, maximum residual likelihood method (OptMRL) for image reconstruction. OptMRL is a Bayesian image reconstruction technique for removing point-spread function blurring. OptMRL uses both a goodness-of-fit criterion (GOF) and an 'image prior', i.e., a function which quantifies the a priori probability of the image. Unlike standard maximum entropy methods, which typically reconstruct the image on the data pixel grid, OptMRL varies the image model in order to find the optimal functional basis with which to represent the image. We show how an optimal basis for image representation can be selected and in doing so, develop the concept of the 'pixon' which is a generalized image cell from which this basis is constructed. By allowing both the image and the image representation to be variable, the OptMRL method greatly increases the volume of solution space over which the image is optimized. Hence the likelihood of the final reconstructed image is greatly increased. For the goodness-of-fit criterion, OptMRL uses the maximum residual likelihood probability distribution introduced previously by Pina and Puetter (1992). This GOF probability distribution, which is based on the spatial autocorrelation of the residuals, has the advantage that it ensures spatially uncorrelated image reconstruction residuals.

Pina, R. K.↗

The M-dwarf Ultraviolet Spectroscopic Sample. I. Determining Stellar Parameters for Field Stars

Accurate stellar properties are essential for precise stellar astrophysics and exoplanetary science. In the M-dwarf regime, much effort has gone into defining empirical relations that can use readily accessible observables to assess physical stellar properties. Often, these relations for the quantity of interest are cast as a nonlinear function of available data; in Bayesian modeling, however, the reverse is needed. In this article, we introduce a new Bayesian framework to self-consistently and simultaneously apply multiple empirical calibrations to fully characterize the mass, luminosity, radius, and effective temperature of a field age M-dwarf. This framework includes a new M-dwarf mass–radius relation with a scatter of 3.1% at fixed mass. We further introduce the M-dwarf Ultraviolet Spectroscopic Sample (MUSS), and apply our methodology to provide consistent stellar parameters for these nearby low-mass stars, selected as having available spectroscopic data in the ultraviolet. These targets are of interest largely as either exoplanet hosts or benchmarks in multiwavelength stellar activity. We use the field MUSS stars to define a low-mass main sequence in the solar neighborhood through Gaussian Process (GP) regression. These results enable us to empirically measure a feature in the GP derivative at M ⊙ that indicates where the MUSS transitions from fully to partly convective interiors.

J. Sebastian Pineda↗

Efficient Calibration of Expensive Computational Models

Accounting for uncertainty when calibrating expensive computational models is a common challenge faced by scientists and engineers. Often Bayesian techniques are adopted to estimate a probability density function over the model parameters given noisy empirical data. The methods used to perform this type of probabilistic calibration are computationally prohibitive in that they require a large number of evaluations of the expensive model. In these cases, surrogate modeling -- that is, using a fast-to-evaluate, lower fidelity stand-in for the original computational model -- may be the only option to alleviate this computational burden. However, the upfront cost of generating training data to build a surrogate model can itself be expensive. As such, it is important to be judicious when selecting training points at which the full-fidelity model is evaluated. Here, an active learning approach is proposed that enables efficient selection of training points using approximate samples of the calibrated parameter probability density function. In this way, the training points can be concentrated in regions where the calibration algorithm requires high model accuracy.

active learning↗

Astrophysical Model Selection in Gravitational Wave Astronomy

Theoretical studies in gravitational wave astronomy have mostly focused on the information that can be extracted from individual detections, such as the mass of a binary system and its location in space. Here we consider how the information from multiple detections can be used to constrain astrophysical population models. This seemingly simple problem is made challenging by the high dimensionality and high degree of correlation in the parameter spaces that describe the signals, and by the complexity of the astrophysical models, which can also depend on a large number of parameters, some of which might not be directly constrained by the observations. We present a method for constraining population models using a hierarchical Bayesian modeling approach which simultaneously infers the source parameters and population model and provides the joint probability distributions for both. We illustrate this approach by considering the constraints that can be placed on population models for galactic white dwarf binaries using a future space-based gravitational wave detector. We find that a mission that is able to resolve approximately 5000 of the shortest period binaries will be able to constrain the population model parameters, including the chirp mass distribution and a characteristic galaxy disk radius to within a few percent. This compares favorably to existing bounds, where electromagnetic observations of stars in the galaxy constrain disk radii to within 20%.

hierarchical↗

Bayesian Geostatistical Modelling of PM10 and PM2.5 Surface Level Concentrations in Europe Using High-Resolution Satellite-Derived Products

Air quality monitoring across Europe is mainly based on in situ ground stations which are too sparse to accurately assess the exposure effects of air pollution for the entire continent. The demand for precise predictive modelsthat estimate gridded geophysical parameters of ambient air at high spatial resolution has rapidly grown. Here, we investigate the potential of satellite derived products to improve particulate matter (PM) estimates. Bayesiangeostatistical models addressing confounding between the spatial distribution of pollutants and remotely sensed predictors were developed to estimate yearly averages of both, fine (PM2.5) and coarse (PM10) surface PM concentrations at 1 sq.km spatial resolution over 46 European countries and were compared to geostatistical, geographically weighted and land-use regression formulations. Rigorous model selection identified the Earth observation data which contribute most to pollutants' estimation. Geostatistical models outperformed the predictive ability of the frequently employed land-use regression. The resulting estimates of PM10 and PM2.5, which represent the main air quality indicators for the urban Sustainable Development Goal, indicate that in 2016, 66.2% of the European population was breathing air above the WHO Air Quality Guidelines thresholds. Our estimates are readily available to policy makers and scientists assessing the effects of long-term exposure to pollution on human and ecosystem health.

Beloconi, Anton↗