Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Model selection”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Factorization of Binary Matrices: Rank Relations, Uniqueness and Model Selection of Boolean Decomposition

The application of binary matrices are numerous. Representing a matrix as a mixture of a small collection of latent vectors via low-rank decomposition is often seen as an advantageous method to interpret and analyze data. In this work, we examine the factorizations of binary matrices using standard arithmetic (real and nonnegative) and logical operations (Boolean and $\mathbb{Z}$ 2 ). We examine the relationships between the different ranks, and discuss when factorization is unique. In particular, we characterize when a Boolean factorization X = W$\land$H has a unique W, a unique H (for a fixed W), and when both W and H are unique, given a rank constraint. We introduce a method for robust Boolean model selection, called BMFk, and show on numerical examples that BMFk not only accurately determines the correct number of Boolean latent features but reconstruct the pre-determined factors accurately.

97 MATHEMATICS AND COMPUTING↗

Bayesian Model Selection for Reducing Bloat and Overfitting in Genetic Programming for Symbolic Regression

When performing symbolic regression using genetic programming, overfitting and bloat can negatively impact generalizability and interpretability of the resulting equations as well as increase computation times. A Bayesian fitness metric is introduced and its impact on bloat and overfitting during population evolution is studied and compared to common alternatives in the literature. The proposed approach was found to be more robust to noise and data sparsity in numerical experiments, guiding evolution to a level of complexity appropriate to the dataset. Further evolution of the population resulted not in overfitting or bloat, but rather in slight simplifications in model form. The ability to identify an equation of complexity appropriate to the scale of noise in the training data was also demonstrated. In general, the Bayesian model selection algorithm was shown to be an effective means of regularization which resulted in less bloat and overfitting when any amount of noise was present in the training data.

G F Bomarito↗

Application of a Bayesian Framework for Plasticity Model Selection

Interpretable Machine Learning (IML) has performed well when tasked with deriving constitutive material models. However, IML has been shown to prefer models that overfit noise in data, which tends to lead to bloat and a decrease in interpretability. Due to these issues, the ability of IML to reliably derive models that fit the data and are both interpretable and generalizable is limited. A method developed recently has shown promise to improve upon traditional IML by using a Bayesian fitness definition for the evolution of free-form models with non-deterministic parameters. This framework was developed for genetic-programming-based symbolic regression(GPSR) and involves model parameter estimation using Sequential Monte Carlo sampling (SMC).The method has demonstrated a reduction in bloat when dealing with noisy data in comparison to conventional GPSR. The results of this framework applied to stress-strain data for copper show models that more effectively predict the experimental data better than was previously shown with GPSR.

plasticity↗

Modeling Selective Availability of the NAVSTAR Global Positioning System

As the development of the NAVSTAR Global Positioning System (GPS) continues, there will increasingly be the need for a software centered signal model. This model must accurately generate the observed pseudorange which would typically be encountered. The observed pseudorange varies from the true geometric (slant) range due to range measurement errors. Errors in range measurement stem from a variety of hardware and environment factors. These errors are classified as either deterministic or random and, where appropriate, their models are summarized. Of particular interest is the model for Selective Availability which is derived from actual GPS data. The procedure for the determination of this model, known as the System Identification Theory, is briefly outlined. The synthesis of these error sources into the final signal model is given along with simulation results.

Braasch, Michael↗

Hybrid Parameter Search and Dynamic Model Selection for Mixed-Variable Bayesian Optimization

Herein this article presents a new type of hybrid model for Bayesian optimization (BO) adept at managing mixed variables, encompassing both quantitative (continuous and integer) and qualitative (categorical) types. Our proposed new hybrid models (named hybridM) merge the Monte Carlo Tree Search structure (MCTS) for categorical variables with Gaussian Processes (GP) for continuous ones. hybridM leverages the upper confidence bound tree search (UCTS) for MCTS strategy, showcasing the tree architecture’s integration into Bayesian optimization. Our innovations, including dynamic online kernel selection in the surrogate modeling phase and a unique UCTS search strategy, position our hybrid models as an advancement in mixed-variable surrogate models. Numerical experiments underscore the superiority of hybrid models, highlighting their potential in Bayesian optimization.

97 MATHEMATICS AND COMPUTING↗

Sequential Decision Making (SDM) for Mesh Refinement and Model Selection in Multiscale, Multi-Physics Applications

Intelligent automation and decision support are needed to enhance computational efficiency and robustness in multiscale and multi-physics problems, including materials science, manufacturing, and climate and weather modeling. Current scientific computing approaches for enabling decisions by scientists fail to explore the role of learning, reasoning, and probabilistic planning. Often these decisions are not performed in real-time during the computation but are made prior to the start of the computation, which must be interrupted in order to make changes to the prior choices. Such interruptions at different stages of the computation increase the total computing time and the need for a human expert to frequently monitor the results. State of art scientific computing methods consist of rule-based algorithms that cannot automatically adapt to a dynamically changing computing environment. The development of a Sequential Decision Making (SDM) framework will automate scientific computing by optimizing the policies for mesh refinement, time-stepping, model and algorithm selection, resource allocation, and pre and post-processing. Our agent SDM framework for scientific computing will consist of data-driven learning (Classifier), automated reasoning (contextual knowledge), and probabilistic planning (Reinforcement Learning). In this project, we focused on three problems to demonstrate our SDM framework on a set of ordinary and partial differential equations. Classification of Lorenz system regions using Feed-Forward Neural Networks examined learning in the SDM framework. On the other hand, reasoning and planning in the SDM framework were used in two problems: adaptive time-stepping for nonlinear ODEs using on-policy RL algorithms, and adaptive mesh refinement for 2-D PDEs using off-policy RL algorithms.

97 MATHEMATICS AND COMPUTING↗

Development of a research project selection model: Application to a civil helicopter research program

A model is described for planning and decision making in research project selection. Evaluations of each project's direct and indirect benefits, uncertainty in achieving these benefits, and schedule priority with resource budget and program balance constraints are considered. The combination of the interactive effect of project selection, resource allocation and scheduling considerations into one model permits tradeoff alternatives to be studied. Clients' value judgments are used in evaluating the benefits from each proposed project. The model is applied to the NASA Civil Helicopter Technology Program. Research project priorities for this program are established, strengths and weaknesses of the model are discussed, and areas of future development are recommended.

Schoultz, M. B.↗

The Dark Energy Survey Supernova Programme: Modelling Selection Efficiency and Observed Core-collapse Supernova Contamination

The analysis of current and future cosmological surveys of Type Ia supernovae (SNe Ia) at high redshift depends on the accuratephotometric classification of the SN events detected. Generating realistic simulations of photometric SN surveys constitutes anessential step for training and testing photometric classification algorithms, and for correcting biases introduced by selectioneffects and contamination arising from core-collapse SNe in the photometric SN Ia samples. We use published SN time-seriesspectrophotometric templates, rates, luminosity functions, and empirical relationships between SNe and their host galaxies toconstruct a framework for simulating photometric SN surveys. We present this framework in the context of the Dark EnergySurvey (DES) 5-yr photometric SN sample, comparing our simulations of DES with the observed DES transient populations.We demonstrate excellent agreement in many distributions, including Hubble residuals, between our simulations and data.We estimate the core collapse fraction expected in the DES SN sample after selection requirements are applied and beforephotometric classification. After testing different modelling choices and astrophysical assumptions underlying our simulation,we find that the predicted contamination varies from 7.2 to 11.7 per cent, with an average of 8.8 per cent and an r.m.s. of 1.1 percent. Our simulations are the first to reproduce the observed photometric SN and host galaxy properties in high-redshift surveyswithout fine-tuning the input parameters. The simulation methods presented here will be a critical component of the cosmologyanalysis of the DES photometric SN Ia sample: correcting for biases arising from contamination, and evaluating the associatedsystematic uncertainty.

M Vincenzi↗

Probing Sensitivity of Discharge Characteristics to Model Selection using Uncertainty Quantification in an aprotic Li-Oxygen Battery

Currently, there are several models in the literature, such as kinetic models, microstructural models, and mass transport models that describe a Li-air battery's discharge behavior. Many of these models are calibrated and tested at low current densities and cannot be easily transferred to high current densities. Even at low current densities, there is no quantitative method for a researcher to choose a reaction kinetic model such as classical Butler-Volmer and its derivatives, and modified Marcus-Hush-Chidsey, a resistance model for lithium peroxide such as electron transport via tunneling or linear resistivity, a surface coverage model (lithium peroxide growth) such as partial coverage or full coverage, and mass transport model (discussed in Ref. [1]). Also, it is time-consuming to test different models at high current density (1C) due to a lack of well-tested models and well-calibrated model parameters. For this presentation, we will develop an analytical model, which acts as a surrogate model for a sophisticated finite element model to predict discharge time and discharge voltage. Next, we use an uncertainty quantifying technique called reduced-order stochastic optimization [2, 3] to determine the uncertainty in model parameters for rate kinetics, lithium peroxide resistivity, and parasitic resistance. Finally, a finite element simulation is performed to determine the error introduced by the surrogate model and its influence on the uncertainty in the model parameters.

M Mehta↗

The Dark Energy Survey supernova programme: modelling selection efficiency and observed core-collapse supernova contamination

ABSTRACT The analysis of current and future cosmological surveys of Type Ia supernovae (SNe Ia) at high redshift depends on the accurate photometric classification of the SN events detected. Generating realistic simulations of photometric SN surveys constitutes an essential step for training and testing photometric classification algorithms, and for correcting biases introduced by selection effects and contamination arising from core-collapse SNe in the photometric SN Ia samples. We use published SN time-series spectrophotometric templates, rates, luminosity functions, and empirical relationships between SNe and their host galaxies to construct a framework for simulating photometric SN surveys. We present this framework in the context of the Dark Energy Survey (DES) 5-yr photometric SN sample, comparing our simulations of DES with the observed DES transient populations. We demonstrate excellent agreement in many distributions, including Hubble residuals, between our simulations and data. We estimate the core collapse fraction expected in the DES SN sample after selection requirements are applied and before photometric classification. After testing different modelling choices and astrophysical assumptions underlying our simulation, we find that the predicted contamination varies from 7.2 to 11.7 per cent, with an average of 8.8 per cent and an r.m.s. of 1.1 per cent. Our simulations are the first to reproduce the observed photometric SN and host galaxy properties in high-redshift surveys without fine-tuning the input parameters. The simulation methods presented here will be a critical component of the cosmology analysis of the DES photometric SN Ia sample: correcting for biases arising from contamination, and evaluating the associated systematic uncertainty.

79 ASTRONOMY AND ASTROPHYSICS↗

Diagnosing Hybrid Systems: a Bayesian Model Selection Approach

In this paper we examine the problem of monitoring and diagnosing noisy complex dynamical systems that are modeled as hybrid systems-models of continuous behavior, interleaved by discrete transitions. In particular, we examine continuous systems with embedded supervisory controllers that experience abrupt, partial or full failure of component devices. Building on our previous work in this area (MBCG99;MBCG00), our specific focus in this paper ins on the mathematical formulation of the hybrid monitoring and diagnosis task as a Bayesian model tracking algorithm. The nonlinear dynamics of many hybrid systems present challenges to probabilistic tracking. Further, probabilistic tracking of a system for the purposes of diagnosis is problematic because the models of the system corresponding to failure modes are numerous and generally very unlikely. To focus tracking on these unlikely models and to reduce the number of potential models under consideration, we exploit logic-based techniques for qualitative model-based diagnosis to conjecture a limited initial set of consistent candidate models. In this paper we discuss alternative tracking techniques that are relevant to different classes of hybrid systems, focusing specifically on a method for tracking multiple models of nonlinear behavior simultaneously using factored sampling and conditional density propagation. To illustrate and motivate the approach described in this paper we examine the problem of monitoring and diganosing NASA's Sprint AERCam, a small spherical robotic camera unit with 12 thrusters that enable both linear and rotational motion.

McIlraith, Sheila A.↗

Effect of model selection on combustor performance and stability using ROCCID

The ROCket Combustor Interactive Design (ROCCID) methodology is an interactive computer program that combines previously developed combustion analysis models to calculate the combustion performance and stability of liquid rocket engines. Test data from a 213 kN (48,000 lbf) Liquid Oxygen (LOX)/RP-1 combustor with a O-F-O (oxidizer-fuel-oxidizer) triplet injector were used to characterize the predictive capabilities of the ROCCID analysis models for this injector/propellant configuration. Thirteen combustion performance and stability models have been incorporated into ROCCID, and ten of them, which have options for triplet injectors, were examined in this study. Calculations using different combinations of analysis models, with little or no anchoring, were carried out on a test matrix of operating conditions matching those of the test program. Results of the computer analyses were compared to test data, and the ability of the model combinations to correctly predict combustion stability or instability was determined. For the best model combination(s), sensitivity of the calculations to fuel drop size and mixing efficiency was examined. Error in the stability calculations due to uncertainty in the pressure interaction index (N) was examined. The recommended model combinations for this O-F-O triplet LOX/RP-1 configuration are proposed.

Giuliani, James E.↗

Effect of model selection on combustor performance and stability predictions using ROCCID

The ROCket Combustor Interactive Design (ROCCID) methodology is an interactive computer program that combines previously developed combustion analysis models to calculate the combustion performance and stability of liquid rocket engines. Test data from 213 kN (48,000 lbf) Liquid Oxygen (LOX)/RP-1 combustor with an O-F-O (oxidizer-fuel-oxidizer) triplet injector were used to characterize the predictive capabilities of the ROCCID analysis models for this injector/propellant configuration. Thirteen combustion performance and stability models were incorporated into ROCCID, and ten of them, which have options for triplet injectors, were examined. Calculations using different combinations of analysis models, with little or no anchoring, were carried out on a test matrix of operating combinations matching those of the test program. Results of the computer analyses were compared to test data, and the ability of the model combinations to correctly predict combustion stability or instability was determined. For the best model combination(s), sensitivity of the calculations to fuel drop size and mixing efficiency was examined. Error in the stability calculations due to uncertainty in the pressure interaction index (N) was examined. The recommended model combinations for this O-F-O triplet LOX/RP-1 configuration are proposed.

Giuliani, James E.↗

Influence of World and Gravity Model Selection on Surface Interacting Vehicle Simulations

A vehicle simulation is surface-interacting if the state of the vehicle (position, velocity, and acceleration) relative to the surface is important. Surface-interacting simulations perform ascent, entry, descent, landing, surface travel, or atmospheric flight. Modeling of gravity is an influential environmental factor for surface-interacting simulations. Gravity is the free-fall acceleration observed from a world-fixed frame that rotates with the world. Thus, gravity is the sum of gravitation and the centrifugal acceleration due to the world s rotation. In surface-interacting simulations, the fidelity of gravity at heights above the surface is more significant than gravity fidelity at locations in inertial space. A surface-interacting simulation cannot treat the gravity model separately from the world model, which simulates the motion and shape of the world. The world model's simulation of the world's rotation, or lack thereof, produces the centrifugal acceleration component of gravity. The world model s reproduction of the world's shape will produce different positions relative to the world center for a given height above the surface. These differences produce variations in the gravitation component of gravity. This paper examines the actual performance of world and gravity/gravitation pairs in a simulation using the Earth.

Madden, Michael M.↗

Effects of turbulence model selection on the prediction of complex aerodynamic flows

Numerical simulations of viscous transonic flow over a circular-arc airfoil and in a diffuser are described. The simulations are made with a new computer program designed to serve as a tool in the development of improved turbulence models for complex flows. The program incorporates zero-, one-, and two-equation eddy viscosity models and includes a variety of subsonic and supersonic boundary conditions. The airfoil flow contains a shock-separated boundary-layer interaction that has resisted previous attempts at simulation. The diffuser flow also contains a shock-boundary-layer interaction, which has not been simulated previously. Calculations using standard turbulence models, developed originally for incompressible unseparated flows, are described. Results indicate that although there are interesting differences in predictions between the various models, none of them predict the flows accurately. Suggestions for improved turbulence models are discussed.

Coakley, T. J.↗

Astrophysical Model Selection in Gravitational Wave Astronomy

Theoretical studies in gravitational wave astronomy have mostly focused on the information that can be extracted from individual detections, such as the mass of a binary system and its location in space. Here we consider how the information from multiple detections can be used to constrain astrophysical population models. This seemingly simple problem is made challenging by the high dimensionality and high degree of correlation in the parameter spaces that describe the signals, and by the complexity of the astrophysical models, which can also depend on a large number of parameters, some of which might not be directly constrained by the observations. We present a method for constraining population models using a hierarchical Bayesian modeling approach which simultaneously infers the source parameters and population model and provides the joint probability distributions for both. We illustrate this approach by considering the constraints that can be placed on population models for galactic white dwarf binaries using a future space-based gravitational wave detector. We find that a mission that is able to resolve approximately 5000 of the shortest period binaries will be able to constrain the population model parameters, including the chirp mass distribution and a characteristic galaxy disk radius to within a few percent. This compares favorably to existing bounds, where electromagnetic observations of stars in the galaxy constrain disk radii to within 20%.

hierarchical↗

Model selection and signal extraction using Gaussian Process regression

We present a novel computational approach for extracting localized signals from smooth background distributions. We focus on datasets that can be naturally presented as binned integer counts, demonstrating our procedure on the CERN open dataset with the Higgs boson signature, from the ATLAS collaboration at the Large Hadron Collider. Our approach is based on Gaussian Process (GP) regression — a powerful and flexible machine learning technique which has allowed us to model the background without specifying its functional form explicitly and separately measure the background and signal contributions in a robust and reproducible manner. Unlike functional fits, our GP-regression-based approach does not need to be constantly updated as more data becomes available. We discuss how to select the GP kernel type, considering trade-offs between kernel complexity and its ability to capture the features of the background distribution. We show that our GP framework can be used to detect the Higgs boson resonance in the data with more statistical significance than a polynomial fit specifically tailored to the dataset. Finally, we use Markov Chain Monte Carlo (MCMC) sampling to confirm the statistical significance of the extracted Higgs signature.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗