Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Model selection”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

LASSO for CALPHAD Model Selection Enables Data-Efficient Thermodynamic Modeling: An Application in Thermochemical Hydrogen Production Materials

Phenomenological CALPHAD (CALculation of PHAse Diagrams) models, widely used for multicomponent materials, often contain a considerable number of parameters and require fitting using data from a relatively small number of experimental measurements or theoretical calculations. Sometimes these parameters are introduced for the purpose of improving model fits but without clear physical justification, which leads to overparametrized models with poor generalization performance. Automated approaches for optimal model selection based on the available data therefore become critical. Here, in this work, a least absolute shrinkage and selection operator (LASSO)-based approach is developed for model selection by leveraging the linearity of the CALPHAD model with respect to its parameters to convert the model selection and fitting to a LASSO minimization problem. We demonstrate its utility for thermodynamic modeling of thermochemical hydrogen (TCH) production materials using lanthanum strontium manganite (LSM) as an example. Various TCH-relevant properties, including oxygen stoichiometry as a function of oxygen partial pressure, enthalpy of reduction, and entropy of reduction, are successfully predicted with reasonable accuracy using a minimal set of model parameters. Importantly, the model selection and fitting involve minimal human decision; it can therefore be applied to high-throughput DFT defect calculations and yield efficient workflows for TCH material modeling and optimization.

CALPHAD↗

Contextual Active Online Model Selection with Expert Advice

How can we collect the most useful labels to learn a model selection policy, when presented with arbitrary heterogeneous data streams? In this paper, we formulate this task as a contextual active model selection problem, where at each round the learner receives an unlabeled data point along with a context. The goal is to output the best model for any given context without obtaining an excessive amount of labels. In particular, we focus on the task of selecting pre-trained classifiers, and propose a contextual active model selection algorithm (CAMS), which relies on a novel uncertainty sampling query criterion defined on a given policy class for adaptive model selection. In comparison to prior art, our algorithm does not assume a globally optimal model. We provide rigorous theoretical analysis for the regret and query complexity under both adversarial and stochastic settings. Our experiments on several benchmark classification datasets demonstrate the algorithm’s effectiveness in terms of both regret and query complexity. Notably, to achieve the same accuracy, CAMS incurs less than 10% of the label cost when compared to the best online model selection baselines on CIFAR10.

Liu, Xuefeng↗

An empirical approach to model selection: weak lensing and intrinsic alignments

ABSTRACT In cosmology, we routinely choose between models to describe our data, and can incur biases due to insufficient models or lose constraining power with overly complex models. In this paper, we propose an empirical approach to model selection that explicitly balances parameter bias against model complexity. Our method uses synthetic data to calibrate the relation between bias and the χ2 difference between models. This allows us to interpret χ2 values obtained from real data (even if catalogues are blinded) and choose a model accordingly. We apply our method to the problem of intrinsic alignments – one of the most significant weak lensing systematics, and a major contributor to the error budget in modern lensing surveys. Specifically, we consider the example of the Dark Energy Survey Year 3 (DES Y3), and compare the commonly used non-linear alignment (NLA) and tidal alignment and tidal torque (TATT) models. The models are calibrated against bias in the Ωm–S8 plane. Once noise is accounted for, we find that it is possible to set a threshold Δχ2 that guarantees an analysis using NLA is unbiased at some specified level Nσ and confidence level. By contrast, we find that theoretically defined thresholds (based on, e.g. p-values for χ2) tend to be overly optimistic, and do not reliably rule out cosmological biases up to ∼1–2σ. Considering the real DES Y3 cosmic shear results, based on the reported difference in χ2 from NLA and TATT analyses, we find a roughly $30{{\ \rm per\ cent}}$ chance that were NLA to be the fiducial model, the results would be biased (in the Ωm–S8 plane) by more than 0.3σ. More broadly, the method we propose here is simple and general, and requires a relatively low level of resources. We foresee applications to future analyses as a model selection tool in many contexts.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

A meta-learning based distribution system load forecasting model selection framework

This paper presents a meta-learning based, automatic distribution system load forecasting model selection framework. Furthermore, the framework includes the following processes: feature extraction, candidate model preparation and labeling, offline training, and online model recommendation. Using load forecasting needs and data characteristics as input features, multiple metalearners are used to rank the candidate load forecast models based on their forecasting accuracy. Then, a scoring-voting mechanism is proposed to weights recommendations from each meta-leaner and make the final recommendations. Heterogeneous load forecasting tasks with different temporal and technical requirements at different load aggregation levels are set up to train, validate, and test the performance of the proposed framework. Simulation results demonstrate that the performance of the meta-learning based approach is satisfactory in both seen and unseen forecasting tasks.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Multiverse: Bayesian model selection for neural networks

Multiverse is a code repository for a set of tools for Bayesian model selection for neural networks. The goal of the tools is to provide capabilities for selecting among prior and model specifications for Bayesian neural networks. Multiverse will include tools for creating Bayesian neural networks, evaluating the Bayesian model evidence, and performing inference in Bayesian neural networks. These components are written in Python, a high-level programming language that takes advantage of the Python ecosystem of high-quality open-source packages for machine learning and signal processing.

Klein, Natalie↗

AT2017gfo: Bayesian inference and model selection of multicomponent kilonovae and constraints on the neutron star equation of state

The joint detection of the gravitational wave GW170817, of the short γ-ray burst GRB170817A and of the kilonova AT2017gfo, generated by the the binary neutron star (NS) merger observed on 2017 August 17, is a milestone in multimessenger astronomy and provides new constraints on the NS equation of state. We perform Bayesian inference and model selection on AT2017gfo using semi-analytical, multicomponents models that also account for non-spherical ejecta. Observational data favour anisotropic geometries to spherically symmetric profiles, with a log-Bayes’ factor of ~104, and favour multicomponent models against single-component ones. The best-fitting model is an anisotropic three-component composed of dynamical ejecta plus neutrino and viscous winds. Using the dynamical ejecta parameters inferred from the best-fitting model and numerical–relativity relations connecting the ejecta properties to the binary properties, we constrain the binary mass ratio to q < 1.54 and the reduced tidal parameter to $120\lt \tilde{\Lambda }\lt 1110$. Finally, we combine the predictions from AT2017gfo with those from GW170817, constraining the radius of a NS of 1.4 M ⊙ to 12.2 ± 0.5 km (1σ level). This prediction could be further strengthened by improving kilonova models with numerical-relativity information.

79 ASTRONOMY AND ASTROPHYSICS↗

Machine-learning interatomic potentials for interfaces in all-solid-state batteries: Perspectives on training data, model selection, and validation

Interfaces play a pivotal role in dictating the performance and reliability of all-solid-state batteries (ASSBs), where complex electro-chemo-mechanical phenomena at grain boundaries (GBs) and interfaces can lead to degradation and failure. Traditional atomistic simulation methods, such as first-principles calculations and classical molecular dynamics, face limitations in modeling these interfaces due to either high computational cost or insufficient transferability to the diverse atomic environments evolving at interfaces. Machine-learning interatomic potentials (MLIPs) have emerged as a transformative approach, enabling large-scale, high-accuracy simulations of disordered and chemically complex systems by leveraging the predictability of machine learning models trained on first-principles data. Recent applications of MLIPs have demonstrated their ability to capture intricate behaviors at ASSB interfaces, including ion transport, interfacial evolution, and degradation mechanisms, with accuracy and efficiency unattainable by conventional methods. This prospective paper presents comprehensive analysis and practical guidance for MLIP development for GBs and interfaces in ASSBs, with a focus on three key pillars: data generation, model selection, and validation. Here, we review the current state of MLIP applications for GBs and interfaces in both general and ASSB-specific materials, highlighting best practices and challenges in constructing diverse and representative datasets, choosing appropriate machine learning architectures, and rigorously validating model performance. We also discuss emerging strategies and opportunities for improved reliability and efficiency of MLIPs to simulate realistic interfaces in ASSBs.

Energy - Storage↗

Factorization of Binary Matrices: Rank Relations, Uniqueness and Model Selection of Boolean Decomposition

The application of binary matrices are numerous. Representing a matrix as a mixture of a small collection of latent vectors via low-rank decomposition is often seen as an advantageous method to interpret and analyze data. In this work, we examine the factorizations of binary matrices using standard arithmetic (real and nonnegative) and logical operations (Boolean and $\mathbb{Z}$ 2 ). We examine the relationships between the different ranks, and discuss when factorization is unique. In particular, we characterize when a Boolean factorization X = W$\land$H has a unique W, a unique H (for a fixed W), and when both W and H are unique, given a rank constraint. We introduce a method for robust Boolean model selection, called BMFk, and show on numerical examples that BMFk not only accurately determines the correct number of Boolean latent features but reconstruct the pre-determined factors accurately.

97 MATHEMATICS AND COMPUTING↗

Hybrid Parameter Search and Dynamic Model Selection for Mixed-Variable Bayesian Optimization

Herein this article presents a new type of hybrid model for Bayesian optimization (BO) adept at managing mixed variables, encompassing both quantitative (continuous and integer) and qualitative (categorical) types. Our proposed new hybrid models (named hybridM) merge the Monte Carlo Tree Search structure (MCTS) for categorical variables with Gaussian Processes (GP) for continuous ones. hybridM leverages the upper confidence bound tree search (UCTS) for MCTS strategy, showcasing the tree architecture’s integration into Bayesian optimization. Our innovations, including dynamic online kernel selection in the surrogate modeling phase and a unique UCTS search strategy, position our hybrid models as an advancement in mixed-variable surrogate models. Numerical experiments underscore the superiority of hybrid models, highlighting their potential in Bayesian optimization.

97 MATHEMATICS AND COMPUTING↗

Sequential Decision Making (SDM) for Mesh Refinement and Model Selection in Multiscale, Multi-Physics Applications

Intelligent automation and decision support are needed to enhance computational efficiency and robustness in multiscale and multi-physics problems, including materials science, manufacturing, and climate and weather modeling. Current scientific computing approaches for enabling decisions by scientists fail to explore the role of learning, reasoning, and probabilistic planning. Often these decisions are not performed in real-time during the computation but are made prior to the start of the computation, which must be interrupted in order to make changes to the prior choices. Such interruptions at different stages of the computation increase the total computing time and the need for a human expert to frequently monitor the results. State of art scientific computing methods consist of rule-based algorithms that cannot automatically adapt to a dynamically changing computing environment. The development of a Sequential Decision Making (SDM) framework will automate scientific computing by optimizing the policies for mesh refinement, time-stepping, model and algorithm selection, resource allocation, and pre and post-processing. Our agent SDM framework for scientific computing will consist of data-driven learning (Classifier), automated reasoning (contextual knowledge), and probabilistic planning (Reinforcement Learning). In this project, we focused on three problems to demonstrate our SDM framework on a set of ordinary and partial differential equations. Classification of Lorenz system regions using Feed-Forward Neural Networks examined learning in the SDM framework. On the other hand, reasoning and planning in the SDM framework were used in two problems: adaptive time-stepping for nonlinear ODEs using on-policy RL algorithms, and adaptive mesh refinement for 2-D PDEs using off-policy RL algorithms.

97 MATHEMATICS AND COMPUTING↗

The Dark Energy Survey supernova programme: modelling selection efficiency and observed core-collapse supernova contamination

ABSTRACT The analysis of current and future cosmological surveys of Type Ia supernovae (SNe Ia) at high redshift depends on the accurate photometric classification of the SN events detected. Generating realistic simulations of photometric SN surveys constitutes an essential step for training and testing photometric classification algorithms, and for correcting biases introduced by selection effects and contamination arising from core-collapse SNe in the photometric SN Ia samples. We use published SN time-series spectrophotometric templates, rates, luminosity functions, and empirical relationships between SNe and their host galaxies to construct a framework for simulating photometric SN surveys. We present this framework in the context of the Dark Energy Survey (DES) 5-yr photometric SN sample, comparing our simulations of DES with the observed DES transient populations. We demonstrate excellent agreement in many distributions, including Hubble residuals, between our simulations and data. We estimate the core collapse fraction expected in the DES SN sample after selection requirements are applied and before photometric classification. After testing different modelling choices and astrophysical assumptions underlying our simulation, we find that the predicted contamination varies from 7.2 to 11.7 per cent, with an average of 8.8 per cent and an r.m.s. of 1.1 per cent. Our simulations are the first to reproduce the observed photometric SN and host galaxy properties in high-redshift surveys without fine-tuning the input parameters. The simulation methods presented here will be a critical component of the cosmology analysis of the DES photometric SN Ia sample: correcting for biases arising from contamination, and evaluating the associated systematic uncertainty.

79 ASTRONOMY AND ASTROPHYSICS↗

Model selection and signal extraction using Gaussian Process regression

We present a novel computational approach for extracting localized signals from smooth background distributions. We focus on datasets that can be naturally presented as binned integer counts, demonstrating our procedure on the CERN open dataset with the Higgs boson signature, from the ATLAS collaboration at the Large Hadron Collider. Our approach is based on Gaussian Process (GP) regression — a powerful and flexible machine learning technique which has allowed us to model the background without specifying its functional form explicitly and separately measure the background and signal contributions in a robust and reproducible manner. Unlike functional fits, our GP-regression-based approach does not need to be constantly updated as more data becomes available. We discuss how to select the GP kernel type, considering trade-offs between kernel complexity and its ability to capture the features of the background distribution. We show that our GP framework can be used to detect the Higgs boson resonance in the data with more statistical significance than a polynomial fit specifically tailored to the dataset. Finally, we use Markov Chain Monte Carlo (MCMC) sampling to confirm the statistical significance of the extracted Higgs signature.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Bayesian model selection for GRB 211211A through multiwavelength analyses

ABSTRACT Although GRB 211211A is one of the closest gamma-ray bursts (GRBs), its classification is challenging because of its partially inconclusive electromagnetic signatures. In this paper, we investigate four astrophysical scenarios as possible progenitors for GRB 211211A: a binary neutron star merger, a black hole–neutron star merger, a core-collapse supernova, and an r-process enriched core collapse of a rapidly rotating massive star (a collapsar). We perform a large set of Bayesian multiwavelength analyses based on different models describing these scenarios and priors to investigate which astrophysical scenarios and processes might be related to GRB 211211A. Our analysis supports previous studies in which the presence of an additional component, likely related to r-process nucleosynthesis, is required to explain the observed light curves of GRB 211211A, as it cannot be explained solely as a GRB afterglow. Fixing the distance to about $350~\rm Mpc$, namely the distance of the possible host galaxy SDSS J140910.47+275320.8, we find a statistical preference for a binary neutron star merger scenario.

(transients:) gamma-ray bursts↗

Boosting Noise2Inverse via enhanced model selection for denoising computed tomography data

Synchrotron-based x-ray tomographic imaging enables the examination of the internal structure of materials at high spatial and temporal resolution. Experimental constraints can impose dose and time limits on the measurements, introducing a higher level of noise and artifacts in the reconstructed images. Deep learning has emerged as a powerful tool to remove noise from reconstructed images. Recently, the Noise2Inverse method was designed specifically for denoising reconstructed images without requiring paired noisy and clean images. This method creates multiple statistically independent reconstructions used to pair the data in which training involves transforming one reconstruction into the other, and vice versa. Originally designed to be used after a fixed number of epochs, we see in practice that this approach may not produce the optimal model and may unnecessarily waste computational resources. Therefore, we propose an alternative method of identifying the best model during training that aligns with the Noise2Inverse method. During validation, we compare the model output of the multiple reconstructions among each other. We hypothesize that the best model is the one that produces images with the highest similarity, implying a convergence in the predicted material properties and absorption values. To compare model outputs, we consider the absolute error, square error, structural similarity index (SSIM), peak signal-to-noise ratio (PSNR), and cosine similarity. We evaluate our method on two simulated tomography datasets and two, real-world, low-contrast, high-energy x-ray tomography datasets. We show our approach is more effective at determining the best model, up to an increase of 12.50% and 12.53% in SSIM and PSNR, respectively, while only requiring a fifth of the training time compared to the original approach.

CT↗