Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Bayesian model selection”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Footprints of the QCD Crossover on Cosmological Gravitational Waves at Pulsar Timing Arrays

Pulsar timing arrays (PTAs) have reported evidence for a stochastic gravitational wave (GW) background at nanohertz frequencies, possibly originating in the early Universe. We show that the spectral shape of the low-frequency (causality) tail of GW signals sourced at temperatures around T ≳ 1 GeV is distinctively affected by confinement of strong interactions (QCD), due to the corresponding sharp decrease in the number of relativistic species, and significantly deviates from ∼ f 3 commonly adopted in the literature. Bayesian analyses in the NANOGrav 15 years and the previous international PTA datasets reveal a significant improvement in the fit with respect to cubic power-law spectra, previously employed for the causality tail. While no conclusion on the nature of the signal can be drawn at the moment, our results show that the inclusion of standard model effects on cosmological GWs can have a decisive impact on model selection. Published by the American Physical Society 2024

Franciolini, Gabriele (ORCID:0000000268929145)↗

Measurement of $\nu_\mu$ CC Interactions With Two-Proton Final State in MINERvA

This dissertation presents a measurement of charged–current (CC) muon–neutrino interactions with exactly two protons and no pions in the final state (CC~$2p\,0\pi$), using data collected by the MINERvA detector in the NuMI medium–energy beam at Fermilab. Such two–proton topologies are a sensitive probe of nuclear dynamics in the few–GeV regime, including multi–nucleon correlations (npnh, notably $2p2h$) and intranuclear final–state interactions (FSI) such as pion absorption and nucleon rescattering. A precise experimental characterization of these processes is essential both for neutrino–interaction theory and for reducing systematic uncertainties in oscillation experiments that rely on accurate modeling of neutrino–nucleus interactions. Events are selected by requiring a $\nu_\mu$ CC interaction with a reconstructed $\mu^-$ and two proton tracks originating from a common vertex in MINERvA’s finely segmented scintillator tracker, with no reconstructed mesons. Muon charge and momentum are constrained by matching to the MINOS Near Detector, while proton identification exploits energy–loss profiles and stopping–proton features. Backgrounds from pion–producing channels that enter the signal region through FSI or reconstruction effects are constrained with data–driven sidebands (Michel–electron and isolated–cluster “blob” samples) and tuned via a simultaneous fit across signal and sideband regions. To correct detector resolution and acceptance effects, the analysis employs iterative Bayesian unfolding with extensive validation: statistical pseudo–experiments, and robustness checks against generator systematic “universes” and additional strong shape warps. Single–differential cross sections are reported for three observables tailored to the two–proton final state: the opening–angle cosine $\cos\!\left(\theta_{pp}\right)$, the leading–proton momentum, and the subleading–proton momentum. Systematic uncertainties include contributions from neutrino flux, interaction modeling (e.g., npnh and resonance parameters, pion FSI), and detector response (calibration, reconstruction efficiencies). The resulting distributions provide targeted constraints on the interplay of multi–nucleon dynamics and FSI that shape CC~$2p\,0\pi$ final states on hydrocarbon. Comparisons to modern GENIE–based simulations highlight kinematic regions where model components require refinement. These measurements thus inform generator tuning and improve the reliability of neutrino–energy reconstruction strategies for current and future long–baseline oscillation programs.

Syrotenko, Vladyslav S. [Tufts U.]↗

The eROSITA Final Equatorial-Depth Survey (eFEDS): The AGN catalog and its X-ray spectral properties

The eROSITA Final Equatorial Depth Survey (eFEDS), observed with eROSITA ahead of its planned 4-yr all-sky survey, is the largest contiguous-field X-ray survey at present. It yielded a large sample of X-ray sources with very rich multiband photometric and spectroscopic coverage. We present here the eFEDS active galactic nuclei (AGN) catalog and the eROSITA X-ray spectral properties of the eFEDS sources. Using a Bayesian method, we performed a systematic X-ray spectral analysis for all the eFEDS sources. We adopted multiple spectral models, including single-component power-law or hot-plasma models and double-component models of a power law plus soft excess. We investigated the capacity of eROSITA X-ray spectra for constraining AGN spectral shapes through a detailed analysis of the posterior parameter probability distribution functions. Hierarchical Bayesian modeling was used to recover the spectral parameter distribution of the sample. The source fluxes and luminosities were measured from the posterior of the spectral fitting. The eFEDS AGN catalog (22 079 sources) comprises ~80% of the eFEDS point sources. Despite a large number of faint sources, our spectral fitting provides reasonable measurements of spectral shapes and intrinsic luminosities for a majority of the sources. Because of sample selection bias, this AGN catalog is dominated by X-ray unobscured sources, with an obscured (logN H > 21.5) fraction of 8%; the power-law emission of the hot corona is also relatively soft, with a typical slope of 2.0. For type-I AGN, the X-ray emission is well correlated with the UV emission with the usual anticorrelation between the X-ray to UV spectral slope α OX and the UV luminosity. The X-ray spectral properties measured with various models are presented for all the eFEDS sources.

79 ASTRONOMY AND ASTROPHYSICS↗

BatteryPro: A Python Toolkit for Battery Data Analysis and Machine Learning Predictions

Analyzing battery test data for research & development can be time-consuming since battery tests often run on the order of months to years, generating large volumes of data. BatteryPro is a comprehensive Python package and software designed to facilitate advanced analysis and performance predictions for battery test data. Developed for battery researchers, it supports data types from widely used battery testing instruments, including MACCOR and Biologic cycling systems. The software provides a variety of tools for extracting and plotting key battery parameters such as time, voltage, capacity, current, and pressure. In addition to its extensive data analysis capabilities, BatteryPro features a dedicated machine learning module that employs a Bayesian Gaussian Mixture Model (GMM) to predict battery performance and degradation. Users can generate synthetic capacity fade data, calculate fade metrics, and leverage predictive models to forecast long-term battery behavior. The software's graphical user interface (GUI) enhances usability, allowing researchers to upload, merge, and analyze multiple data files with full customizability. The GUI also supports machine learning predictions, enabling users to fit models and make predictions based on selected data and parameters. BatteryPro is built using QtDesigner, scikit-learn, matplotlib, and pandas, ensuring a high level of customization, flexibility, and accuracy in battery data analysis. This tool aims to empower researchers with the ability to perform detailed battery analysis and make informed predictions, ultimately advancing the field of battery research.

25 - ENERGY STORAGE↗

Model form and sensitivity analysis of CALPHAD-based nucleation models in b-stabilized Ti alloys

Accurate prediction of α-phase nucleation and growth in β-stabilized titanium alloys is crucial for designing heat treatments to optimize mechanical properties in additively manufactured lightweight components. Ideally, predictions of nucleation and growth would incorporate both top-down observations of past experimental heat treatments and bottom-up modeling of phase transformations; however, the appropriate method of combining these information sources is not self-evident. Combining top-down and bottom-up information requires a unified form of model that can connect between spatiotemporal scales, as well as sets of fitting parameters that can be identified by each data source. The selection of which parameters to fit to which data source can be made based on expert opinion, or by performing a sensitivity analysis. In solid-solid nucleation, direct observation of the nucleation and growth process is challenging. Most data on the heat treatment-controlled phase transformations are not in-situ. To predict the process and outcome of the nucleation, growth and coarsening of precipitates, theoretical models of the nucleation pathway are used to bridge the gap. Many sources of uncertainty affect the modeling of this nucleation process. It can be influenced by small variations in the thermomechanical processing history, chemical composition, and initial microstructure. If molecular dynamics (MD) simulations are used to determine thermodynamic quantities and inform CALPHAD modeling, additional uncertainty can be introduced and accounted for using Bayesian methods. Top-down uncertainties require additional steps to quantify. The influence of nucleation model form on the sensitivity of predictions to input parameters and physical conditions is the focus of this study. Classical nucleation theory (CNT) allows modeling to formulate the nucleation as homogeneous or, more commonly, heterogeneous. Non-classical nucleation models are also increasingly explored as a means of reconciling top-down and bottom-up data. In this study, the sensitivity of the intragranular nucleation of α in a β-annealed, slow-cooled aging (BASCA) heat treatment of β-stabilized Ti5553 alloy is explored using CNT and both heterogeneous and homogeneous assumptions. The Kampmann-Wagner Numerical model of precipitate nucleation and growth is employed. Using open-source tools (pyCalphad and thermodynamic modeling of TiMo as a surrogate system, a sensitivity analysis is performed to measure variations in key parameters, including chemical driving force, interfacial energy, and diffusivity, as they relate to predictions of precipitate number density. The inclusion of top-down and bottom-up data in selection of nucleation model form is discussed.

Rodriguez Negron, A. M.↗

The dark matter halo masses of elliptical galaxies as a function of observationally robust quantities

Context. The assembly history of the stellar component of a massive elliptical galaxy is closely related to that of its dark matter halo. Measuring how the properties of galaxies correlate with their halo mass can therefore help to understand their evolution. Aims. We investigate how the dark matter halo mass of elliptical galaxies varies as a function of their properties, using weak gravitational lensing observations. To minimise the chances of biases, we focus on the following galaxy properties that can be determined robustly: the surface brightness profile and the colour. Methods. We selected 2409 central massive elliptical galaxies (log M*/M ⊙ ≳ 11.4) from the Sloan Digital Sky Survey spectroscopic sample. We first measured their surface brightness profile and colours by fitting Sérsic models to photometric data from the Kilo-Degree Survey (KiDS). We fitted their halo mass distribution as a function of redshift, rest-frame r-band luminosity, half-light radius, and rest-frame u - g colour, using KiDS weak lensing measurements and a Bayesian hierarchical approach. For the sake of robustness with respect to assumptions on the large-radii behaviour of the surface brightness, we repeated the analysis replacing the total luminosity and half-light radius with the luminosity within a 10 kpc aperture, L r, 10 , and the light-weighted surface brightness slope, Γ 10 . Results. We did not detect any correlation between the halo mass and either the half-light radius or colour at fixed redshift and luminosity. Using the robust surface brightness parameterisation, we found that the halo mass correlates weakly with L r,10 and anti-correlates with Γ 10 . At fixed redshift, L r, 10 and Γ 10 , the difference in the average halo mass between galaxies at the 84th percentile and 16th percentile of the colour distribution is 0.00 ± 0.11 dex. Conclusion. Our results indicate that the average star formation efficiency of massive elliptical galaxies has little dependence on their final size or colour. This suggests that the origin of the diversity in the size and colour distribution of these objects lies with properties other than the halo mass.

79 ASTRONOMY AND ASTROPHYSICS↗

Optimal sizing of battery energy storage systems for peak shaving and demand response using a degradation-aware Bayesian Optimization-Mixed-Integer Linear Programming framework

The increasing integration of renewable energy and rising electricity demand highlight the importance of battery energy storage systems for peak shaving and demand response. Unlike prior approaches that overlook operational impacts on degradation, this study proposes a Bayesian Optimization–Mixed Integer Linear Programming framework for optimal battery energy storage system sizing. In this framework, Mixed Integer Linear Programming determines short-term scheduling while a calibrated electrochemical model iteratively evaluates degradation. The central hypothesis is that the framework can efficiently identify optimal sizes that yield realistic and economically robust outcomes. The method is tested across three scenarios: peak shaving, peak shaving with energy-reduction demand response, and peak shaving with power-reduction demand response. Results show that the framework converge to the optimum within 20 iterations out of 150 possible sizes. Under baseline conditions, the framework consistently selects the smallest feasible system, minimizing unnecessary degradation costs from oversized storage. Sensitivity analyses reveal that larger systems are favored as demand rates or incentives increase. Comparisons of demand response programs indicate that power-reduction demand response offers greater economic benefits than energy-reduction demand response, although demand savings from peak shaving remain the dominant contributor to overall performance. This study demonstrates that the proposed framework balances computational tractability with degradation fidelity, identifies critical economic thresholds for investment, and offers a practical, flexible tool to guide industrial stakeholders in cost-effective battery energy storage system deployment.

Batteries↗

Noise-aware optimization in nominally identical manufacturing and measuring systems for high-throughput parallel workflows

Device-to-device variability in experimental noise critically impacts reproducibility, especially in automated, high-throughput systems like additive manufacturing farms. While manageable in small labs, such variability can escalate into serious risks at larger scales, such as architectural 3D printing, where noise may cause structural or economic failures. This contribution presents a noise-aware decision-making algorithm that quantifies and models device-specific noise profiles to manage variability adaptively. It uses distributional analysis and pairwise divergence metrics with clustering to choose between single-device and robust multi-device Bayesian optimization strategies. Unlike conventional methods that assume homogeneous devices or enforce generic robustness, the proposed framework explicitly determines whether shared optimization across devices is appropriate based on the degree of inter-device noise heterogeneity. This enables improved performance, reproducibility, and efficiency. An experimental case study involving three nominally identical 3D printers (same brand, model, and close serial numbers) demonstrates reduced redundancy, lower resource usage, and improved reliability, along with improved convergence stability and solution quality through the selection of the appropriate optimization strategy based on the degree of inter-device noise heterogeneity. Overall, this framework establishes a general approach for precision- and resource-aware optimization in scalable, automated experimental platforms, demonstrated here on a representative multi-device 3D printing case study.

Schenk, Christina↗

Hybrid Data‐Driven Discovery of High‐Performance Silver Selenide‐Based Thermoelectric Composites

Optimizing material compositions often enhances thermoelectric performances. However, the large selection of possible base elements and dopants results in a vast composition design space that is too large to systematically search using solely domain knowledge. To address this challenge, a hybrid data-driven strategy that integrates Bayesian optimization (BO) and Gaussian process regression (GPR) is proposed to optimize the composition of five elements (Ag, Se, S, Cu, and Te) in AgSe-based thermoelectric materials. Data is collected from the literature to provide prior knowledge for the initial GPR model, which is updated by actively collected experimental data during the iteration between BO and experiments. Within seven iterations, the optimized AgSe-based materials prepared using a simple high-throughput ink mixing and blade coating method deliver a high power factor of 2100 µW m −1 K −2 , which is a 75% improvement from the baseline composite (nominal composition of Ag 2 Se 1 ). In conclusion, the success of this study provides opportunities to generalize the demonstrated active machine learning technique to accelerate the development and optimization of a wide range of material systems with reduced experimental trials.

36 MATERIALS SCIENCE↗

Autoclass: An automatic classification system

The task of inferring a set of classes and class descriptions most likely to explain a given data set can be placed on a firm theoretical foundation using Bayesian statistics. Within this framework, and using various mathematical and algorithmic approximations, the AutoClass System searches for the most probable classifications, automatically choosing the number of classes and complexity of class descriptions. A simpler version of AutoClass has been applied to many large real data sets, has discovered new independently-verified phenomena, and has been released as a robust software package. Recent extensions allow attributes to be selectively correlated within particular classes, and allow classes to inherit, or share, model parameters through a class hierarchy. The mathematical foundations of AutoClass are summarized.

Stutz, John↗

Bayesian classification theory

The task of inferring a set of classes and class descriptions most likely to explain a given data set can be placed on a firm theoretical foundation using Bayesian statistics. Within this framework and using various mathematical and algorithmic approximations, the AutoClass system searches for the most probable classifications, automatically choosing the number of classes and complexity of class descriptions. A simpler version of AutoClass has been applied to many large real data sets, has discovered new independently-verified phenomena, and has been released as a robust software package. Recent extensions allow attributes to be selectively correlated within particular classes, and allow classes to inherit or share model parameters though a class hierarchy. We summarize the mathematical foundations of AutoClass.

Hanson, Robin↗

Assimilation of satellite microwave observations over the rainbands of tropical cyclones

A novel Bayesian Monte Carlo integration (BMCI) technique was developed to retrieve geophysical variables from satellite microwave radiometer data in the presence of tropical cyclones. The BMCI technique includes three steps: generating a stochastic database, simulating satellite brightness temperatures using a radiative transfer model, and retrieving geophysical variables such as profiles of temperature, relative humidity, and cloud liquid and ice water content from real observations. The technique also provides uncertainty estimates for each retrieval and can output the error covariance matrix of selected parameters. The measurements from the Advanced Technology Microwave Sounder (ATMS) on board Suomi National Polar-Orbiting Partnership (Suomi NPP) and the Global Precipitation Measurement (GPM) Microwave Imager (GMI) were used as input. A new technique was developed to correct the ATMS and GMI observations for the beam-filling effect, which is due to small-scale variability of precipitation and clouds when compared with the instrument footprint and also the nonlinear relation between the brightness temperature and precipitation. In addition, the assimilation of the BMCI retrievals into the NASA GEOS model is discussed for Hurricane Maria. The results show that assimilating the BMCI retrievals can influence the dynamical features of the cyclone, including a stronger warm core, a symmetric eye, and vertically aligned wind columns. Two possible factors that may limit the impact of the BMCI retrievals include 1) the resolution of the model (about 25 km), which was too coarse to show the potential of the BMCI data in improving the representation of tropical storms in the model forecast, and 2) the data assimilation system not being able to consider vertically correlated observation errors.

Isaac Moradi↗

Atacama Cosmology Telescope measurements of a large sample of candidates from the Massive and Distant Clusters of WISE Survey: Sunyaev-Zeldovich effect confirmation of MaDCoWS candidates using ACT

Context. Galaxy clusters are an important tool for cosmology, and their detection and characterization are key goals for current and future surveys. Using data from the Wide-field Infrared Survey Explorer (WISE), the Massive and Distant Clusters of WISE Survey (MaDCoWS) located 2839 significant galaxy overdensities at redshifts 0.7 . z . 1.5, which included extensive follow-up imaging from the Spitzer Space Telescope to determine cluster richnesses. Concurrently, the Atacama Cosmology Telescope (ACT) has produced large area millimeter-wave maps in three frequency bands along with a large catalog of Sunyaev-Zeldovich (SZ)-selected clusters as part of its Data Release 5 (DR5). Aims. We aim to verify and characterize MaDCoWS clusters using measurements of, or limits on, their thermal SZ effect signatures. We also use these detections to establish the scaling relation between SZ mass and the MaDCoWS-defined richness. Methods. Using the maps and cluster catalog from DR5, we explore the scaling between SZ mass and cluster richness. We do this by comparing cataloged detections and extracting individual and stacked SZ signals from the MaDCoWS cluster locations. We use complementary radio survey data from the Very Large Array, submillimeter data from Herschel, and ACT 224 GHz data to assess the impact of contaminating sources on the SZ signals from both ACT and MaDCoWS clusters. We use a hierarchical Bayesian model to fit the mass-richness scaling relation, allowing for clusters to be drawn from two populations: one, a Gaussian centered on the mass-richness relation, and the other, a Gaussian centered on zero SZ signal. Results. We find that MaDCoWS clusters have submillimeter contamination that is consistent with a gray-body spectrum, while the ACT clusters are consistent with no submillimeter emission on average. Additionally, the intrinsic radio intensities of ACT clusters are lower than those of MaDCoWS clusters, even when the ACT clusters are restricted to the same redshift range as the MaDCoWS clusters. We find the best-fit ACT SZ mass versus MaDCoWS richness scaling relation has a slope of p1 = 1.84+0.15 −0.14, where the slope is defined as M ∝ λ p1 15 and λ15 is the richness. We also find that the ACT SZ signals for a significant fraction (∼57%) of the MaDCoWS sample can statistically be described as being drawn from a noise-like distribution, indicating that the candidates are possibly dominated by low-mass and unvirialized systems that are below the mass limit of the ACT sample. Further, we note that a large portion of the optically confirmed ACT clusters located in the same volume of the sky as MaDCoWS are not selected by MaDCoWS, indicating that the MaDCoWS sample is not complete with respect to SZ selection. Finally, we find that the radio loud fraction of MaDCoWS clusters increases with richness, while we find no evidence that the submillimeter emission of the MaDCoWS clusters evolves with richness. Conclusions. We conclude that the original MaDCoWS selection function is not well defined and, as such, reiterate the MaDCoWS collaboration’s recommendation that the sample is suited for probing cluster and galaxy evolution, but not cosmological analyses. We find a best-fit mass-richness relation slope that agrees with the published MaDCoWS preliminary results. Additionally, we find that while the approximate level of infill of the ACT and MaDCoWS cluster SZ signals (1–2%) is subdominant to other sources of uncertainty for current generation experiments, characterizing and removing this bias will be critical for next-generation experiments hoping to constrain cluster masses at the sub-percent level.

large↗

Bayesian image reconstruction - The pixon and optimal image modeling

In this paper we describe the optimal image model, maximum residual likelihood method (OptMRL) for image reconstruction. OptMRL is a Bayesian image reconstruction technique for removing point-spread function blurring. OptMRL uses both a goodness-of-fit criterion (GOF) and an 'image prior', i.e., a function which quantifies the a priori probability of the image. Unlike standard maximum entropy methods, which typically reconstruct the image on the data pixel grid, OptMRL varies the image model in order to find the optimal functional basis with which to represent the image. We show how an optimal basis for image representation can be selected and in doing so, develop the concept of the 'pixon' which is a generalized image cell from which this basis is constructed. By allowing both the image and the image representation to be variable, the OptMRL method greatly increases the volume of solution space over which the image is optimized. Hence the likelihood of the final reconstructed image is greatly increased. For the goodness-of-fit criterion, OptMRL uses the maximum residual likelihood probability distribution introduced previously by Pina and Puetter (1992). This GOF probability distribution, which is based on the spatial autocorrelation of the residuals, has the advantage that it ensures spatially uncorrelated image reconstruction residuals.

Pina, R. K.↗

Optimizing Optical Searches for Supermassive Black Hole Binaries in Active Galactic Nuclei Light Curves: Fourier versus Bayesian Periodicity Detection

Simulations predict that supermassive black hole binaries (SMBHBs) will exhibit periodic brightness variations that may exceed the stochastic variability intrinsic to active galactic nuclei (AGN). In this paper, we simulate SMBHBs with damped random walk (DRW) AGN variability and an added sinusoidal signal from the orbital motion, and test three methods—a generalized Lomb–Scargle periodogram (GLSP), a nested Bayesian sampler (NBS), and a weighted wavelet z-transform (or WWZ)—to determine which is best at recovering the periodicity. Our simulated light curves follow the properties of the Catalina Real-Time Transient Survey (or CRTS), Legacy Survey of Space and Time (LSST), and Zwicky Transient Facility (ZTF) to best inform current and future SMBHB searches. We map a broad range of parameter space and identify which DRW-only light curves best mimic periodicity and pass each method’s model selection. The NBS performs best at detecting periodicity and filtering out DRW-only light curves. Combined candidate selection with both the NBS and GLSP significantly reduces false-positive rates (FPRs) with marginal impact on true-positive rates (TPRs). With this joint model selection pipeline, we find the lowest FPRs in ZTF-like simulations and the highest detection rates in LSST-like simulations. Using a modified computation of the false-alarm probability with GLSP, we efficiently triage LSST AGN light curves (∼10 7 light curves in ∼10–30 hr) and achieve TPRs and FPRs of ∼40% and ∼0.5%, respectively.

Banaszak, Sebastian M. [Vanderbilt Univ., Nashvill↗

Accelerated Discovery of CH 4 Uptake Capacity Metal–Organic Frameworks Using Bayesian Optimization

Abstract High‐throughput computational studies for discovery of metal–organic frameworks (MOFs) for separations and storage applications are often limited by the costs of computing thermodynamic quantities. Recent such studies at the time of writing may use ab initio results for a narrow selection of MOFs or empirical force‐field methods for larger selections. Here, a proof‐of‐concept study is conducted using Bayesian optimization on CH 4 uptake capacity of hypothetical MOFs for an existing dataset (Wilmer et al., Nature Chem. 2012, 4 , 83). It is shown that less than 0.1% of the database needs to be screened with the Bayesian optimization approach to recover the top candidate MOFs. This opens the possibility for efficient screening of MOF databases using accurate ab initio calculations for future adsorption studies on a minimal subset of MOFs. Furthermore, Bayesian optimization and the surrogate model presented here can offer interpretable material design insights and the framework will be applicable in the context of other target properties.

Taw, Eric↗

Optimal experimental design: Formulations and computations

Questions of ‘how best to acquire data’ are essential to modelling and prediction in the natural and social sciences, engineering applications, and beyond. Optimal experimental design (OED) formalizes these questions and creates computational methods to answer them. This article presents a systematic survey of modern OED, from its foundations in classical design theory to current research involving OED for complex models. We begin by reviewing criteria used to formulate an OED problem and thus to encode the goal of performing an experiment. We emphasize the flexibility of the Bayesian and decision-theoretic approach, which encompasses information-based criteria that are well-suited to nonlinear and non-Gaussian statistical models. We then discuss methods for estimating or bounding the values of these design criteria; this endeavour can be quite challenging due to strong nonlinearities, high parameter dimension, large per-sample costs, or settings where the model is implicit. A complementary set of computational issues involves optimization methods used to find a design; we discuss such methods in the discrete (combinatorial) setting of observation selection and in settings where an exact design can be continuously parametrized. Finally we present emerging methods for sequential OED that build non-myopic design policies, rather than explicit designs; these methods naturally adapt to the outcomes of past experiments in proposing new experiments, while seeking coordination among all experiments to be performed. Throughout, we highlight important open questions and challenges.

97 MATHEMATICS AND COMPUTING↗

Counterpart identification and classification for eRASS1 and characterisation of the active galactic nuclei content

Context. Accurately accounting for the Active Galactic Nucleus (AGN) phase in galaxy evolution requires a large, clean AGN sample. This is now possible with SRG/eROSITA, which completed its first all-sky X-ray survey (eRASS1) on June 12, 2020. The public Data Release 1 (DR1, Jan 31, 2024) includes 930,203 sources from the western Galactic hemisphere. Aims. The data enable the selection of a large AGN sample and the discovery of rare sources. However, scientific return depends on accurate characterisation of the X-ray emitters, requiring high-quality multi-wavelength data. This paper presents the identification and classification of optical and infrared counterparts to eRASS1 sources. Methods. Counterparts to eRASS1 X-ray point sources were identified using Gaia DR3, CatWISE2020, and Legacy Survey DR10 (LS10) with the Bayesian NWAY algorithm and trained priors. Sources were classified as Galactic or extragalactic via a machine-learning model combining optical/IR and X-ray properties, trained on a reference sample. For extragalactic LS10 sources, photometric redshifts were computed using CIRCLEZ. Results. Within the LS10 footprint, all 656,614 eROSITA/DR1 sources have at least one possible optical counterpart; ∼570 000 are extragalactic and likely AGN. Half are new detections compared to AllWISE, Gaia, and Quaia AGN catalogues. Gaia and CatWISE2020 counterparts are less reliable, due to the survey’s shallowness and the limited amount of features available to assess the probability of being an X-ray emitter. In the Galactic plane, where the overdensity of stellar sources also increases the chance of associations, using conservative reliability cuts, we identified approximately 18 000 Gaia and 55 000 CatWISE2020 extragalactic sources. Conclusions. We have released three high-quality counterpart catalogues – plus the training and validation sets – as a benchmark for the field. These datasets have many applications, but in particular, they empower researchers to build AGN samples tailored for completeness and purity, accelerating the hunt for the Universe’s most energetic engines.

X-rays: general↗