Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Bayesian model selection”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

A Gaussian Process Enhancement to Linear Parameter Varying Models

Simulation and analysis for modern engineering systems now routinely requires the merging of multiple disciplines, physical-domains, time-scales, and data sets — all at ever increasing levels. These capabilities are especially needed in the domain of Advanced Air Mobility, where rapidly emerging vehicle designs are significantly more complex, while having to be both cost-effective and safe. To meet these engineering challenges, machine learning methods are an attractive option for merging models and data across multiple areas while providing uncertainty quantification and maintaining computational efficiency. This paper examines the use of Gaussian process machine learning to generalize and enhance the commonly used class of quasi-Linear Parameter Varying models for fast full-envelope simulation while also supporting control system design and analysis with model uncertainty. Gaussian process machine learning is selected because it: can fuse multiple data sets, enables an easy trade-off between data fitting and smoothing, provides model uncertainty quantification, scales well with increasing complexity, and does not generally require starting from a large training data set. To demonstrate the benefits of the approach, a robust stability analysis with Gaussian process uncertainty is shown for a NASA reference design of an electric quad-rotor air-taxi concept vehicle with motor parameter uncertainty.

Gaussian Process↗

A Swing of the Pendulum: The Chemodynamics of the Local Stellar Halo Indicate Contributions from Several Radial Merger Events

We find that the chemical abundances and dynamics of APOGEE and GALAH stars in the local stellar halo are inconsistent with a scenario in which the inner halo is primarily composed of debris from a single massive, ancient merger event, as has been proposed to explain the Gaia-Enceladus/Gaia Sausage (GSE) structure. The data contain trends of chemical composition with energy that are opposite to expectations for a single massive, ancient merger event, and multiple chemical evolution paths with distinct dynamics are present. We use a Bayesian Gaussian mixture model regression algorithm to characterize the local stellar halo, and find that the data are fit best by a model with four components. We interpret these components as the Virgo Radial Merger (VRM), Cronus, Nereus, and Thamnos; however, Nereus and Thamnos likely represent more than one accretion event because the chemical abundance distributions of their member stars contain many peaks. Although the Cronus and Thamnos components have different dynamics, their chemical abundances suggest they may be related. We show that the distinct low- and high-α halo populations from Nissen & Schuster are explained by VRM and Cronus stars, as well as some in situ stars. Because the local stellar halo contains multiple substructures, different popular methods of selecting GSE stars will actually select different mixtures of these substructures, which may change the apparent chemodynamic properties of the selected stars. We also find that the Splash stars in the Solar region are shifted to higher v $\phi$ and slightly lower [Fe/H] than previously reported.

79 ASTRONOMY AND ASTROPHYSICS↗

Predicting Patterns of Solar Energy Buildout to Identify Opportunities for Biodiversity Conservation

The construction of solar energy facilities can have positive or negative impacts on biodiversity depending on siting and associated land use transitions. We identified drivers of solar siting and quantified patterns of buildout in states surrounding the Chesapeake Bay watershed – a biodiversity hotspot with numerous ecosystem services. Using a convolutional neural network, we mapped the footprints of ground-mounted solar arrays present in satellite imagery annually from 2017 to 2021 in Delaware, Maryland, Pennsylvania, New York, Virginia, and West Virginia. As of 2021, we identified 958 solar arrays covering 52.3 km2 built primarily on previously cultivated land, while avoiding natural landcover. We fit a binomial-Weibull model to these solar timeseries data in a hierarchical, Bayesian framework to quantify the relationship between geospatial covariates and rate of solar development. Solar array construction rate increased in cultivated areas, areas of lower agricultural suitability, lower slope, lower forest cover, lower biodiversity protection, and greater distances from roads. We also estimated changes in the rate of solar construction over time and found differences among states: acceleration in Virginia and deceleration in New York. We used parameter estimates to map the relative likelihood of future solar development across the study area. This methodology can be used to anticipate where solar is likely to be built in different landscapes and how these patterns align with conservation goals. Around the Chesapeake Bay watershed, the selection of lower quality agricultural areas for solar energy minimizes removal of important habitat and provides opportunities for native plant and pollinator restoration.

Artificial Intelligence↗

Countermeasures for Mitigation of Sensorimotor Decrements Following Head-Down Tilt Bed Rest

BACKGROUND Astronauts experience postflight disturbances in postural and locomotor control due to sensorimotor adaptations during spaceflight. These alterations may have adverse consequences if a rapid egress is required after landing. Although exercise is partially effective for mitigating cardiovascular and muscular deconditioning, additional countermeasures are needed to further preserve sensorimotor function. Proprioception training and electrical muscle stimulation (EMS) are two promising in-flight countermeasures. Since prolonged head down tilt bed rest (HDTBR) is a spaceflight analog for body unloading and causes postural and locomotor control decrements that parallel those observed after spaceflight, it can be used to facilitate the development of these countermeasures. METHODS This study will determine the effects of proprioception training and EMS on functional task performance and sensorimotor function following 60 days of 6° HDTBR. Subjects will be randomly assigned to one of four groups: 1) an EMS arm, 2) a proprioception training arm, 3) an exercise plus proprioceptive training arm, and 4) a control arm. The EMS countermeasure will include daily bilateral stimulation of selected bilateral lower extremity muscles (30 minutes per session). Proprioception training will be performed three days per week (25 minutes per session) consisting of body-loaded postural tasks in the horizontal position on an air bearing sled. Exercise training will mimic current protocols used on the International Space Station, but treadmill aerobic exercise will be replaced with additional cycling aerobic exercise. Primary outcome measures will include pre- and post- HDTBR functional tests that require high demand for dynamic control of postural stability. Secondary measures will be used to explore key physiological changes that underlie countermeasure benefits. All HDTBR and data collection activities will be completed by the German Aerospace Center (DLR) at the :envihab facility. Given the constrained samples size, a Bayesian modelling approach will be used to quantify the probability that there is an effect of a given magnitude. HARDWARE AND PROTOCOL DEVELOPMENT Final hardware modifications and protocol developments were completed in preparation for Campaign 1, which began in September 2024. These included shipment, setup, and operator training for the transportable gravity bed, horizontal squat device, foam obstacle course, EMS devices, leg dexterity system, foot sole skin sensitivity system, and Radiofrequency Echographic Multi Spectrometry (REMS) ultrasound device. In addition, specialized protocols were developed for data collection using DLR’s equipment, including muscle morphology magnetic resonance imaging (MRI), optical coherence tomography, venous blood flow MRI and ultrasound, muscle ultrasound and impedance, and skin blood flow ultrasound. We will present early data from the first campaign, which concluded in November, 2024. These will be compared with previous data from the recent 30-day HDTBR campaigns (Spaceflight associated neuro-ocular syndrome countermeasures (SANS-CM)) conducted at DLR. RELEVANCE The deliverable from this project will be proof-of-concept sensorimotor countermeasure designs for functional task performance with full assessment of efficacy in a spaceflight analog. If one or more countermeasures are effective, they will be translated for validation with the suite of operationally implemented in-flight countermeasures.

T R Macaulay↗

Securing Grid-interactive Efficient Buildings (GEB) through Cyber Defense and Resilient System (CYDRES)

The DOE CYDRES project is driven by the urgent need to address critical research gaps in the domain of cyber-physical security of smart buildings, including Grid-interactive Efficient Buildings (GEBs). CYDRES, a real-time advanced building resilient platform, aims to enhance the cyber-attack-immune capabilities of buildings through multi-layered prevention, detection, and adaptation mechanisms. CYDRES consists of five key modules: a multi-layer network analyzer, an Automatic Fault Detection, Diagnosis, and Prognosis (AFDDP) framework, an intelligent mode selector, a cyber-resilient control framework, and a situation awareness platform. The Network Analyzer employs a data-driven framework that includes a protocol state learning tool and a CRF (Conditional Random Field) command validator. In Hardware-In-the-Loop (HIL) testbeds, it achieved 100% detection accuracy with a false alarm rate of 3%, validating its efficacy in identifying selected cyber-attacks. The AFDDP framework leverages pattern matching, PCA (Principal Component Analysis)-based strategies, and a DBN (Dynamic Bayesian Network)-based fault diagnosis approach to pinpoint the causes of physical system abnormalities using Building Automation System (BAS) data. In HIL experiments, the AFDDP module attained a detection accuracy of over 95% with a false alarm rate below 7%. Additionally, the fault detector utilized machine learning (Random Forest) and deep learning (Multi-Layer Perceptron) methods with acoustic sensor data to achieve a 100% fault detection accuracy in Heating, Ventilation, and Air-Conditioning (HVAC) equipment. The Mode Selector offered real-time impact analysis, allowing immediate actions to protect BASs in the face of emerging threats. The cyber-resilient control framework included an adaptive Model Predictive Control (MPC) and a measurement compensator, reducing temperature violations by up to 94% and improving the total demand flexibility by up to 70% in HIL experiments. Such HIL experiments covered a cyber-attack case and a physical fault case, showcasing CYDRES’ efficiency in maintaining operational continuity during threats. The situation awareness platform in Grafana enhanced real-time threat detection and response visualization, augmenting the operational awareness for building operators. CYDRES demonstrated high technical effectiveness in various test scenarios, particularly in HIL environments. The project's phased development approach ensured efficient use of resources, highlighting its practical feasibility and readiness for commercialization. By enhancing the security and resilience of building operations, CYDRES represents a significant advance in mitigating risks associated with cyber-physical systems, thereby enhancing public confidence in the safety of modern building infrastructure. Future directions for the project include expanding testing protocols, refining AFDDP methodologies, exploring more comprehensive resilient control strategies, and testing in real commercial buildings.

42 ENGINEERING↗

Sensor Selection and Data Validation for Reliable Integrated System Health Management

For new access to space systems with challenging mission requirements, effective implementation of integrated system health management (ISHM) must be available early in the program to support the design of systems that are safe, reliable, highly autonomous. Early ISHM availability is also needed to promote design for affordable operations; increased knowledge of functional health provided by ISHM supports construction of more efficient operations infrastructure. Lack of early ISHM inclusion in the system design process could result in retrofitting health management systems to augment and expand operational and safety requirements; thereby increasing program cost and risk due to increased instrumentation and computational complexity. Having the right sensors generating the required data to perform condition assessment, such as fault detection and isolation, with a high degree of confidence is critical to reliable operation of ISHM. Also, the data being generated by the sensors needs to be qualified to ensure that the assessments made by the ISHM is not based on faulty data. NASA Glenn Research Center has been developing technologies for sensor selection and data validation as part of the FDDR (Fault Detection, Diagnosis, and Response) element of the Upper Stage project of the Ares 1 launch vehicle development. This presentation will provide an overview of the GRC approach to sensor selection and data quality validation and will present recent results from applications that are representative of the complexity of propulsion systems for access to space vehicles. A brief overview of the sensor selection and data quality validation approaches is provided below. The NASA GRC developed Systematic Sensor Selection Strategy (S4) is a model-based procedure for systematically and quantitatively selecting an optimal sensor suite to provide overall health assessment of a host system. S4 can be logically partitioned into three major subdivisions: the knowledge base, the down-select iteration, and the final selection analysis. The knowledge base required for productive use of S4 consists of system design information and heritage experience together with a focus on components with health implications. The sensor suite down-selection is an iterative process for identifying a group of sensors that provide good fault detection and isolation for targeted fault scenarios. In the final selection analysis, a statistical evaluation algorithm provides the final robustness test for each down-selected sensor suite. NASA GRC has developed an approach to sensor data qualification that applies empirical relationships, threshold detection techniques, and Bayesian belief theory to a network of sensors related by physics (i.e., analytical redundancy) in order to identify the failure of a given sensor within the network. This data quality validation approach extends the state-of-the-art, from red-lines and reasonableness checks that flag a sensor after it fails, to include analytical redundancy-based methods that can identify a sensor in the process of failing. The focus of this effort is on understanding the proper application of analytical redundancy-based data qualification methods for onboard use in monitoring Upper Stage sensors.

Garg, Sanjay↗

Quantum model learning agent: characterisation of quantum systems through machine learning

Accurate models of real quantum systems are important for investigating their behaviour, yet are difficult to distil empirically. Here, we report an algorithm—the quantum model learning agent (QMLA)—to reverse engineer Hamiltonian descriptions of a target system. We test the performance of QMLA on a number of simulated experiments, demonstrating several mechanisms for the design of candidate Hamiltonian models and simultaneously entertaining numerous hypotheses about the nature of the physical interactions governing the system under study. QMLA is shown to identify the true model in the majority of instances, when provided with limited a priori information, and control of the experimental setup. Our protocol can explore Ising, Heisenberg and Hubbard families of models in parallel, reliably identifying the family which best describes the system dynamics. We demonstrate QMLA operating on large model spaces by incorporating a genetic algorithm to formulate new hypothetical models. The selection of models whose features propagate to the next generation is based upon an objective function inspired by the Elo rating scheme, typically used to rate competitors in games such as chess and football. In all instances, our protocol finds models that exhibit F 1 score ≥ 0.88 when compared with the true model, and it precisely identifies the true model in 72% of cases, whilst exploring a space of over 250 000 potential models. By testing which interactions actually occur in the target system, QMLA is a viable tool for both the exploration of fundamental physics and the characterisation and calibration of quantum devices.

97 MATHEMATICS AND COMPUTING↗

Constraining the Milky Way Mass Profile with Phase-space Distribution of Satellite Galaxies

We estimate the Milky Way (MW) halo properties using satellite kinematic data including the latest measurements from Gaia DR2. With a simulation-based 6D phase-space distribution function (DF) of satellite kinematics, we can infer halo properties efficiently and without bias, and handle the selection function and measurement errors rigorously in the Bayesian framework. Applying our DF from the EAGLE simulation to 28 satellites, we obtain an MW halo mass of $M={1.23}_{-0.18}^{+0.21}\times {10}^{12}{M}_{\odot }$ and a concentration of $c={9.4}_{-2.1}^{+2.8}$ with the prior based on the M–c relation. The inferred mass profile is consistent with previous measurements but with better precision and reliability due to the improved methodology and data. Potential improvement is illustrated by combining satellite data and stellar rotation curves. Using our EAGLE DF and best-fit MW potential, we provide much more precise estimates of the kinematics for those satellites with uncertain measurements. Compared to the EAGLE DF, which matches the observed satellite kinematics very well, the DF from the semi-analytical model based on the dark-matter-only simulation Millennium II (SAM-MII) over-represents satellites with small radii and velocities. We attribute this difference to less disruption of satellites with small pericenter distances in the SAM-MII simulation. Finally, by varying the disruption rate of such satellites in this simulation, we estimate a ~5% scatter in the inferred MW halo mass among hydrodynamics-based simulations.

79 ASTRONOMY AND ASTROPHYSICS↗

Cholesky-based experimental design for Gaussian process and kernel-based emulation and calibration.

Gaussian processes and other kernel-based methods are used extensively to construct approximations of multivariate data sets. The accuracy of these approximations is dependent on the data used. This paper presents a computationally efficient algorithm to greedily select training samples that minimize the weighted L p error of kernel-based approximations for a given number of data. The method successively generates nested samples, with the goal of minimizing the error in high probability regions of densities specified by users. The algorithm presented is extremely simple and can be implemented using existing pivoted Cholesky factorization methods. Training samples are generated in batches which allows training data to be evaluated (labeled) in parallel. For smooth kernels, the algorithm performs comparably with the greedy integrated variance design but has significantly lower complexity. Numerical experiments demonstrate the efficacy of the approach for bounded, unbounded, multi-modal and non-tensor product densities. We also show how to use the proposed algorithm to efficiently generate surrogates for inferring unknown model parameters from data using Bayesian inference.

97 MATHEMATICS AND COMPUTING↗

Assimilating microwave cloudy observations into NASA GEOS model using a novel Bayesian Monte Carlo technique

Despite the importance of clouds and their influence on atmospheric water and energy balance, Numerical Weather Prediction (NWP) centers systematically exclude cloud information from the assimilation process and only assimilate clear-sky radiances (Janiskov´a et al. 2012). In order to ensure that only clear sky radiances are assimilated, strict cloud detection thresholds are applied before radiances are fed into data assimilation (DA) systems. This process not only excludes a large portion of satellite radiances, but causes loss of information in the regions that are of high interest to meteorologists and are most challenging for weather forecasts (Errico et al. 2007; Haddad et al. 2015). Although, in recent years there has been great advances in the operational weather forecasting, the prediction of tropical cyclones (TC), especially the intensity of TCs, remains challenging. According to Aksoy et al. (2013), in addition to the model deficiencies, another important factor that contributes to this challenge includes lack of observations in the peripheral environment (rain- bands) of TCs mainly because of the selective assimilation of existing observations. Satellite observations provide more than 90 % of the input data for the initialization of NWP models but more than 75 % of satellite observations are discarded due to the cloud contamination as well as land, snow, and ice emissivity issues (Bauer et al. 2010).

Rainband↗

Constraining neutrino oscillation and interaction parameters with the NOvA Near Detector and Far Detector data using Markov Chain Monte Carlo

This thesis reports a constraint of the neutrino oscillation parameters $\Delta m^{2}_{32}$, $\sin^2 \theta_{23}$, and $\delta_{CP}$ using the NuMI Off-Axis $\nu$ Appearance (NOvA) experiment's Near Detector (ND) data and Far Detector (FD) fake data set simultaneously. This thesis also reports a constraint on NOvA's systematic uncertainty model solely with its Near Detector data. The Hamiltonian Monte Carlo algorithm is used to estimate Bayesian Credible Intervals for the oscillation and interaction parameters. The $1\sigma$ Credible Intervals for $\sin^2 \theta_{23}$ are $(0.44, 0.512)$ $\cup$ $(0.536, 0.56)$, for $\Delta m^{2}_{32}$ $(2.41 \times 10^{-3}$ eV$^2,\ 2.52 \times 10^{-3}$ eV$^2)$, and for $\delta_{CP}$ $(0.74\pi,\ 1.1\pi)$ $\cup$ $(1.38\pi,\ 1.58\pi)$. The statistical power of the ND data constrains NOvA's interaction parameters, while the FD fake data constrains the oscillation parameters. This is the first analysis within NOvA to constrain the ND and FD prediction sim ultaneously, and to investigate the neutrino interaction modeling in the context of constraining the oscillation parameters. To constrain the ND data requires a sophisticated understanding of the neutrino interaction modeling and its uncertainties. The interested reader is advised to focus on Chapters 4 and 6, which discuss the ND selection, uncertainties, and ND-only fits to data. The reader interested in oscillation parameter constraints will find this in Chapter 7.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Photometric redshift uncertainties in weak gravitational lensing shear analysis: models and marginalization

ABSTRACT Recovering credible cosmological parameter constraints in a weak lensing shear analysis requires an accurate model that can be used to marginalize over nuisance parameters describing potential sources of systematic uncertainty, such as the uncertainties on the sample redshift distribution n(z). Due to the challenge of running Markov chain Monte Carlo (MCMC) in the high-dimensional parameter spaces in which the n(z) uncertainties may be parametrized, it is common practice to simplify the n(z) parametrization or combine MCMC chains that each have a fixed n(z) resampled from the n(z) uncertainties. In this work, we propose a statistically principled Bayesian resampling approach for marginalizing over the n(z) uncertainty using multiple MCMC chains. We self-consistently compare the new method to existing ones from the literature in the context of a forecasted cosmic shear analysis for the HSC three-year shape catalogue, and find that these methods recover statistically consistent error bars for the cosmological parameter constraints for predicted HSC three-year analysis, implying that using the most computationally efficient of the approaches is appropriate. However, we find that for data sets with the constraining power of the full HSC survey data set (and, by implication, those upcoming surveys with even tighter constraints), the choice of method for marginalizing over n(z) uncertainty among the several methods from the literature may modify the 1σ uncertainties on Ωm–S8 constraints by ∼4 per cent, and a careful model selection is needed to ensure credible parameter intervals.

Zhang, Tianqing (ORCID:000000025596198X)↗

Rao-Blackwellization for Adaptive Gaussian Sum Nonlinear Model Propagation

When dealing with imperfect data and general models of dynamic systems, the best estimate is always sought in the presence of uncertainty or unknown parameters. In many cases, as the first attempt, the Extended Kalman filter (EKF) provides sufficient solutions to handling issues arising from nonlinear and non-Gaussian estimation problems. But these issues may lead unacceptable performance and even divergence. In order to accurately capture the nonlinearities of most real-world dynamic systems, advanced filtering methods have been created to reduce filter divergence while enhancing performance. Approaches, such as Gaussian sum filtering, grid based Bayesian methods and particle filters are well-known examples of advanced methods used to represent and recursively reproduce an approximation to the state probability density function (pdf). Some of these filtering methods were conceptually developed years before their widespread uses were realized. Advanced nonlinear filtering methods currently benefit from the computing advancements in computational speeds, memory, and parallel processing. Grid based methods, multiple-model approaches and Gaussian sum filtering are numerical solutions that take advantage of different state coordinates or multiple-model methods that reduced the amount of approximations used. Choosing an efficient grid is very difficult for multi-dimensional state spaces, and oftentimes expensive computations must be done at each point. For the original Gaussian sum filter, a weighted sum of Gaussian density functions approximates the pdf but suffers at the update step for the individual component weight selections. In order to improve upon the original Gaussian sum filter, Ref. [2] introduces a weight update approach at the filter propagation stage instead of the measurement update stage. This weight update is performed by minimizing the integral square difference between the true forecast pdf and its Gaussian sum approximation. By adaptively updating each component weight during the nonlinear propagation stage an approximation of the true pdf can be successfully reconstructed. Particle filtering (PF) methods have gained popularity recently for solving nonlinear estimation problems due to their straightforward approach and the processing capabilities mentioned above. The basic concept behind PF is to represent any pdf as a set of random samples. As the number of samples increases, they will theoretically converge to the exact, equivalent representation of the desired pdf. When the estimated qth moment is needed, the samples are used for its construction allowing further analysis of the pdf characteristics. However, filter performance deteriorates as the dimension of the state vector increases. To overcome this problem Ref. [5] applies a marginalization technique for PF methods, decreasing complexity of the system to one linear and another nonlinear state estimation problem. The marginalization theory was originally developed by Rao and Blackwell independently. According to Ref. [6] it improves any given estimator under every convex loss function. The improvement comes from calculating a conditional expected value, often involving integrating out a supportive statistic. In other words, Rao-Blackwellization allows for smaller but separate computations to be carried out while reaching the main objective of the estimator. In the case of improving an estimator's variance, any supporting statistic can be removed and its variance determined. Next, any other information that dependents on the supporting statistic is found along with its respective variance. A new approach is developed here by utilizing the strengths of the adaptive Gaussian sum propagation in Ref. [2] and a marginalization approach used for PF methods found in Ref. [7]. In the following sections a modified filtering approach is presented based on a special state-space model within nonlinear systems to reduce the dimensionality of the optimization problem in Ref. [2]. First, the adaptive Gaussian sum propagation is explained and then the new marginalized adaptive Gaussian sum propagation is derived. Finally, an example simulation is presented.

state estimation↗

Closed-loop optimization of fast-charging protocols for batteries with machine learning

Simultaneously optimizing many design parameters in time-consuming experiments causes bottlenecks in a broad range of scientific and engineering disciplines. One such example is process and control optimization for lithium-ion batteries during materials selection, cell manufacturing and operation. A typical objective is to maximize battery lifetime; however, conducting even a single experiment to evaluate lifetime can take months to years. Furthermore, both large parameter spaces and high sampling variability necessitate a large number of experiments. As such, the key challenge is to reduce both the number and the duration of the experiments required. Here we develop and demonstrate a machine learning methodology to efficiently optimize a parameter space specifying the current and voltage profiles of six-step, ten-minute fast-charging protocols for maximizing battery cycle life, which can alleviate range anxiety for electric-vehicle users. We combine two key elements to reduce the optimization cost: an early-prediction model, which reduces the time per experiment by predicting the final cycle life using data from the first few cycles, and a Bayesian optimization algorithm, which reduces the number of experiments by balancing exploration and exploitation to efficiently probe the parameter space of charging protocols. Using this methodology, we rapidly identify high-cycle-life charging protocols among 224 candidates in 16 days (compared with over 500 days using exhaustive search without early prediction), and subsequently validate the accuracy and efficiency of our optimization approach. Our closed-loop methodology automatically incorporates feedback from past experiments to inform future decisions and can be generalized to other applications in battery design and, more broadly, other scientific domains that involve time-intensive experiments and multi-dimensional design spaces.

25 ENERGY STORAGE↗

Adaptive hyperparameter updating for training restricted Boltzmann machines on quantum annealers

Restricted Boltzmann Machines (RBMs) have been proposed for developing neural networks for a variety of unsupervised machine learning applications such as image recognition, drug discovery, and materials design. The Boltzmann probability distribution is used as a model to identify network parameters by optimizing the likelihood of predicting an output given hidden states trained on available data. Training such networks often requires sampling over a large probability space that must be approximated during gradient based optimization. Quantum annealing has been proposed as a means to search this space more efficiently which has been experimentally investigated on D-Wave hardware. D-Wave implementation requires selection of an effective inverse temperature or hyperparameter (β) within the Boltzmann distribution which can strongly influence optimization. Here, we show how this parameter can be estimated as a hyperparameter applied to D-Wave hardware during neural network training by maximizing the likelihood or minimizing the Shannon entropy. We find both methods improve training RBMs based upon D-Wave hardware experimental validation on an image recognition problem. Neural network image reconstruction errors are evaluated using Bayesian uncertainty analysis which illustrate more than an order magnitude lower image reconstruction error using the maximum likelihood over manually optimizing the hyperparameter. The maximum likelihood method is also shown to out-perform minimizing the Shannon entropy for image reconstruction.

97 MATHEMATICS AND COMPUTING↗

Constraints on neutrino physics from DESI DR2 BAO and DR1 full shape

The Dark Energy Spectroscopic Instrument (DESI) Collaboration has obtained robust measurements of baryon acoustic oscillations in the redshift range 0.1 < 𝑧 < 4.2, based on the Lyman-𝛼 forest and galaxies from data release 2. We combine these measurements with cosmic microwave background (CMB) data from Planck and the Atacama Cosmology Telescope to place our tightest constraints yet on the sum of neutrino masses. Assuming the cosmological Λ⁢ CDM model and three degenerate neutrino states, we find ∑𝑚 𝜈 < 0.0642 eV (95%) with a marginalized error of 𝜎⁡(∑𝑚 𝜈 ) = 0.020 eV. We also constrain the effective number of neutrino species, finding 𝑁 eff = 3.2⁢3$^{+0.35}_{−0.34}$ (95%), in line with the Standard Model prediction. When accounting for neutrino oscillation constraints, we find a preference for the normal mass ordering and an upper limit on the lightest neutrino mass of 𝑚 𝑙 < 0.023 eV (95%). However, we determine using frequentist and Bayesian methods that our constraints are in tension with the lower limits derived from neutrino oscillations. Correcting for the physical boundary at zero mass, we report a 95% Feldman-Cousins upper limit of ∑𝑚 𝜈 < 0.053 eV, breaching the lower limit from neutrino oscillations. Considering a more general Bayesian analysis with an effective cosmological neutrino mass parameter, ∑𝑚 𝜈,eff , that allows for negative energy densities and removes unsatisfactory prior weight effects, we derive constraints that are in 3⁢𝜎 tension with the same oscillation limit, while the error rises to 𝜎⁡(∑𝑚 𝜈,eff ) = 0.053 eV. In the absence of unknown systematics, this finding could be interpreted as a hint of new physics not necessarily related to neutrinos. The preference of DESI and CMB data for an evolving dark energy model offers one possible solution. In the 𝑤 0 ⁢𝑤 𝑎 ⁢CDM model, we find ∑𝑚 𝜈 < 0.163 eV (95%), relaxing the neutrino tension. These constraints all rely on the effects of neutrinos on the cosmic expansion history. Using full-shape power spectrum measurements of data release 1 galaxies, we place complementary constraints that rely on neutrino free streaming. Our strongest such limit in Λ ⁢CDM, using selected CMB priors, is ∑𝑚 𝜈 < 0.193 eV (95%).

79 ASTRONOMY AND ASTROPHYSICS↗

Beyond maximum entropy: Fractal Pixon-based image reconstruction

We have developed a new Bayesian image reconstruction method that has been shown to be superior to the best implementations of other competing methods, including Goodness-of-Fit methods such as Least-Squares fitting and Lucy-Richardson reconstruction, as well as Maximum Entropy (ME) methods such as those embodied in the MEMSYS algorithms. Our new method is based on the concept of the pixon, the fundamental, indivisible unit of picture information. Use of the pixon concept provides an improved image model, resulting in an image prior which is superior to that of standard ME. Our past work has shown how uniform information content pixons can be used to develop a 'Super-ME' method in which entropy is maximized exactly. Recently, however, we have developed a superior pixon basis for the image, the Fractal Pixon Basis (FPB). Unlike the Uniform Pixon Basis (UPB) of our 'Super-ME' method, the FPB basis is selected by employing fractal dimensional concepts to assess the inherent structure in the image. The Fractal Pixon Basis results in the best image reconstructions to date, superior to both UPB and the best ME reconstructions. In this paper, we review the theory of the UPB and FPB pixon and apply our methodology to the reconstruction of far-infrared imaging of the galaxy M51. The results of our reconstruction are compared to published reconstructions of the same data using the Lucy-Richardson algorithm, the Maximum Correlation Method developed at IPAC, and the MEMSYS ME algorithms. The results show that our reconstructed image has a spatial resolution a factor of two better than best previous methods (and a factor of 20 finer than the width of the point response function), and detects sources two orders of magnitude fainter than other methods.

Puetter, Richard C.↗