The variational method 2 - Perturbation theory
Variational method to approximate solutions of perturbation equations, perturbation analyses, and use of perturbation theory to infer variational principle
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Variational method to approximate solutions of perturbation equations, perturbation analyses, and use of perturbation theory to infer variational principle
Neural posterior estimation (NPE), a type of amortized variational inference, is a computationally efficient means of constructing probabilistic catalogs of light sources from astronomical images. To date, NPE has not been used to perform inference in models with spatially varying covariates. However, ground-based astronomical images exhibit spatially varying sky backgrounds and point spread functions (PSFs), and accounting for this variation is essential for constructing accurate catalogs of imaged light sources. In this work, we introduce a novel NPE-based cataloging method that trains an inference network with semisynthetic astronomical images generated using PSFs and backgrounds sampled from the Sloan Digital Sky Survey. In experiments with semisynthetic images, we evaluate the method on key cataloging tasks: light source detection, star/galaxy separation, and flux measurement. A “generalist” inference network—trained with diverse PSFs and backgrounds—performs as well as a “specialist” network even when both are evaluated on the specialist’s particular PSF/background combination. This result suggests that a single NPE network can generalize across spatial variations, eliminating the need for retraining on each observational condition.
The application of neural network models to scientific machine learning tasks has proliferated in recent years. In particular, neural networks have proved to be adept at modeling processes with spatial–temporal complexity. Nevertheless, these highly parameterized models have garnered skepticism in their ability to produce outputs with quantified error bounds over the regimes of interest. Hence there is a need to find uncertainty quantification methods that are suitable for neural networks. In this work we present comparisons of the parametric uncertainty quantification of neural networks modeling complex spatial–temporal processes with Hamiltonian Monte Carlo and Stein variational gradient descent and its projected variant. Specifically we apply these methods to graph convolutional neural network models of evolving systems modeled with recurrent neural network and neural ordinary differential equations architectures. We show that Stein variational inference is a viable alternative to Monte Carlo methods with some clear advantages for complex neural network models. For our exemplars, Stein variational interference gave similar pushed forward uncertainty profiles through time compared to Hamiltonian Monte Carlo, albeit with generally more generous variance. As a result, projected Stein variational gradient descent also produced similar uncertainty profiles to the non-projected counterpart, but large reductions in the active weight space were confounded by the stability of the neural network predictions and the convoluted likelihood landscape.
Bayesian Metabolic Inference is a platform to integrate genome-scale multi-omics data into an iterative method to obtain optimized recommendations for optimal target compound production. This method uses a variational inference approach to approximate new posterior data to analyze in order to obtain useful correlations and control coefficients.
In situ mass spectrometric measurements of ion and neutral particle thermospheric compositions have been used to infer the latitudinal and diurnal variations of thermospheric atomic hydrogen for solstice conditions. Local time-dependent and local time-independent components of the observed hydrogen distribution were separated on the basis of a model generated by expanding the log of neutral hydrogen concentration in terms of spherical harmonics. Results are compared with analogous data on N2 concentration; the comparison reveals an anticorrelation between gas temperature and hydrogen concentration. The slope of this anticorrelation line represents the 'zero flux condition' from the exosphere theory of Hodges (1973), thus further bearing out the conclusion that exospheric flow is the dominant process governing the global distribution of hydrogen.
We present GIGA-Lens: a gradient-informed, GPU-accelerated Bayesian framework for modeling strong gravitational lensing systems, implemented in TensorFlow and JAX. The three components, optimization using multistart gradient descent, posterior covariance estimation with variational inference, and sampling via Hamiltonian Monte Carlo, all take advantage of gradient information through automatic differentiation and massive parallelization on graphics processing units (GPUs). We test our pipeline on a large set of simulated systems and demonstrate in detail its high level of performance. The average time to model a single system on four Nvidia A100 GPUs is 105 s. The robustness, speed, and scalability offered by this framework make it possible to model the large number of strong lenses found in current surveys and present a very promising prospect for the modeling of ${ \mathcal O }({10}^{5})$ lensing systems expected to be discovered in the era of the Vera C. Rubin Observatory, Euclid, and the Nancy Grace Roman Space Telescope.
Hamiltonian Monte Carlo (HMC) is a powerful and accurate method to sample from the posterior distribution in Bayesian inference. However, HMC techniques are computationally demanding for Bayesian neural networks due to the high dimensionality of the network’s parameter space and the non-convexity of their posterior distributions. Therefore, various approximation techniques, such as variational inference (VI) or stochastic gradient MCMC, are often employed to infer the posterior distribution of the network parameters. Such approximations introduce inaccuracies in the inferred distributions, resulting in unreliable uncertainty estimates. In this work, we propose a hybrid approach that combines inexpensive VI and accurate HMC methods to efficiently and accurately quantify uncertainties in neural networks and neural operators. The proposed approach leverages an initial VI training on the full network. We examine the influence of individual parameters on the prediction uncertainty, which shows that a large proportion of the parameters do not contribute substantially to uncertainty in the network predictions. This information is then used to significantly reduce the dimension of the parameter space, and HMC is performed only for the subset of network parameters that strongly influence prediction uncertainties. This yields a framework for accelerating the full batch HMC for posterior inference in neural networks. We demonstrate the efficiency and accuracy of the proposed framework on deep neural networks and operator networks, showing that inference can be performed for large networks with tens to hundreds of thousands of parameters. Finally, we show that this method can effectively learn surrogates for complex physical systems by modeling the operator that maps from upstream conditions to wall-pressure data on a cone in hypersonic flow.
In scientific applications, predictive modeling is often of limited use without accurate uncertainty quantification (UQ) to indicate when a model may be extrapolating or when more data needs to be collected. Bayesian Neural Networks (BNNs) produce predictive uncertainty by propagating uncertainty in neural network (NN) weights and offer the promise of obtaining not only an accurate predictive model but also accurate UQ. However, in practice, obtaining accurate UQ with BNNs is difficult due in part to the approximations used for model training (such as those made in variational inference) and in part to the need to choose a suitable set of hyperparameters; these hyperparameters outnumber those needed for traditional NNs and often have opaque effects on the results. We aim to shed light on the effects of hyperparameter choices for variational BNNs by performing a global sensitivity analysis of variational BNN performance under varying hyperparameter settings. Our results indicate that many of the hyperparameters interact with each other to affect both predictive accuracy and UQ. For improved usage of variational BNNs in real-world applications, we suggest that thorough hyperparameter tuning, including tuning of prior hyperparameters and loss function parameters, is essential for accurate UQ in variational BNNs.
Here, we present the marginal unbiased score expansion (MUSE) method, an algorithm for generic high-dimensional hierarchical Bayesian inference. MUSE performs approximate marginalization over arbitrary non-Gaussian latent parameter spaces, yielding Gaussianized asymptotically unbiased and near-optimal constraints on global parameters of interest. It is computationally much cheaper than exact alternatives like Hamiltonian Monte Carlo (HMC), excelling on funnel problems which challenge HMC, and does not require any problem-specific user supervision like other approximate methods such as variational inference or many simulation-based inference methods. MUSE makes possible the first joint Bayesian estimation of the delensed Cosmic Microwave Background (CMB) power spectrum and gravitational lensing potential power spectrum, demonstrated here on a simulated data set as large as the upcoming South Pole Telescope 3G 1500 deg 2 survey, corresponding to a latent dimensionality of ~6 million and of order 100 global bandpower parameters. On a subset of the problem where an exact but more expensive HMC solution is feasible, we verify that MUSE yields nearly optimal results. We also demonstrate that existing spectrum-based forecasting tools which ignore pixel-masking underestimate predicted error bars by only ~10%. This method is a promising path forward for fast lensing and delensing analyses which will be necessary for future CMB experiments such as SPT-3G, Simons Observatory, or CMB-S4, and can complement or supersede existing HMC approaches. The success of MUSE on this challenging problem strengthens its case as a generic procedure for a broad class of high-dimensional inference problems.
Scientific machine learning (SciML) often operates in ill-conditioned, weakly identifiable regimes due to limited data or indirect observations. In such settings, optimization and inference are highly sensitive to the starting point, making initialization--often under-reported--a consequential degree of freedom. Random initialization is not a neutral default as it induces an implicit prior over candidate solutions and can systematically bias the result, producing large run-to-run variability. Here, we formalize this view by treating initialization as a hidden confounder in SciML and develop a unifying theory for structure-aware initialization via numerical continuation, constructing warm starts from related problem instances. Across representative tasks, including physics-informed neural networks, maximum likelihood estimation, and variational inference, warm starts have been shown to consistently reduce optimization effort and improve reliability.
This report describes a physics-based model for creep and thermal aging in Alloy 709. Alloy 709 is an advanced austenitic alloy, targeted for use in future Sodium Fast Reactors (SFRs) and other advanced reactors. The material has superior high temperature properties compared to currently qualified 316 and 304 stainless steels. However, the available creep and thermal aging test database for Alloy 709 is significantly more limited compared to the historical materials. The physics-based model developed here is one way to accelerate the qualification of the material by providing more accurate long-term predictions for creep properties and thermal aging, compared to current empirical time-extrapolate techniques. The crystal plasticity finite element model is used to predict the deformation and failure of alloy 709. The same setup for the CPFE model is used in both the baseline model calibration process and the simulation campaigns for parameter inference. Specific constitutive choices are made for Alloy 709 to capture the primary deformation mechanisms. The dislocation creep formulation developed by Hu and Cocks is extended to account for coupled precipitation formation and the grain boundary cavitation model developed by Sham, Needleman, et al. is used to model grain boundary cavitation-induced failure. A novel update algorithm is proposed to render the semi-discrete constitutive update for the Sham-Needleman model unconditionally stable. A progressive calibration approach is adopted based on the observations that several types of material responses can be effectively decoupled. A surrogate model is trained based on full-fledged CPFE simulations to accelerate the forward model evaluations, and stochastic variational inference (SVI) is used to calibrate the unknown microstructural model parameters. The calibrated mechanistic model is used to predict the long-term creep life of Alloy 709, and the predictions are compared against classical empirical approaches.
Estimation of riverbed profiles, also known as bathymetry, plays a vital role in many applications, such as safe and efficient inland navigation, prediction of bank erosion, land subsidence, and flood risk management. The high cost and complex logistics of direct bathymetry surveys, i.e, depth imaging, have encouraged the use of indirect measurements such as surface flow velocities. However, estimating high-resolution bathymetry from indirect measurements is an inverse problem that can be computationally challenging. Here, we propose a reduced-order model (ROM) based approach that utilizes a variational autoencoder (VAE), a type of deep neural network with a narrow layer in the middle, to compress bathymetry and flow velocity information and accelerate bathymetry inverse problems from flow velocity measurements. In our application, the shallow-water equations (SWE) with appropriate boundary conditions (BCs), e.g., the discharge and/or the free surface elevation, constitute the forward problem, to predict flow velocity. Then, ROMs of the SWEs are constructed on a nonlinear manifold of low dimensionality through a variational encoder and the bathymetry inversion problem is derived on the low-dimensional latent space in a Hierarchical Bayesian setting. Further, the reformulation allows variational inference with a small number (e.g., $\mathscr{O}$ (100) of ROM runs and efficient uncertainty quantification. We have tested our inversion approach on a one-mile reach of the Savannah River, GA, USA. Once the neural network is trained (offline stage), the proposed technique can perform the inversion operation orders of magnitude faster than traditional inversion methods that are commonly based on linear projections, such as principal component analysis (PCA), or the principal component geostatistical approach (PCGA). Furthermore, tests show that the algorithm can estimate the bathymetry with good accuracy even with sparse flow velocity measurements.
Here we propose a new class of Bayesian neural networks (BNNs) that can be trained using noisy data of variable fidelity, and we apply them to learn function approximations as well as to solve inverse problems based on partial differential equations (PDEs). These multi-fidelity BNNs consist of three neural networks: The first is a fully connected neural network, which is trained following the maximum a posteriori probability (MAP) method to fit the low-fidelity data; the second is a Bayesian neural network employed to capture the cross-correlation with uncertainty quantification between the low- and high-fidelity data; and the last one is the physics-informed neural network, which encodes the physical laws described by PDEs. For the training of the last two neural networks, we first employ the mean-field variational inference (VI) to maximize the evidence lower bound (ELBO) to obtain informative prior distributions for the hyperparameters in the BNNs, and subsequently we use the Hamiltonian Monte Carlo (HMC) method to estimate accurately the posterior distributions for the corresponding hyperparameters. We demonstrate the accuracy of the present method using synthetic data as well as real measurements. Specifically, we first approximate a one- and four-dimensional function, and then infer the reaction rates in one- and two-dimensional diffusion-reaction systems. Moreover, we infer the sea surface temperature (SST) in the Massachusetts and Cape Cod Bays using satellite images and in-situ measurements. Taken together, our results demonstrate that the present method can capture both linear and nonlinear correlation between the low- and high-fidelity data adaptively, identify unknown parameters in PDEs, and quantify uncertainties in predictions, given a few scattered noisy high-fidelity data. Finally, we demonstrate that we can effectively and efficiently reduce the uncertainties and hence enhance the prediction accuracy with an active learning approach, using as examples a specific one-dimensional function approximation and an inverse PDE problem.
Novel X-ray methods are transforming the study of the functional dynamics of biomolecules. Key to this revolution is detection of often subtle conformational changes from diffraction data. Diffraction data contain patterns of bright spots known as reflections. To compute the electron density of a molecule, the intensity of each reflection must be estimated, and redundant observations reduced to consensus intensities. Systematic effects, however, lead to the measurement of equivalent reflections on different scales, corrupting observation of changes in electron density. Here, we present a modern Bayesian solution to this problem, which uses deep learning and variational inference to simultaneously rescale and merge reflection observations. We successfully apply this method to monochromatic and polychromatic single-crystal diffraction data, as well as serial femtosecond crystallography data. We find that this approach is applicable to the analysis of many types of diffraction experiments, while accurately and sensitively detecting subtle dynamics and anomalous scattering.
With the increased use of data-driven approaches and machine learning-based methods in material science, the importance of reliable uncertainty quantification (UQ) of the predicted variables for informed decision-making cannot be overstated. UQ in material property prediction poses unique challenges, including multi-scale and multi-physics nature of materials, intricate interactions between numerous factors, limited availability of large curated datasets, etc. In this work, we introduce a physics-informed Bayesian Neural Networks (BNNs) approach for UQ, which integrates knowledge from governing laws in materials to guide the models toward physically consistent predictions. To evaluate the approach, we present case studies for predicting the creep rupture life of steel alloys. Experimental validation with three datasets of creep tests demonstrates that this method produces point predictions and uncertainty estimations that are competitive or exceed the performance of conventional UQ methods such as Gaussian Process Regression. Additionally, we evaluate the suitability of employing UQ in an active learning scenario and report competitive performance. The most promising framework for creep life prediction is BNNs based on Markov Chain Monte Carlo approximation of the posterior distribution of network parameters, as it provided more reliable results in comparison to BNNs based on variational inference approximation or related NNs with probabilistic outputs.
Galaxy clusters are one of the most powerful probes to study extensions of General Relativity and the Standard Cosmological Model. Upcoming surveys like the Vera Rubin Observatory’s Legacy Survey of Space and Time are expected to revolutionise the field, by enabling the analysis of cluster samples of unprecedented size and quality. To reach this era of high-precision cluster cosmology, the mitigation of sources of systematic error is crucial. A particularly important challenge is bias in cluster mass measurements induced by the mismodelling of photometric redshift estimates of source galaxies. This work proposes a method to optimise the source sample selection in cluster weak lensing analyses drawn from wide-field survey lensing catalogs to reduce the bias on reconstructed cluster masses. We use a combinatorial optimisation scheme and methods from variational inference to select galaxies in latent space to produce a probabilistic galaxy source sample catalog for highly accurate cluster mass estimation. We show that our method reduces the critical surface mass density Σ crit relative modelling bias on the 60-70% level, while maintaining up to 90% of galaxies. We highlight that our methodology has applications beyond cluster mass estimation as an approach to jointly combine galaxy selection and model inference under sources of systematics.
The abnormal events, such as the unprecedented COVID-19 pandemic, can significantly change the load behaviors, leading to huge challenges for traditional short-term forecasting methods. This article proposes a robust deep Gaussian processes (DGP)-based probabilistic load forecasting method using a limited number of data. Since the proposed method only requires a limited number of training samples for load forecasting, it allows us to deal with extreme scenarios that cause short-term load behavior changes. In particular, the load forecasting at the beginning of abnormal event is cast as a regression problem with limited training samples and solved by double stochastic variational inference DGP. The mobility data are also utilized to deal with the uncertainties and pattern changes and enhance the flexibility of the forecasting model. The proposed method can quantify the uncertainties of load forecasting outcomes, which would be essential under uncertain inputs. Extensive comparison results with other state-of-the-art point and probabilistic forecasting methods show that our proposed approach can achieve high forecasting accuracies with only a limited number of data while maintaining the excellent performance of capturing the forecasting uncertainties.
This work proposes a physics-informed sparse Gaussian process (SGP) for probabilistic stability assessment of large-scale power systems in the presence of uncertain dynamic PVs and loads. The differential and algebraic equations considering uncertainties from dynamic PVs and loads are reformulated to a nonlinear mapping relationship that allows the application of SGP. Thanks to the nonparametric characteristic of Gaussian process, the proposed framework does not require distributions of uncertain inputs and this distinguishes it from existing approaches. As the original Gaussian process is not scalable to large-scale systems with high dimensional uncertain inputs, this paper develops the SGP with a stochastic variational inference technique. It leads to approximately two orders of complex reduction. A data pre-processing step is also introduced to tackle the coexistence of stable and unstable cases by sample clustering and constructing separate SGPs. The probabilistic transient stability index is analyzed to assess system stability under different uncertain dynamics loads and PVs. Comparisons are performed with the sampling-based, the polynomial chaos expansion-based, and traditional Gaussian process-based methods on the modified IEEE 118-bus and Texas 2000-bus systems under various scenarios, including different levels of uncertainties and the existence of nonlinear correlations among dynamic PVs. The impacts of data quality and quantity issues are also investigated. It is shown that the proposed SGP achieves significantly improved computational efficiency while maintaining high accuracy with a limited number of data.