Probability maximization via Minkowski functionals: convex representations and tractable resolution
Not provided.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Not provided.
We analyze the convergence of a nonlocal gradient descent method for minimizing a class of high-dimensional non-convex functions, where a directional Gaussian smoothing (DGS) is proposed to define the nonlocal gradient (also referred to as the DGS gradient). The method was first proposed in [Zhang et al., Enabling long-range exploration in minimization of multimodal functions, UAI 2021], in which multiple numerical experiments showed that replacing the traditional local gradient with the DGS gradient can help the optimizers escape local minima more easily and significantly improve their performance. However, a rigorous theory for the efficiency of the method on nonconvex landscape is lacking. In this work, we investigate the scenario where the objective function is composed of a convex function, perturbed by deterministic oscillating noise. We provide a convergence theory under which the iterates exponentially converge to a tightened neighborhood of the solution, whose size is characterized by the noise wavelength. Here, we also establish a correlation between the optimal values of the Gaussian smoothing radius and the noise wavelength, thus justifying the advantage of using moderate or large smoothing radii with the method. Furthermore, if the noise level decays to zero when approaching the global minimum, we prove that DGS-based optimization converges to the exact global minimum with linear rates, similarly to standard gradient-based methods in optimizing convex functions. Several numerical experiments are provided to confirm our theory and illustrate the superiority of the approach over those based on the local gradient.
Many applications require minimizing the sum of smooth and nonsmooth functions. For example, basis pursuit denoising problems in data science require minimizing a measure of data misfit plus an $\ell^1$-regularizer. Similar problems arise in the optimal control of partial differential equations (PDEs) when sparsity of the control is desired. Here, we develop a novel trust-region method to minimize the sum of a smooth nonconvex function and a nonsmooth convex function. Our method is unique in that it permits and systematically controls the use of inexact objective function and derivative evaluations. When using a quadratic Taylor model for the trust-region subproblem, our algorithm is an inexact, matrix-free proximal Newton-type method that permits indefinite Hessians. We prove global convergence of our method in Hilbert space and demonstrate its efficacy on three examples from data science and PDE-constrained optimization.
Dynamic optimization problems arise in many applications including flow control, full waveform inversion, and medical imaging. These problems are plagued by significant computational challenges. One such challenge — and the focus of this work — is the memory limitation induced by the size of the underlying dynamical system. In particular, the entire dynamic trajectory is required for derivative computation and therefore must be stored or recomputed using, e.g., checkpointing. Although recent work demonstrated the use of adaptive randomized sketching to overcome the memory challenge, that work only applies to smooth unconstrained problems, prohibiting its use for nonsmooth regularized and constrained problems. The inclusion of nonsmooth regularizers and constraints is critical as they often arise in an attempt to preserve certain physical properties or to promote sparsity. To solve these problems, we introduce a trust-region algorithm for minimizing the sum of a smooth nonconvex function and a nonsmooth convex function that leverages randomized sketching to compress the dynamical system trajectories and adaptively adjust the sketch rank to satisfy a gradient inexactness condition. We prove convergence of this algorithm and demonstrate that it achieves substantial memory reduction on three discretized PDE-constrained optimization applications.
In Baraldi, we introduced an inexact trust-region algorithm for minimizing the sum of a smooth nonconvex function and a nonsmooth convex function in Hilbert space—a class of problems that is ubiquitous in data science, learning, optimal control, and inverse problems. Furthermore, this algorithm has demonstrated excellent performance and scalability with problem size. In this paper, we enrich the convergence analysis for this algorithm, proving strong convergence of the iterates with guaranteed rates. In particular, we demonstrate that the trust-region algorithm recovers superlinear, even quadratic, convergence rates when using a second-order Taylor approximation of the smooth objective function term.
Total variation (TV) is a widely used function for regularizing imaging inverse problems that is particularly appropriate for images whose underlying structure is piecewise constant. TV regularized optimization problems are typically solved using proximal methods, but the way in which they are applied is constrained by the absence of a closed-form expression for the proximal operator of the TV function. A closed-form approximation of the TV proximal operator has previously been proposed, but its accuracy was not theoretically explored in detail. Here, we address this gap by making several new theoretical contributions, proving that the approximation leads to a proximal operator of some convex function, it is equivalent to a gradient descent step on a smoothed version of TV, and that its error can be fully characterized and controlled with its scaling parameter. We experimentally validate our theoretical results on image denoising and sparse-view computed tomography (CT) image reconstruction.
Input convex neural networks (ICNNs) are a family of deep learning models where the outputs are constructed to be convex functions of the inputs. By parameterizing system models using ICNNs, optimization-based control problems can be solved as convex optimization problems, leading to improved performance and robustness. This work proposes a novel framework where the control objective function and constraints are modelled using ICNNs. A case study of optimization-based control with output constraints is conducted on a process with input multiplicity and nonminimum phase behavior. The simulation results demonstrate improved economic yield compared with normal neural networks. Additionally, the input convexity formulation is compared with simple regularization techniques, and unique benefits such as improved data efficiency and robustness of the proposed formulation are shown. Finally, by explicitly incorporating prior knowledge about convexity, this framework provides a good balance between the universal approximation power of deep learning and computational feasibility required by control.
We consider the Seiberg-Witten solution of pure N = 2 gauge theory in four dimensions, with gauge group SU(N). A simple exact series expansion for the dependence of the 2(N – 1) Seiberg-Witten periods a I (u), a DI (u) on the N – 1 Coulomb-branch moduli un is obtained around the Z 2N -symmetric point of the Coulomb branch, where all u n vanish. This generalizes earlier results for N = 2 in terms of hypergeometric functions, and for N = 3 in terms of Appell functions. Using these and other analytical results, combined with numerical computations, we explore the global structure of the Kähler potential K = 1/2Σ I Im(a¯ I a DI ), which is single valued on the Coulomb branch. Evidence is presented that K is a convex function, with a unique minimum at the Z 2N -symmetric point. Finally, we explore candidate walls of marginal stability in the vicinity of this point, and their relation to the surface of vanishing Kähler potential.
In [R. J. Baraldi and D. P. Kouri, Mathematical Programming, (2022), pp. 1-40], we introduced an inexact trust-region algorithm for minimizing the sum of a smooth nonconvex and nonsmooth convex function. The principle expense of this method is in computing a trial iterate that satisfies the so-called fraction of Cauchy decrease condition—a bound that ensures the trial iterate produces sufficient decrease of the subproblem model. In this paper, we expound on various proximal trust-region subproblem solvers that generalize traditional trust-region methods for smooth unconstrained and convex-constrained problems. We introduce a simplified spectral proximal gradient solver, a truncated nonlinear conjugate gradient solver, and a dogleg method. Finally, we compare algorithm performance on examples from data science and PDE-constrained optimization.
Here we study the problem of minimizing a convex function on a nonempty, finite subset of the integer lattice when the function cannot be evaluated at noninteger points. We propose a new underestimator that does not require access to (sub)gradients of the objective; such information is unavailable when the objective is a blackbox function. Rather, our underestimator uses secant linear functions that interpolate the objective function at previously evaluated points. These linear mappings are shown to underestimate the objective in disconnected portions of the domain. Therefore, the union of these conditional cuts provides a nonconvex underestimator of the objective. We propose an algorithm that alternates between updating the underestimator and evaluating the objective function. We prove that the algorithm converges to a global minimum of the objective function on the feasible set. We present two approaches for representing the underestimator and compare their computational effectiveness. We also compare implementations of our algorithm with existing methods for minimizing functions on a subset of the integer lattice. We discuss the difficulty of this problem class and provide insights into why a computational proof of optimality is challenging even for moderate problem sizes.
We consider minimizing a sum of agent-specific nondifferentiable merely convex functions over the solution set of a variational inequality (VI) problem in that each agent is associated with a local monotone mapping. This problem finds an application in computation of the best equilibrium in nonlinear complementarity problems arising in transportation networks. We develop an iteratively regularized incremental gradient method where at each iteration, agents communicate over a directed cycle graph to update their solution iterates using their local information about the objective and the mapping. The proposed method is single-timescale in the sense that it does not involve any excessive hard-to-project computation per iteration. We derive nonasymptotic agent-wise convergence rates for the suboptimality of the global objective function and infeasibility of the VI constraints measured by a suitably defined dual gap function. Finally, the proposed method appears to be the first fully iterative scheme equipped with iteration complexity that can address distributed optimization problems with VI constraints over cycle graphs.
Optimization problems become fundamentally challenging as the number of variables increases. Because the volume of the search space grows exponentially, classical algorithms frequently fail to locate the global minimum of non-convex functions. While quantum optimization offers a potential alternative, mapping continuous problems onto near-term quantum hardware introduces severe scaling limits and barren plateaus. To bridge this gap, we propose the Distributed Quantum-Enhanced Optimization (D-QEO) framework. Instead of forcing the quantum processor to find the exact minimum, we use it simply as a topographical preconditioner. The QPU maps the landscape to locate the most promising basin of attraction, generating high-quality seed points for a classical GPU-accelerated solver to refine. To make this approach viable for utility-scale problems, we exploit the mathematical structure of separable functions. This allows us to cut a 50-qubit (i.e., $2^{50}$) global search space into independent and manageable sub-spaces using 5-qubit subcircuits. By executing these fragments concurrently with CUDA-Q, we completely bypass the overhead of cross-register entanglement and classical tensor knitting for separable functions. Benchmarks on the 10-dimensional Rastrigin and Ackley functions show that D-QEO prevents the exponential failure rates observed in purely classical algorithms. Furthermore, this quantum warm-start significantly reduces the number of classical BFGS iterations required to converge, providing a highly practical blueprint for utilizing near-term quantum resources in complex global search.
Two anomaly-detection criteria are proposed for transactive energy systems with competitive markets. Participants of transactive energy systems seek an optimal power allocation through hybrid economic-control methods to facilitate the integration of various types of distributed energy resources to power distribution systems. In transactive energy systems, every participant is assumed to be a rational entity, and consumers have diminishing marginal utility and suppliers have increasing marginal cost. With the first proposed anomaly-detection criterion, the monotonicity of marginal cost and marginal utility are examined. The impact of line flow constraints is also taken into consideration. Then, the second anomaly-detection criterion is proposed for TESs with marginal cost and marginal utility which change faster than a certain rate. The second criterion is more accurate than the first one for TES with marginal cost and marginal utility which change faster than a certain rate, but it requires the knowledge of that rate. Neither criteria requires more data than those necessary to find the optimal power allocation and the market-clear price in a transactive energy system. Therefore, the proposed criteria do not disclose any more data than necessary. As the monotonicity of marginal cost and marginal utility in a TES with competitive markets results in convex objective functions in an optimization problem and strongly convex ones when marginal cost and marginal utility changes faster than a certain rate, the two detection criteria are also applicable to anomaly detection of general convex optimization problems. Simulations are carried out to show the efficacy of the proposed criteria to detect anomalies caused by cyberattacks.
Not Available
We introduce and investigate matrix approximation by decomposition into a sum of radial basis function (RBF) components. An RBF component is a generalization of the outer product between a pair of vectors, where an RBF function replaces the scalar multiplication between individual vector elements. Even though the RBF functions are positive definite, the summation across components is not restricted to convex combinations and allows us to compute the decomposition for any real matrix that is not necessarily symmetric or positive definite. We formulate the problem of seeking such a decomposition as an optimization problem with a nonlinear and non-convex loss function. Several modern versions of the gradient descent method, including their scalable stochastic counterparts, are used to solve this problem. We provide extensive empirical evidence of the effectiveness of the RBF decomposition and that of the gradient-based fitting algorithm. While being conceptually motivated by singular value decomposition (SVD), our proposed nonlinear counterpart outperforms SVD by drastically reducing the memory required to approximate a data matrix with the same L2 error for a wide range of matrix types. For example, it leads to 2 to 6 times memory save for Gaussian noise, graph adjacency matrices, and kernel matrices. Moreover, this proximity-based decomposition can offer additional interpretability in applications that involve, e.g., capturing the inner low-dimensional structure of the data, retaining graph connectivity structure, and preserving the acutance of images.
An optimization-based strategy for solving nonlinear mechanics problems is proposed. In contrast to typical nonlinear equation solver algorithms that aim to find zeros in the residual force function, we minimize an energy (or energy-like) function to encourage solutions which are locally stable equilibria. These smooth and potentially non-convex objective functions are minimized using a preconditioned conjugate-gradient trust-region algorithm. Contact is formulated as an inequality constrained minimization problem, and is solved with an augmented Lagrangian algorithm. Friction is included in the approach via a regularized quasi-potential energy, and other dissipative behavior is included through the use of variational constitutive updates. Finally, to accelerate convergence rates for the Lagrange multipliers, we propose a novel multiplier update algorithm utilizing the Fischer-Burmeister function, and demonstrate super-linear solver convergence for some applications.
An optimization algorithm for nonsmooth nonconvex constrained optimization problems with upper- \({\boldsymbol{\mathcal{C}^2}}\) objective functions is proposed and analyzed. Upper- \({\boldsymbol{\mathcal{C}^2}}\) is a weakly concave property that exists in difference of convex (DC) functions and arises naturally in many applications, particularly certain classes of solutions to parametric optimization problems e.g., recourse of stochastic programming and projection onto closed sets. The algorithm can be viewed as an extension of sequential quadratic programming (SQP) to nonsmooth problems with upper- \({\boldsymbol{\mathcal{C}^2}}\) objectives or a simplified bundle method. It is globally convergent with bounded algorithm parameters that are updated with a trust-region criterion. The algorithm handles general smooth constraints through linearization and uses a line search to ensure progress. The potential inconsistencies from the linearization of the constraints are addressed through a penalty method. In conclusion, the capabilities of the algorithm are demonstrated by solving both simple upper- \({\boldsymbol{\mathcal{C}^2}}\) problems and a real-world optimal power flow problem used in current power grid industry practices.
We present a novel method for modeling time-dependent statistics in the power signal generated by a photovoltaic (PV) system. Our white-box machine learning method is interpretable and auditable, based on principles of multiperiodic basis functions and convex optimization. Our proposed method of time dilating the daily signal to remove night time values results in a novel representation of PV power signals, evocative of a ‘Bundt cake’. The proposed model describes the marginal distribution of power output as a function of date and time. The resulting probabilistic model of a PV system can be used to perform a variety of tasks, and here, we demonstrate the application of clear sky detection.