Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Convex functions”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

RAISHIN: A High-Resolution Three-Dimensional General Relativistic Magnetohydrodynamics Code

We have developed a new three-dimensional general relativistic magnetohydrodynamic (GRMHD) code, RAISHIN, using a conservative, high resolution shock-capturing scheme. The numerical fluxes are calculated using the Harten, Lax, & van Leer (HLL) approximate Riemann solver scheme. The flux-interpolated, constrained transport scheme is used to maintain a divergence-free magnetic field. In order to examine the numerical accuracy and the numerical efficiency, the code uses four different reconstruction methods: piecewise linear methods with Minmod and MC slope-limiter function, convex essentially non-oscillatory (CENO) method, and piecewise parabolic method (PPM) using multistep TVD Runge-Kutta time advance methods with second and third-order time accuracy. We describe code performance on an extensive set of test problems in both special and general relativity. Our new GRMHD code has proven to be accurate in second order and has successfully passed with all tests performed, including highly relativistic and magnetized cases in both special and general relativity.

Mizuno, Yosuke↗

Random Variables with Moment-Matching Staircase Density Functions

This paper proposes a family of random variables for uncertainty modeling. The variables of interest have a bounded support set, and prescribed values for the first four moments. We present the feasibility conditions for the existence of any of such variables, and propose a class of variables that conforms to such constraints. This class is called staircase because the density of its members is a piecewise constant function. Convex optimization is used to calculate their distributions according to several optimality criteria, including maximal entropy and maximal log-likelihood. The flexibility and efficiency of staircases enable modeling phenomena having a possibly skewed and/or multimodal response at a low computational cost. Furthermore, we provide a means to account for the uncertainty in the distribution caused by estimating staircases from data. These ideas are illustrated by generating empirical staircase predictor models. We consider the case in which the predictor matches the sample moments exactly (a setting applicable to large datasets), as well as the case in which the predictor accounts for the sampling error in such moments (a setting applicable to sparse datasets). A predictor model for the dynamics of an aeroelastic airfoil subject to flutter instability is used as an example. The resulting predictor not only describes the system's response accurately, but also enables carrying out a risk analysis for safe flight.

Luis G. Crespo↗

Distributed Quantum-Enhanced Optimization: A Topographical Preconditioning Approach for High-Dimensional Search

Optimization problems become fundamentally challenging as the number of variables increases. Because the volume of the search space grows exponentially, classical algorithms frequently fail to locate the global minimum of non-convex functions. While quantum optimization offers a potential alternative, mapping continuous problems onto near-term quantum hardware introduces severe scaling limits and barren plateaus. To bridge this gap, we propose the Distributed Quantum-Enhanced Optimization (D-QEO) framework. Instead of forcing the quantum processor to find the exact minimum, we use it simply as a topographical preconditioner. The QPU maps the landscape to locate the most promising basin of attraction, generating high-quality seed points for a classical GPU-accelerated solver to refine. To make this approach viable for utility-scale problems, we exploit the mathematical structure of separable functions. This allows us to cut a 50-qubit (i.e., $2^{50}$) global search space into independent and manageable sub-spaces using 5-qubit subcircuits. By executing these fragments concurrently with CUDA-Q, we completely bypass the overhead of cross-register entanglement and classical tensor knitting for separable functions. Benchmarks on the 10-dimensional Rastrigin and Ackley functions show that D-QEO prevents the exponential failure rates observed in purely classical algorithms. Furthermore, this quantum warm-start significantly reduces the number of classical BFGS iterations required to converge, providing a highly practical blueprint for utilizing near-term quantum resources in complex global search.

Soos, Dominik [Old Dominion U.]↗

On Anomaly Detection for Transactive Energy Systems with Competitive Market

Two anomaly-detection criteria are proposed for transactive energy systems with competitive markets. Participants of transactive energy systems seek an optimal power allocation through hybrid economic-control methods to facilitate the integration of various types of distributed energy resources to power distribution systems. In transactive energy systems, every participant is assumed to be a rational entity, and consumers have diminishing marginal utility and suppliers have increasing marginal cost. With the first proposed anomaly-detection criterion, the monotonicity of marginal cost and marginal utility are examined. The impact of line flow constraints is also taken into consideration. Then, the second anomaly-detection criterion is proposed for TESs with marginal cost and marginal utility which change faster than a certain rate. The second criterion is more accurate than the first one for TES with marginal cost and marginal utility which change faster than a certain rate, but it requires the knowledge of that rate. Neither criteria requires more data than those necessary to find the optimal power allocation and the market-clear price in a transactive energy system. Therefore, the proposed criteria do not disclose any more data than necessary. As the monotonicity of marginal cost and marginal utility in a TES with competitive markets results in convex objective functions in an optimization problem and strongly convex ones when marginal cost and marginal utility changes faster than a certain rate, the two detection criteria are also applicable to anomaly detection of general convex optimization problems. Simulations are carried out to show the efficacy of the proposed criteria to detect anomalies caused by cyberattacks.

Wang, Peng↗

Nonlinear Matrix Approximation with Radial Basis Function Components

We introduce and investigate matrix approximation by decomposition into a sum of radial basis function (RBF) components. An RBF component is a generalization of the outer product between a pair of vectors, where an RBF function replaces the scalar multiplication between individual vector elements. Even though the RBF functions are positive definite, the summation across components is not restricted to convex combinations and allows us to compute the decomposition for any real matrix that is not necessarily symmetric or positive definite. We formulate the problem of seeking such a decomposition as an optimization problem with a nonlinear and non-convex loss function. Several modern versions of the gradient descent method, including their scalable stochastic counterparts, are used to solve this problem. We provide extensive empirical evidence of the effectiveness of the RBF decomposition and that of the gradient-based fitting algorithm. While being conceptually motivated by singular value decomposition (SVD), our proposed nonlinear counterpart outperforms SVD by drastically reducing the memory required to approximate a data matrix with the same L2 error for a wide range of matrix types. For example, it leads to 2 to 6 times memory save for Gaussian noise, graph adjacency matrices, and kernel matrices. Moreover, this proximity-based decomposition can offer additional interpretability in applications that involve, e.g., capturing the inner low-dimensional structure of the data, retaining graph connectivity structure, and preserving the acutance of images.

Rebrova, Elizaveta↗

Optimization-based algorithms for nonlinear mechanics and frictional contact

An optimization-based strategy for solving nonlinear mechanics problems is proposed. In contrast to typical nonlinear equation solver algorithms that aim to find zeros in the residual force function, we minimize an energy (or energy-like) function to encourage solutions which are locally stable equilibria. These smooth and potentially non-convex objective functions are minimized using a preconditioned conjugate-gradient trust-region algorithm. Contact is formulated as an inequality constrained minimization problem, and is solved with an augmented Lagrangian algorithm. Friction is included in the approach via a regularized quasi-potential energy, and other dissipative behavior is included through the use of variational constitutive updates. Finally, to accelerate convergence rates for the Lagrange multipliers, we propose a novel multiplier update algorithm utilizing the Fischer-Burmeister function, and demonstrate super-linear solver convergence for some applications.

42 ENGINEERING↗

A Sequential Quadratic Programming Algorithm for Nonsmooth Problems with Upper- \({\boldsymbol{\mathcal{C}^2}}\) Objective

An optimization algorithm for nonsmooth nonconvex constrained optimization problems with upper- \({\boldsymbol{\mathcal{C}^2}}\) objective functions is proposed and analyzed. Upper- \({\boldsymbol{\mathcal{C}^2}}\) is a weakly concave property that exists in difference of convex (DC) functions and arises naturally in many applications, particularly certain classes of solutions to parametric optimization problems e.g., recourse of stochastic programming and projection onto closed sets. The algorithm can be viewed as an extension of sequential quadratic programming (SQP) to nonsmooth problems with upper- \({\boldsymbol{\mathcal{C}^2}}\) objectives or a simplified bundle method. It is globally convergent with bounded algorithm parameters that are updated with a trust-region criterion. The algorithm handles general smooth constraints through linearization and uses a line search to ensure progress. The potential inconsistencies from the linearization of the constraints are addressed through a penalty method. In conclusion, the capabilities of the algorithm are demonstrated by solving both simple upper- \({\boldsymbol{\mathcal{C}^2}}\) problems and a real-world optimal power flow problem used in current power grid industry practices.

97 MATHEMATICS AND COMPUTING↗

Time Dilated Bundt Cake Analysis of PV Output [Poster]

We present a novel method for modeling time-dependent statistics in the power signal generated by a photovoltaic (PV) system. Our white-box machine learning method is interpretable and auditable, based on principles of multiperiodic basis functions and convex optimization. Our proposed method of time dilating the daily signal to remove night time values results in a novel representation of PV power signals, evocative of a ‘Bundt cake’. The proposed model describes the marginal distribution of power output as a function of date and time. The resulting probabilistic model of a PV system can be used to perform a variety of tasks, and here, we demonstrate the application of clear sky detection.

14 SOLAR ENERGY↗

Triangle Method for Dense ReLU Layers [SWR-25-72]

This software is an implementation of the methods for initializing and training neural networks to be more efficient per parameter, described more fully below and in the related publication: In theory, depth should make a ReLU network EXPONENTIALLY more efficient by enabling it to produce an exponential number of piecewise linear sections in its output. This reasoning is largely based on the work of mathematicians that have hand-constructed networks that make good use of depth. In practice however, even very deep ReLU networks that have been randomly initialized will behave identically to their shallow counterparts - missing an entire exponential dimension of efficiency. The triangle method is a first attempt at realizing the exponential potential of deep networks. Instead of randomly setting weights, we force pairs of neurons in each layer learn to build triangles (i.e. functions from [0,1] -> [0,1] that look like triangles). This is a very efficient pattern for generating lots of linear pieces because composing two triangular functions doubles the number of pieces with each composition. The triangle method is more than just a different initialization, it is a new paradigm of training. Instead of making direct updates to the matrix weights, we do an extra step of backpropagation to collect the derivatives of the loss function with respect to the shapes of the triangles, training them to tilt left or right. This process essentially holds the networks hand throughout the loss landscape and forces it to always use depth effectively by producing triangular shapes internally. This can produce several orders of magnitude of improvement on convex one-dimensional regression problems. Much more theoretical work is needed to realize its full potential beyond this context, but the implementation in this repository will still work in arbitrary numbers of dimensions. The file Triangle_Method.py is a generalized form of the method that will build each neuron its own custom 1-d convex activation function (with exponential efficiency). Example usage on one dimensional problems can be found in Example_Usage.ipynb and an example of using this in a real neural network can be found in Example_VGG16_CIFAR10.ipynb.

Milkert, Max [National Renewable Energy Laboratory↗

Rao-Blackwellization for Adaptive Gaussian Sum Nonlinear Model Propagation

When dealing with imperfect data and general models of dynamic systems, the best estimate is always sought in the presence of uncertainty or unknown parameters. In many cases, as the first attempt, the Extended Kalman filter (EKF) provides sufficient solutions to handling issues arising from nonlinear and non-Gaussian estimation problems. But these issues may lead unacceptable performance and even divergence. In order to accurately capture the nonlinearities of most real-world dynamic systems, advanced filtering methods have been created to reduce filter divergence while enhancing performance. Approaches, such as Gaussian sum filtering, grid based Bayesian methods and particle filters are well-known examples of advanced methods used to represent and recursively reproduce an approximation to the state probability density function (pdf). Some of these filtering methods were conceptually developed years before their widespread uses were realized. Advanced nonlinear filtering methods currently benefit from the computing advancements in computational speeds, memory, and parallel processing. Grid based methods, multiple-model approaches and Gaussian sum filtering are numerical solutions that take advantage of different state coordinates or multiple-model methods that reduced the amount of approximations used. Choosing an efficient grid is very difficult for multi-dimensional state spaces, and oftentimes expensive computations must be done at each point. For the original Gaussian sum filter, a weighted sum of Gaussian density functions approximates the pdf but suffers at the update step for the individual component weight selections. In order to improve upon the original Gaussian sum filter, Ref. [2] introduces a weight update approach at the filter propagation stage instead of the measurement update stage. This weight update is performed by minimizing the integral square difference between the true forecast pdf and its Gaussian sum approximation. By adaptively updating each component weight during the nonlinear propagation stage an approximation of the true pdf can be successfully reconstructed. Particle filtering (PF) methods have gained popularity recently for solving nonlinear estimation problems due to their straightforward approach and the processing capabilities mentioned above. The basic concept behind PF is to represent any pdf as a set of random samples. As the number of samples increases, they will theoretically converge to the exact, equivalent representation of the desired pdf. When the estimated qth moment is needed, the samples are used for its construction allowing further analysis of the pdf characteristics. However, filter performance deteriorates as the dimension of the state vector increases. To overcome this problem Ref. [5] applies a marginalization technique for PF methods, decreasing complexity of the system to one linear and another nonlinear state estimation problem. The marginalization theory was originally developed by Rao and Blackwell independently. According to Ref. [6] it improves any given estimator under every convex loss function. The improvement comes from calculating a conditional expected value, often involving integrating out a supportive statistic. In other words, Rao-Blackwellization allows for smaller but separate computations to be carried out while reaching the main objective of the estimator. In the case of improving an estimator's variance, any supporting statistic can be removed and its variance determined. Next, any other information that dependents on the supporting statistic is found along with its respective variance. A new approach is developed here by utilizing the strengths of the adaptive Gaussian sum propagation in Ref. [2] and a marginalization approach used for PF methods found in Ref. [7]. In the following sections a modified filtering approach is presented based on a special state-space model within nonlinear systems to reduce the dimensionality of the optimization problem in Ref. [2]. First, the adaptive Gaussian sum propagation is explained and then the new marginalized adaptive Gaussian sum propagation is derived. Finally, an example simulation is presented.

state estimation↗

Subcell limiting strategies for discontinuous Galerkin spectral element methods

Here, we present a general family of subcell limiting strategies to construct robust high-order accurate nodal discontinuous Galerkin (DG) schemes. The main strategy is to construct compatible low order finite volume (FV) type discretizations that allow for convex blending with the high-order variant with the goal of guaranteeing additional properties, such as bounds on physical quantities and/or guaranteed entropy dissipation. For an implementation of this main strategy, four main ingredients are identified that may be combined in a flexible manner: (i) a nodal high-order DG method on Legendre–Gauss–Lobatto nodes, (ii) a compatible robust subcell FV scheme, (iii) a convex combination strategy for the two schemes, which can be element-wise or subcell-wise, and (iv) a strategy to compute the convex blending factors, which can be either based on heuristic troubled-cell indicators, or using ideas from flux-corrected transport methods. By carefully designing the metric terms of the subcell FV method, the resulting methods can be used on unstructured curvilinear meshes, are locally conservative, can handle strong shocks efficiently while directly guaranteeing physical bounds on quantities such as density, pressure or entropy. We further show that it is possible to choose the four ingredients to recover existing methods such as a provably entropy dissipative subcell shock-capturing approach or a sparse invariant domain preserving approach. We test the versatility of the presented strategies and mix and match the four ingredients to solve challenging simulation setups, such as the KPP problem (a hyperbolic conservation law with non-convex flux function), turbulent and hypersonic Euler simulations, and MHD problems featuring shocks and turbulence.

97 MATHEMATICS AND COMPUTING↗

Multi-variance replica exchange SGMCMC for inverse and forward problems via Bayesian PINN

Physics-informed neural network (PINN) has been successfully applied in solving a variety of nonlinear non-convex forward and inverse problems. However, the training is challenging because of the non-convex loss functions and the multiple optima in the Bayesian inverse problem. In this work, we propose a multi-variance replica exchange stochastic gradient Langevin dynamics method to tackle the challenge of the multiple local optima in the optimization and the challenge of the multiple modal posterior distribution in the inverse problem. Replica exchange methods are capable of escaping from the local traps and accelerating the convergence; two chains with different temperatures are designed where the low temperature chain aims for the local convergence, and the target of the high temperature chain is to travel globally and explore the whole loss function entropy landscape. However, it may not be efficient to solve mathematical inversion problems by using the vanilla replica method directly since the method doubles the computational cost in evaluating the forward solvers (likelihood functions) in the two chains. To address this issue, we propose to make different assumptions on the energy function estimation and this facilities one to use solvers of different fidelities in the likelihood function evaluation. More precisely, one can use a solver with low fidelity in the high temperature chain while using a solver with high fidelity in the low temperature chain. Our proposed method significantly lowers the computational cost in the high temperature chain, meanwhile preserving the accuracy and converging very fast. Here we give an unbiased estimate of the swapping rate and give an estimation of the discretization error of the scheme. To verify our idea, we design and solve four inverse problems which have multiple modes. The proposed method is also employed to train the Bayesian PINN to solve the forward and inverse problems; faster and more accurate convergence has been observed when compared to the stochastic gradient Langevin dynamics (SGLD) method and vanilla replica exchange methods.

97 MATHEMATICS AND COMPUTING↗

Accelerating gradient descent and Adam via fractional gradients

Here we propose a class of novel fractional-order optimization algorithms. We define a fractional-order gradient via the Caputo fractional derivatives that generalizes integer-order gradient. We refer it to as the Caputo fractional-based gradient, and develop an efficient implementation to compute it. A general class of fractional-order optimization methods is then obtained by replacing integer-order gradients with the Caputo fractional-based gradients. To give concrete algorithms, we consider gradient descent (GD) and Adam, and extend them to the Caputo fractional GD (CfGD) and the Caputo fractional Adam (CfAdam). We demonstrate the superiority of CfGD and CfAdam on several large scale optimization problems that arise from scientific machine learning applications, such as ill-conditioned least squares problem on real-world data and the training of neural networks involving non-convex objective functions. Numerical examples show that both CfGD and CfAdam result in acceleration over GD and Adam, respectively. We also derive error bounds of CfGD for quadratic functions, which further indicate that CfGD could mitigate the dependence on the condition number in the rate of convergence and results in significant acceleration over GD.

97 MATHEMATICS AND COMPUTING↗

Stochastic exciton-scattering theory of optical line shapes: Renormalized many-body contributions

Spectral line shapes provide a window into the local environment coupled to a quantum transition in the condensed phase. In this paper, we build upon a stochastic model to account for non-stationary background processes produced by broad-band pulsed laser stimulation, as distinguished from those for stationary phonon bath. In particular, we consider the contribution of pair-fluctuations arising from the full bosonic many-body Hamiltonian within a mean-field approximation, treating the coupling to the system as a stochastic noise term. Herein, using the Itô transformation, we consider two limiting cases for our model, which lead to a connection between the observed spectral fluctuations and the spectral density of the environment. In the first case, we consider a Brownian environment and show that this produces spectral dynamics that relax to form dressed excitonic states and recover an Anderson–Kubo-like form for the spectral correlations. In the second case, we assume that the spectrum is Anderson–Kubo like and invert to determine the corresponding background. Using the Jensen inequality, we obtain an upper limit for the spectral density for the background. The results presented here provide the technical tools for applying the stochastic model to a broad range of problems.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Computational optimal transport for molecular spectra: The fully continuous case

Computational optimal transport is used to analyze the difference between pairs of continuous molecular spectra. It is demonstrated that transport distances which are derived from this approach may be a more appropriate measure of the difference between two continuous spectra than more familiar measures of distance under many common circumstances. Associated with the transport distances is the transport map which provides a detailed analysis of the difference between two molecular spectra and is a key component of our study of quantitative differences between two continuous spectra. The use of optimal transport for comparing molecular spectra is developed in detail here with a set of model spectra, so that the discussion is self-contained. The difference between the transport distance and more common definitions of distance is elucidated for some well-chosen examples and it is shown where transport distances may be very useful alternatives to standard definitions of distance. The transport distance between a theoretical and experimental electronic absorption spectrum for SO 2 is studied and it is shown how the theoretical spectrum can be modified to fit the experimental spectrum better adjusting the theoretical band origin and the resolution of the theoretical spectrum. In conclusion, this analysis includes the calculation of transport maps between the theoretical and experimental spectra suggesting future applications of the methodology.

74 ATOMIC AND MOLECULAR PHYSICS↗

Iterative Linearization for Phasor-Defined Optimal Power Dispatch

Optimal power flow (OPF) problems, which dispatch power targets to controllable generating units across a network, must generally account for non-convex constraints on power flow. Furthermore, adapting those problems so as to make them solvable with convex optimization techniques is an area of much academic and operational interest. In this paper, we present a method for solving OPF as a quadratic program by iteratively refining and re-initializing a linearized model of power flow based on the outputs of an associated nonlinear solver. The linear model on which we demonstrate this method is an adapted version of an approximation designed for use with unbalanced distribution networks. As an important benefit, the model allows for the explicit inclusion of nodal voltage phasor values in both the OPF problem's objective and its constraints, which opens the door to the idea of phasor-based control (PBC) design. We show in simulations on the IEEE 13-node test feeder that our method quickly converges to a set of phasor targets that are sufficiently precise for use in operations at the distribution level.

24 POWER TRANSMISSION AND DISTRIBUTION↗