Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “optimized Schwarz method”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Quasi-Optimal Schwarz Methods for the Conforming Spectral Element Discretization

Fast methods are proposed for solving the system K(sub N)x = b resulting from the discretization of self-adjoint elliptic equations in three dimensional domains by the spectral element method. The domain is decomposed into hexahedral elements, and in each of these elements the discretization space is formed by polynomials of degree N in each variable. Gauss-Lobatto-Legendre (GLL) quadrature rules replace the integrals in the Galerkin formulation. This system is solved by the preconditioned conjugate gradients method. The conforming finite element space on the GLL mesh consisting of piecewise Q(sub 1) elements produces a stiffness matrix K(sub h) that is spectrally equivalent to the spectral element stiffness matrix K(sub N). The action of the inverse of K(sub h) is expensive for large problems, and is therefore replaced by a Schwarz preconditioner B(sub h) of this finite element stiffness matrix. The preconditioned operator then becomes B(sub h)(exp -l)K(sub N). The technical difficulties stem from the nonregularity of the mesh. Tools to estimate the convergence of a large class of new iterative substructuring and overlapping Schwarz preconditioners are developed. This technique also provides a new analysis for an iterative substructuring method proposed by Pavarino and Widlund for the spectral element discretization.

Casarin, Mario↗

Asynchronous Iterative Solvers for Extreme-Scale Computing

The Asynchronous Iterative Solvers for Extreme-Scale Computing (AsyncIS) project aims to explore more efficient numerical algorithms by decreasing their overhead. AsyncIS does this by replacing the outer Krylov subspace solver with an asynchronous optimized Schwarz method, thereby removing the global synchronization and bulk synchronous operations typically used in numerical codes. AsyncIS—a U.S. Department of Energy (DOE)-funded collaboration between Georgia Tech, the University of Tennessee, Knoxville, Temple University, and Sandia National Laboratories—also focuses on the development and optimization of asynchronous preconditioners (i.e., preconditioners that are generated and/or applied in an asynchronous fashion). The novel preconditioning algorithms that provide fine-grained parallelism enable preconditioned Krylov solvers to run efficiently on large-scale distributed systems and manycore accelerators like GPUs.

97 MATHEMATICS AND COMPUTING↗

A Fast Temporal Decomposition Procedure for Long-Horizon Nonlinear Dynamic Programming

We propose a fast temporal decomposition procedure for solving long-horizon nonlinear dynamic programs. The core of the procedure is sequential quadratic programming (SQP) that utilizes a differentiable exact augmented Lagrangian as the merit function. Within each SQP iteration, we approximately solve the Newton system using an overlapping temporal decomposition strategy. We show that the approximate search direction is still a descent direction of the augmented Lagrangian provided the overlap size and penalty parameters are suitably chosen, which allows us to establish the global convergence. Moreover, we show that a unit step size is accepted locally for the approximate search direction and further establish a uniform, local linear convergence over stages. This local convergence rate matches the rate of the recent Schwarz scheme (Na et al. 2022). However, the Schwarz scheme has to solve nonlinear subproblems to optimality in each iteration, whereas we only perform a single Newton step instead. Numerical experiments validate our theories and demonstrate the superiority of our method.

97 MATHEMATICS AND COMPUTING↗

Accelerating Multivariate Functional Approximation Computation with Domain Decomposition Techniques⋆

Modeling large datasets through Multivariate Functional Approximations (MFA) provide an elegant way to handle many visualization and scientific analysis workflows. The process necessitates scalable data partitioning methods to compute MFA representations efficiently without compromising the accuracy or continuity of the reconstructed solution. We propose a domain -decomposed method for computing the MFA with B -spline bases, which reduces the total work per task and uses a restricted Additive Schwarz (RAS) method to converge the control point data degrees -of -freedom along subdomain boundaries. We provide an in-depth analysis of the parallel approach with domain decomposition solvers, aiming to minimize local subdomain error residuals and recover high -order continuity at subdomain interfaces with appropriate choices of knot overlaps. The communication cost, determined by the overlap regions in the RAS implementation, is optimized to recover the numerical error profile of the single subdomain case. Our proposed method stands in contrast to previous methods, which typically only recover either C 0 or at best C 1 continuity for arbitrary B -spline degree expansions, or those that require post -processing to blend discontinuities in the reconstructed data. We demonstrate the effectiveness of our approach using analytical and real -world datasets in 1D, 2D, and 3D through both strong and weak scaling studies. The performance results indicate that the overall cost of computing the approximation is directly proportional to the underlying nearest -neighbor communication implementation, and is only weakly dependent on the overlap region size that determines the size of the messages. This finding underscores the efficiency and scalability of our proposed method, making it a promising solution for handling large datasets in scientific workflows.

additive Schwarz solvers↗

Optimal Polynomial Smoothers and One‐Sided V‐Cycles for Poisson Problems

The solution to the Poisson equation arising from the spectral element discretization of the incompressible Navier‐Stokes equations needs robust preconditioning strategies. One such strategy is multigrid. To realize the potential of multigrid methods, effective smoothing strategies are needed. Chebyshev polynomial smoothers, in conjunction with pointwise Jacobi or additive Schwarz methods (ASMs), prove to be an effective smoother. Other polynomial smoothers, however, may provide superior convergence to the multigrid preconditioner. The authors compare the standard Chebyshev polynomial smoothers to both the novel fourth‐kind Chebyshev polynomial smoothers proposed by Lottes as well as smoothers based on the polynomial of best uniform approximation to as proposed by Kraus, Vassilevski, and Zikatanov. At the cost of symmetry, further improvements may be made. For example, a order polynomial smoother on both sides of the V‐cycle may be substituted with an order polynomial smoother on one side at no additional cost. The choice of omitting the postsmoother in favor of higher‐order polynomial presmoothing is advantageous in cases where the multigrid approximation property constant is large. The authors consider a 2D model problem based on finite differences to motivate the choice of polynomial smoother, order, and whether to apply postsmoothing for the target application of high‐order ‐geometric multigrid methods for GPU architectures. Results from both domains demonstrate the substantial improvement of these approaches over the standard Chebyshev polynomial smoother with a symmetric V‐cycle.

97 MATHEMATICS AND COMPUTING↗

Local multiplicative Schwarz algorithms for convection-diffusion equations

We develop a new class of overlapping Schwarz type algorithms for solving scalar convection-diffusion equations discretized by finite element or finite difference methods. The preconditioners consist of two components, namely, the usual two-level additive Schwarz preconditioner and the sum of some quadratic terms constructed by using products of ordered neighboring subdomain preconditioners. The ordering of the subdomain preconditioners is determined by considering the direction of the flow. We prove that the algorithms are optimal in the sense that the convergence rates are independent of the mesh size, as well as the number of subdomains. We show by numerical examples that the new algorithms are less sensitive to the direction of the flow than either the classical multiplicative Schwarz algorithms, and converge faster than the additive Schwarz algorithms. Thus, the new algorithms are more suitable for fluid flow applications than the classical additive or multiplicative Schwarz algorithms.

Cai, Xiao-Chuan↗

Optimal Methods for Estimating Cactus Pear Biomass Using Cladode Dimensions of Morphologically Diverse Accessions

Current allometric methods for photosynthetic-stem (cladode) plants, such as cactus pear (Opuntia spp.), require refinement to be used in field settings in which diverse accessions are grown. We analysed cladode dimensional data using 14 accessions representing four species and two hybrids to quantify statistically significant morphological differences among accessions and derived cross-accession models to approximate cladode fresh weight. A Box model using cladode dimensions (e.g., length, width, thickness and diameter) and factorial combinations of these measures (e.g., length*width*thickness*diameter vs. fresh weight) resulted in the highest coefficient of determination (R 2 = 0.95 general fit) across all accessions for estimating fresh weight along with parsimony estimates using the Schwarz–Bayes Criterion (SBC), which assesses the most consistent performance on individual accessions. A Fitting-box modelling approach used the measured cladode area captured using ImageJ (R 2 = 0.93 general fit). Lastly, an Elliptical model used an elliptical approximation for the measured area and performed well over all accessions (R 2 = 0.94 general fit) while avoiding extensive manual measurements. These models meet or exceed the performance of previously published approaches when applied across morphologically diverse accessions, providing efficient tools for nondestructive estimation of cactus pear biomass under the conditions tested.

Opuntia↗

On the Convergence of Overlapping Schwarz Decomposition for Nonlinear Optimal Control

Here, we study the convergence properties of an overlapping Schwarz decomposition algorithm for solving nonlinear optimal control problems (OCPs). The algorithm decomposes the time domain into a set of overlapping subdomains, and solves all subproblems defined over subdomains in parallel. The convergence is attained by updating primal-dual information at the boundaries of overlapping subdomains. We show that the algorithm exhibits local linear convergence, and that the convergence rate improves exponentially with the overlap size. We also establish global convergence results for a general quadratic programming, which enables the application of the Schwarz scheme inside second-order optimization algorithms (e.g., sequential quadratic programming). The theoretical foundation of our convergence analysis is a sensitivity result of nonlinear OCPs, which we call "exponential decay of sensitivity" (EDS). Intuitively, EDS states that the impact of perturbations at domain boundaries (i.e., initial and terminal time) on the solution decays exponentially as one moves into the domain. Here, we expand a previous analysis available in the literature by showing that EDS holds for both primal and dual solutions of nonlinear OCPs, under uniform second-order sufficient condition, controllability condition, and boundedness condition. We conduct experiments with a quadrotor motion planning problem and a partial differential equations (PDE) control problem to validate our theory, and show that the approach is significantly more efficient than alternating direction method of multipliers and as efficient as the centralized interior-point solver.

42 ENGINEERING↗

Two-level overlapping additive Schwarz preconditioner for training scientific machine learning applications

In this work we introduce a novel two-level overlapping additive Schwarz preconditioner for accelerating the training of scientific machine learning applications. The design of the proposed preconditioner is motivated by the nonlinear two-level overlapping additive Schwarz preconditioner. The neural network parameters are decomposed into groups (subdomains) with overlapping regions. In addition, the network’s feed-forward structure is indirectly imposed through a novel subdomain-wise synchronization strategy and a coarse-level training step. Through a series of numerical experiments, which consider physicsinformed neural networks and operator learning approaches, we demonstrate that the proposed two-level preconditioner significantly speeds up the convergence of the standard (LBFGS) optimizer while also yielding more accurate machine learning models. Moreover, the devised preconditioner is designed to take advantage of model-parallel computations, which can further reduce the training time.

97 MATHEMATICS AND COMPUTING↗

Proton reconstruction with the TOTEM Roman pot detectors for high- β * LHC data

The TOTEM Roman pot detectors are used to reconstruct the transverse momentum of scattered protons and to estimate the transverse location of the primary interaction. This paper presents new methods of track reconstruction, measurements of strip-level detection efficiencies, cross-checks of the LHC beam optics, and detector alignment techniques, along with their application in the selection of signal collision events. The track reconstruction is performed by exploiting hit cluster information through a novel method using a common polygonal area in the intercept-slope plane. The technique is applied in the relative alignment of detector layers with μm precision. A tag-and-probe method is used to extract strip-level detection efficiencies. The alignment of the Roman pot system is performed through time-dependent adjustments, resulting in a position accuracy of 3 μm in the horizontal and 60 μm in the vertical directions. The goal is to provide an optimal reconstruction tool for central exclusive physics analyses based on the high-β* data-taking period at $\sqrt({s})$ = 13 TeV in 2018.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Machine learning method for enforcing variable independence in background estimation with LHC data: ABCDisCoTEC

A novel solution is presented for the problem of estimating the backgrounds of a signal search using observed data while simultaneously maximizing the sensitivity of the search to the signal. The 'ABCD method' provides a reliable framework for background estimation by partitioning events into one signal-enhanced region (A) and three background-enhanced control regions (B, C, and D) via two smoothly varying, statistically independent variables. In practice, even slight correlations between the two variables can significantly undermine the method's performance. Thus, choosing appropriate variables by hand can present a formidable challenge, especially when background and signal differ only subtly. To address this issue, the ABCD with distance correlation (ABCDisCo) method was developed to construct two learned variables via a neural network trained to provide strong signal-background discrimination with small values of the distance correlation (DisCo) measure between the two learned variables. However, relying solely on minimizing the DisCo can result in learned variables that may not have distributions of background events that are smoothly varying and localized at extreme values, as necessary for the validity of the background estimation. The ABCDisCo training enhanced with closure (ABCDisCoTEC) method is introduced to solve this issue by directly minimizing the nonclosure, expressed as a dedicated differentiable loss term. This extended method is applied to a data set of proton-proton collisions at a center-of-mass energy of 13 TeV recorded by the CMS detector at the CERN Large Hadron Collider. Additionally, given the complexity of the minimization problem with constraints on multiple loss terms, the modified differential method of multipliers is applied and shown to greatly improve the stability and robustness of the ABCDisCoTEC method, compared to grid search hyperparameter optimization procedures.

Hayrapetyan, Aram [Yerevan Phys. Inst.]↗

Atomic-Scale Surface Studies of Bulk Metallic Glasses. Final Report

Bulk metallic glasses (BMGs) are of both scientific and technological interest because the absence of periodic atomic arrangements provides them with unique physical, chemical, and mechanical properties. Their high strength, superior elasticity, and an ability to be easily formed into virtually unlimited shapes with feature sizes from centimeters to Angstroms by blow molding and thermoplastic forming makes them an attractive choice for more and more practical applications and products. Due to their complex internal structure, however, experiments that yield insight into their exact atomic arrangements have been scarce. As a result, glass physics is one of the last remaining unexplored fields of materials science despite the scientific and technological importance of glasses in our daily lives, and the question how to characterize and control matter away from equilibrium, as glasses are, was listed as one of five Grand Challenges in a recent DoE report. The main reason for the slow progress in glass physics is the lack of experimental tools that enable access to atomic-scale structural and behavioral information for disordered materials. Such atomic-scale knowledge is mandatory to establish structure-property relationships that ultimately could allow to custom-design alloys featuring specific desired characteristics. With no such relationships available, theory development in glass remains basic and the few that exist are often untested. The aim of this research was to enable progress in our understanding of BMGs by developing a new approach that will allow a meaningful application of local surface science methods to specially prepared BMG samples to obtain a wealth of quantitative information on their atomic arrangements. Key was the availability of specially prepared samples whose surfaces feature large atomically flat terraces despite being entirely amorphous, which we have produced from a Pt 57.5 Cu 14.7 Ni 5.3 P 22.5 alloy (‘Pt-BMG’) both under ambient conditions as well as in ultrahigh vacuum using a unique setup that has been specially developed within this grant. Our approach starts with the in-situ preparation of oxide crystals that are terminated by large terraces, from which exact mirror images out of BMG will be produced using thermoplastic forming (TPF); for the research within this grant, we have successfully used (001)-oriented SrTiO 3 single crystal surfaces as well as (100)-, (110)-, and (111)-oriented single crystals made from LaAlO 3 . Since the resulting BMG replicas display all features of the original crystal with sub-Angstrom fidelity, thereby mimicking the original crystal’s termination by atomically flat terraces without being crystalline themselves, they are ideally suited for further investigation. The following atomic-scale local studies were then carried within this proposal: (i) high-resolution surface imaging and local spectroscopy using scanning probe microscopy, which showed disordered atom-like features and revealed changes in the gradient of the local surface potential on a 1-2 nm length scale; (ii) characterization of atomic-scale plastic flow with affected volumes as low as 1000 atoms, which showed local hardness near or above the theoretically predicted maximum and, once plastic deformation was initiated, homogeneous flow of the atoms involved; (iii) characterization of surface relaxation processes and the onset of crystallization induced by annealing, which showed that upon heating over the material’s glass transition temperature, the surface rearranges and relaxes towards a more stable, denser packed glass, which increases on-terrace surface roughness, while surface tension smoothens step edges; and (iv) studies that investigate the dependence of the material’s mechanical properties and structure on processing parameters, revealing that relaxed glasses get denser, harder, and more elastic. In combination, this information allows to combine structural models with mechanical properties and preparation history, thereby facilitating the development of preparation-structure-property relationships for metallic glasses. With the availability of such information, bulk metallic glasses can be further optimized to be used in more and more applications in industry.

36 MATERIALS SCIENCE↗