Engineering Papers⌕ Search

Engineering topics

Yang, Ulrike

Publications and source records attributed to Yang, Ulrike.

Summarizing the interoperabilities between xSDK members

Rapid, efficient production of high-quality, sustainable extreme-scale scientific applications is best accomplished using a rich ecosystem of state-of-the art reusable libraries, tools, lightweight frameworks, and defined software methodologies, developed by a community of scientists who are striving to identify, adapt, and adopt best practices in software engineering. The vision of the xSDK is to provide infrastructure for and interoperability of a collection of related and complementary software elements — developed by diverse, independent teams throughout the high-performance computing (HPC) community — that provide the building blocks, tools, models, processes, and related artifacts for rapid and efficient development of high-quality applications. A primary component, and challenge, of the xSDK is to improve interoperability among software libraries and domain components. This document summarizes the current and planned interoperabilities of the twenty-three xSDK member packages as of March 2021. Additionally, the status of the xSDK example codes that demonstrate and test the interoperabilities within the xSDK is provided.

97 MATHEMATICS AND COMPUTING↗

Low synchronization Gram–Schmidt and generalized minimal residual algorithms

The Gram–Schmidt process uses orthogonal projection to construct the A = QR factorization of a matrix. When Q has linearly independent columns, the operator P = I - Q(QTQ)-1QT defines an orthogonal projection onto Q⊥. In finite precision, Q loses orthogonality as the factorization progresses. A family of approximate projections is derived with the form P = I - QTQT, with correction matrix T. When T = (QTQ)-1, and T is triangular, it is postulated that the best achievable orthogonality is $\mathcal{O}(ε)\mathcal{K}(A)$. We present new variants of modified (MGS) and classical Gram–Schmidt algorithms that require one global reduction step. An interesting form of the projector leads to a compact WY representation for MGS. In particular, the inverse compact WY MGS algorithm is equivalent to a lower triangular solve. Our main contribution is to introduce a backward normalization lag into the compact WY representation, resulting in a $\mathcal{O}(ε)\mathcal{K}[r_0, AV_m])$ stable Generalized Minimal Residual Method (GMRES) algorithm that requires only one global reduce per iteration. Finally, further improvements in performance are achieved by accelerating GMRES on GPUs.

97 MATHEMATICS AND COMPUTING↗