Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “user-level threads”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Analyzing the Performance Trade-Off in Implementing User-Level Threads

User-level threads have been widely adopted as a means of achieving lightweight concurrent execution without the costs of OS-level threads. Nevertheless, the costs of managing user-level threads represent a performance barrier that dictates how fine grained the concurrency exposed by an application can be without incurring significant overheads; this in turn may translate into insufficient parallelism to exploit highly parallel systems. This article is a deep dive into the fundamental costs in implementing user-level threads. We first identify that one of the highest sources of fork-join overheads stems from deviations, events that incur context switching during the execution of a thread and disrupt a run-to-completion execution. We then conduct an in-depth investigation of a wide spectrum of methods with respect to how they handle deviations while covering both parent- and child-first scheduling policies. Our methodology involves a comprehensive instruction- and cache-level analysis of all methods on several modern CPU architectures. Finally, the primary finding of our evaluation is that dynamic promotion methods that assume the absence of deviation and dynamically provide context-switching support offer the best trade-off between performance and capability when the likelihood of deviation is low.

97 MATHEMATICS AND COMPUTING↗

Qthreads Support for MPICH

SAND2024-00944O Qthread Support for MPICH is software that provides additions needed to enable the use of the Qthreads library. The high-performance message passing interface (MPICH) is an open-source implementation of MPI mainly developed and distributed by Argonne National Laboratory. Qthreads is a lightweight, user-level threading library developed and distributed by Sandia National Laboratories. MPICH currently supports Posix threads, Windows threads, and Argobots. This software enables parallel programs built with the MPICH implementation of MPI to use Qthreads user-level threads rather than Posix system-level threads. The software uses existing infrastructure in MPICH to interface to the Qthreads library. The existing interfaces in MPICH allow for creating, destroying, and managing the execution of multiple threads within a process and this software translates these calls to the equivalent Qthreads library functions. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Ciesko, Jan↗

Taking the MPI standard and the open MPI library to exascale

The Open MPI for Exascale (OMPI-X) project was one of two in the Exascale Computing Project (ECP) focused on advancing the MPI ecosystem. The OMPI-X team worked with other MPI Forum members to champion several important features for inclusion in the MPI 4.0, 4.1, and upcoming 5.0 MPI standard versions, in support of the needs of exascale applications and systems. The team also worked with the larger Open MPI community to bring implementations of these new features and other enhancements into Open MPI, one of the leading open-source implementations of the MPI interface. Here, this paper describes the motivation for the work of the OMPI-X project in the context of exascale computing needs, the nature of the resulting new capabilities in the MPI standard, and how they were implemented in the Open MPI library. Features include improved support for “MPI + X” programming models through partitioned communications and support for user-level threading, sessions, fault tolerance through the user-level fault mitigation (ULFM) and Reinit models, and other features. We also discuss enhancements to Open MPI providing improved performance and scalability for existing features, such as collective operations, one-sided operations, support for the Slingshot-11 interconnect of the initial exascale systems, and how the OMPI-X team worked to improve quality assurance for the Open MPI library, particularly on platforms of interest to the Department of Energy community.

97 MATHEMATICS AND COMPUTING↗

Scalable line and plane relaxation in a parallel structured multigrid solver

The efficient solution of sparse, linear systems that arise through the discretization of partial differential equations remains a key challenge for a range of high performance scientific simulations. One approach for reducing data movement and improving performance is by exposing and exploiting structure in a problem through the use of robust structured multilevel solvers. By choosing coarsening that preserves the structure of the problem, these methods maintain efficient structured computation and communication throughout the multigrid hierarchy. However, when coarsening is not permitted to be dependent on the operator, anisotropy must be addressed by the smoother — producing error compatible for coarse-grid correction with structured coarsening. Here, the components required in a scalable parallel structured solver are described with a focus on memory and communication efficiency of robust smoothers. While the implementation of communication and memory reduction techniques in smoothers integrated in a complete 3D solver present a significant engineering challenge, a novel approach is proposed that addresses these challenges systematically through a change to the solver’s execution model. Enabled by user-level threading paired with a set of data and communication abstractions, this approach permits seamless aggregation of communication in plane smoothers — directly reusing code for a 2D distributed multilevel cycle. Results show an effective reduction in communication costs for coarse-grid problems, and result in a speedup of 8.7x in smoothing routines shown in Fig. 12 using this approach. This produces a significant improvement to strong scalability while maintaining favorable weak scaling behavior. Finally, a parallel scaling study using a series of refined meshes is included that demonstrates the effectiveness of this approach in an application of interest.

97 MATHEMATICS AND COMPUTING↗

MPI Partix

SAND2022-6990 O MPI Partix is an application suite to test user-level threading (ULT) and partitioned communication in MPI. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Ciesko, Jan↗

Portable C++ Code that can Look and Feel Like Fortran Code with Yet Another Kernel Launcher (YAKL)

This paper introduces the Yet Another Kernel Launcher (YAKL) C++ portability library, which strives to enable user-level code with the look and feel of Fortran code. The intended audience includes both C++ developers and Fortran developers unfamiliar with C++. The C++ portability approach is briefly explained, YAKL’s main features are described, and code examples are given that demonstrate YAKL’s usage. YAKL fills a niche capability important particularly to scientific applications seeking to port Fortran code quickly to a portable C++ library. YAKL places heavy emphasis on simplicity, readability, and productivity with performance mainly emphasizing Graphics Processing Units (GPUs). Central to YAKL’s ability to allow Fortran-like user-level code are three features: (1) a multi-dimensional Array class that allows Fortran behavior; (2) a limited library of Fortran intrinsic functions; and (3) an efficient pool allocator that transparently enables cheap frequent allocations and deallocations of YAKL Arrays. While YAKL allows Fortran-style code, it also allows Arrays that exhibit C-like behavior as well, including row-major index ordering and lower bounds of “0”. YAKL currently supports CPUs, CPU threading, and Nvidia, AMD, and Intel GPUs.

97 MATHEMATICS AND COMPUTING↗

Multiple social platforms reveal actionable signals for software vulnerability awareness: A study of GitHub, Twitter and Reddit

Software vulnerabilities are flaws in computer systems that leave users open to attack. In many cases, these vulnerabilities go unnoticed and remain unresolved in codebases. Thus, awareness of software vulnerabilities among the public is crucial to ensure effective cybersecurity practices, the development of high quality software, and ultimately national security. This awareness can be better understood by studying the spread and evolution of software vulnerability discussions in online communities. This work is the first to evaluate and contrast how discussions about software vulnerabilities spread on three social platforms -- Twitter, GitHub, and Reddit. To lay the groundwork, we showcase a novel fundamental framework for measuring information spread that identifies the spread mechanisms and observables across platforms, the units of information, and the groups of measurements that can be applied to focus on a specific phenomena e.g., information cascades. We then analyze and contrast social network topologies for three example social networks and measure the scale and speed of the spread of discussion of specific vulnerabilities to understand how far and how widely they spread, how many users participate in discussions, and the duration of their spread. To demonstrate the awareness of more impactful software vulnerabilities, a subset of our analysis focuses on vulnerabilities targeted during recent major cyber attacks as well as vulnerabilities exploited by advanced persistent threat groups. We discover that usually, vulnerability discussions start on GitHub, before occurring on Twitter and Reddit. While studying how some user-level and content-level characteristics influence vulnerability spread, we observe that Twitter discussions started by users predicted to be humans have larger size, breadth, depth, adoption rate, lifetime, and structural virality compared to those started by users predicted to be bots. On Reddit, we contrast the differences in thread structure that originate from posts with positive, negative and neutral polarity. We find that posts that are positive have larger, deeper and wider discussions compared to negative and neutral posts. We anticipate the results of our analysis to not only increase the understanding of software vulnerability awareness but also inform models for simulating information spread across multiple social environments online.

97 MATHEMATICS AND COMPUTING↗