Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “task parallelism”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Contra: A New Language for Task- and Data-Parallellism [Slides]

A new language is beneficial to take advantage of emerging architectures, and existing and new software technologies. Contra is a new language to provide performance portable code and consists of two innovations, which are described. A current status and illustrations of the code are shown.

97 MATHEMATICS AND COMPUTING↗

An Investigation of using FleCSI for Monte Carlo Radiation Transport

This document details the attempt to use FleCSI to provide MPI parallelization and domain decomposition for a Monte Carlo radiation transport application. FleCSI is a framework designed to support multi-physics application development with a focus on task-based parallelism and domain decomposition [1]. FleCSI also supports performance portability via a wrapper around a back-end performance portability layer. The goal of this work is to use FleCSI for domain decomposition and MPI parallelization inside of a Monte Carlo radiation transport application and document any pain points and shortcomings. The following section introduce the nomenclature used in FleCSI and its potential benefit for physics code developers, detail the code used to explore using FleCSI in a Monte Carlo radiation transport solver, and list the issues and concerns discovered during the work.

36 MATERIALS SCIENCE↗

GR-Athena++: Puncture Evolutions on Vertex-centered Oct-tree Adaptive Mesh Refinement

Numerical relativity is central to the investigation of astrophysical sources in the dynamical and strong-field gravity regime, such as binary black hole and neutron star coalescences. Current challenges set by gravitational-wave and multimessenger astronomy call for highly performant and scalable codes on modern massively parallel architectures. We present GR-Athena++, a general-relativistic, high-order, vertex-centered solver that extends the oct-tree, adaptive mesh refinement capabilities of the astrophysical (radiation) magnetohydrodynamics code Athena++. To simulate dynamical spacetimes, GR-Athena++ uses the Z4c evolution scheme of numerical relativity coupled to the moving puncture gauge. We demonstrate stable and accurate binary black hole merger evolutions via extensive convergence testing, cross-code validation, and verification against state-of-the-art effective-one-body waveforms. GR-Athena++ leverages the task-based parallelism paradigm of Athena++ to achieve excellent scalability. We measure strong-scaling efficiencies above 95% for up to ~1.2 × 10 4 CPUs and excellent weak scaling is shown up to ~10 5 CPUs in a production binary black hole setup with adaptive mesh refinement. GR-Athena++ thus allows for the robust simulation of compact binary coalescences and offers a viable path toward numerical relativity at exascale.

79 ASTRONOMY AND ASTROPHYSICS↗

Analyzing the Performance Trade-Off in Implementing User-Level Threads

User-level threads have been widely adopted as a means of achieving lightweight concurrent execution without the costs of OS-level threads. Nevertheless, the costs of managing user-level threads represent a performance barrier that dictates how fine grained the concurrency exposed by an application can be without incurring significant overheads; this in turn may translate into insufficient parallelism to exploit highly parallel systems. This article is a deep dive into the fundamental costs in implementing user-level threads. We first identify that one of the highest sources of fork-join overheads stems from deviations, events that incur context switching during the execution of a thread and disrupt a run-to-completion execution. We then conduct an in-depth investigation of a wide spectrum of methods with respect to how they handle deviations while covering both parent- and child-first scheduling policies. Our methodology involves a comprehensive instruction- and cache-level analysis of all methods on several modern CPU architectures. Finally, the primary finding of our evaluation is that dynamic promotion methods that assume the absence of deviation and dynamically provide context-switching support offer the best trade-off between performance and capability when the likelihood of deviation is low.

97 MATHEMATICS AND COMPUTING↗

EQC: Ensembled Quantum Computing for Variational Quantum Algorithms

Variational quantum algorithms (VQA), which are comprised of a classical optimizer and a parameterized quantum circuit, emerges as one of the most promising approaches of harvesting quantum power in the noisy-intermediate-scale-quantum (NISQ) era. However, the deployment of VQAs on today's NISQ devices often faces considerable system noise and prohibitively slow training speeds. On the other hand, the expensive supporting sources and infrastructure make quantum computers extremely keen on high utilization. In this paper, we propose a novel way of thinking about a quantum backend: rather than relying on one physical device which tends to introduce platform-specific noise and bias, a quantum ensemble, which distributes quantum tasks across parallel devices, can serve as a virtualized quantum computer for offering reduced noise levels through an adaptive mixture and also provide significantly improved training speeds through parallelization. With this idea, we build a distributive VQA optimization framework called DVQA, serving as the first effort in adopting parallel quantum devices for cooperative VQA training. To further constraint noise and speed-up convergence, we design a model for individual NISQ devices concerning their properties and running conditions, and propose a weighting mechanism for regularizing the returned gradients. Extensive evaluations on 10 IBM-Q quantum devices using the VQE example show that the distributive VQA training framework can substantially boost the training speed by 10.5x on average (up to 86x and at least 5.2x) with improved training accuracy.

Stein, Samuel A.↗

Evolution of the SLATE linear algebra library

SLATE (Software for Linear Algebra Targeting Exascale) is a distributed, dense linear algebra library targeting both CPU-only and GPU-accelerated systems, developed over the course of the Exascale Computing Project (ECP). While it began with several documents setting out its initial design, significant design changes occurred throughout its development. In some cases, these were anticipated: an early version used a simple consistency flag that was later replaced with a full-featured consistency protocol. In other cases, performance limitations and software and hardware changes prompted a redesign. Sequential communication tasks were parallelized; host-to-host MPI calls were replaced with GPU device-to-device MPI calls; more advanced algorithms such as Communication Avoiding LU and the Random Butterfly Transform (RBT) were introduced. Early choices that turned out to be cumbersome, error prone, or inflexible have been replaced with simpler, more intuitive, or more flexible designs. Applications have been a driving force, prompting a lighter weight queue class, nonuniform tile sizes, and more flexible MPI process grids. Of paramount importance has been building a portable library that works across several different GPU architectures – AMD, Intel, and NVIDIA – while keeping a clean and maintainable codebase. Here we explore the evolving design choices and their effects, both in terms of performance and software sustainability.

Gates, Mark↗

GLUE Code: A framework handling communication and interfaces between scales

Many scientific applications are inherently multiscale in nature. Such complex physical phenomena often require simultaneous execution and coordination of simulations spanning multiple time and length scales. This is possible by combining expensive small-scale simulations (such as molecular dynamics simulations) with larger scale simulations (such continuum limit/hydro solvers) to allow for considerably larger systems using task and data parallelism. However, the granularity of the tasks can be very large and often leads to load imbalance. Traditionally, we use approximations to streamline the computation of the more costly interactions and this introduces trade-offs between simulation cost and accuracy. In recent years, the available computational power and the advances in machine learning have made computing these scale-bridging interactions and multiscale simulations more feasible. One driving application has been plasma modeling in inertial confinement fusion (ICF), which is fundamentally multiscale in nature. This requires deep understanding of how to extrapolate microscopic information into macroscopically relevant scales. For example, in ICF one needs an accurate understanding of the connection between experimental observables and the underlying microphysics. The properties of the larger scales are often affected by the microscale behavior incorporated usually into the equations of state and ionic and electronic transport coefficients (Liboff, 1959; Rinderknecht et al., 2014; Rosenberg et al., 2015; Ross et al., 2017). Instead of incorporating this information using reliable molecular dynamics (MD) simulations, one often needs to use theoretical models, due to the inability of MD to reach engineering scales (Glosli et al., 2007; Marinak et al., 1998). One approach to resolve this issue is by coupling two MD simulations of different scales via force interpolation, e.g., the AdResS method (Krekeler et al., 2018; Nagarajan et al., 2013). Another approach, which we will pursue in the scope of this work, is by enabling scale bridging between MD simulations and meso/macro-scale models through the development and support of application programming interfaces that these different applications can interact through.

54 ENVIRONMENTAL SCIENCES↗

VA EDH Advanced Software Pipeline Framework Report: Enhancing Automation and Scalability

The VA Environmental Determinants of Health (EDH) Advanced Software Pipeline Framework is designed to enhance the efficiency, scalability, and security of geospatial data processing workflows. This framework integrates modern data orchestration and containerization technologies, including Prefect for workflow automation, Docker for containerization, and PostgreSQL/PostGIS for geospatial data storage and analysis. It ensures standardized, reproducible, and automated data processing, supporting VA objectives related to substance use risk assessment and recovery research. The pipeline addresses key scalability and performance challenges through horizontal and vertical scaling, high-performance computing (HPC) integration, parallel processing, task caching, and dynamic resource allocation. These optimizations improve throughput and reduce latency, allowing the system to efficiently manage large and complex datasets. Additionally, security and compliance measures—such as data encryption (SSL), Role-Based Access Control (RBAC), and adherence to GDPR and HIPAA standards—safeguard sensitive information throughout data transmission and storage. A key implementation of this framework includes the automation of shelter list geolocation workflows, ensuring that up-to-date data is readily available for VA decision-making. Lessons learned from this project include the transition from in-memory processing to incremental storage writes, improving resource management and reliability. Future enhancements aim to expand automation, integrate AI-driven anomaly detection, and incorporate high-performance computing resources. This framework provides a scalable, secure, and adaptable solution for managing geospatial datasets, reinforcing the VA’s ability to support clinical and strategic initiatives through data-driven decision-making.

97 MATHEMATICS AND COMPUTING↗

Distributed Data-Driven Optimization for Voltage Regulation in Distribution Systems

Here, this paper proposes a distributed data-driven optimization framework for voltage regulation in distribution systems. The recursive kernel regression and alternating direction method of multipliers (ADMM) are selected to cover the system learning and distributed optimization tasks. The proposed distributed data-driven framework is capable of having a rapid response to system or load changes while considering the operation optimality. Besides, the distributed algorithm parallels the computation tasks and reduces the computational expense of a single agent. To validate the performance of the proposed method, a hypothetical 7-Bus system and the IEEE 123-Bus system are selected to show the effectiveness of the proposed data-driven framework. According to the numerical study results, the proposed method offers great flexibility for selecting customized kernel models for different regions and can effectively improve the system voltage profile in a distributed manner.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Analysis of Threading Libraries for High Performance Computing

With the appearance of multi-/many core machines, applications and runtime systems have evolved in order to exploit the new on-node concurrency brought by new software paradigms. POSIX threads (Pthreads) was widely-adopted for that purpose and it remains as the most used threading solution in current hardware. Lightweight thread (LWT) libraries emerged as an alternative offering lighter mechanisms to tackle the massive concurrency of current hardware. In this article, we analyze in detail the most representative threading libraries including Pthread- and LWT-based solutions. In addition, to examine the suitability of LWTs for different use cases, we develop a set of microbenchmarks consisting of OpenMP patterns commonly found in current parallel codes, and we compare the results using threading libraries and OpenMP implementations. Moreover, we study the semantics offered by threading libraries in order to expose the similarities among different LWT application programming interfaces and their advantages over Pthreads. This article exposes that LWT libraries outperform solutions based on operating system threads when tasks and nested parallelism are required.

GLT↗

All-electron APW+ lo calculation of magnetic molecules with the SIRIUS domain-specific package

We report APW+lo (augmented plane wave plus local orbital) density functional theory (DFT) calculations of large molecular systems using the domain specific SIRIUS multi-functional DFT package. The APW and FLAPW (full potential linearized APW) task and data parallelism options and the advanced eigen-system solver provided by SIRIUS can be exploited for performance gains in ground state Kohn–Sham calculations on large systems. This approach is distinct from our prior use of SIRIUS as a library backend to another APW+lo or FLAPW code. We benchmark the code and demonstrate performance on several magnetic molecule and metal organic framework systems. We show that the SIRIUS package in itself is capable of handling systems as large as a several hundred atoms in the unit cell without having to make technical choices that result in the loss of accuracy with respect to that needed for the study of magnetic systems.

Chemistry↗

Adaptive Spatially Aware I/O for Multiresolution Particle Data Layouts

Large-scale simulations on nonuniform particle distributions that evolve over time are widely used in cosmology, molecular dynamics, and engineering. Such data are often saved in an unstructured format that neither preserves spatial locality nor provides metadata for accelerating spatial or attribute subset queries, leading to poor performance of visualization tasks. Furthermore, the parallel I/O strategy used typically writes a file per process or a single shared file, neither of which is portable or scalable across different HPC systems. We present a portable technique for scalable, spatially aware adaptive aggregation that preserves spatial locality in the output. We evaluate our approach on two supercomputers, Stampede2 and Summit, and demonstrate that it outperforms prior approaches at scale, achieving up to 2.5× faster writes and reads for nonuniform distributions. Furthermore, the layout written by our method is directly suitable for visual analytics, supporting low-latency reads and attribute-based filtering with little overhead.

Usher, Will↗

Stochastic Vector Techniques in Ground-State Electronic Structure

Herein we review a suite of stochastic vector computational approaches for studying the electronic structure of extended condensed matter systems. These techniques help reduce algorithmic complexity, facilitate efficient parallelization, simplify computational tasks, accelerate calculations, and diminish memory requirements. While their scope is vast, we limit our study to ground-state and finite temperature density functional theory (DFT) and second-order many-body perturbation theory. More advanced topics, such as quasiparticle (charge) and optical (neutral) excitations and higher-order processes, are covered elsewhere. We start by explaining how to use stochastic vectors in computations, characterizing the associated statistical errors. Next, we show how to estimate the electron density in DFT and discuss effective techniques to reduce statistical errors. Finally, we review the use of stochastic vectors for calculating correlation energies within the second-order Møller-Plesset perturbation theory and its finite temperature variational form. Example calculation results are presented and used to demonstrate the efficacy of the methods.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Scalable Second Order Optimization for Machine Learning

Many machine learning (ML) training tasks are essentially optimization processes that would at first glance appear eminently parallelizable and scalable. However, effective acceleration of these tasks with scalable parallel hardware has proven to be elusive. While standard methods for machine learning, e.g., stochastic gradient descent (SGD) for DNNs, tend to be resource efficient, they appear to be fundamentally sequential in nature.

97 MATHEMATICS AND COMPUTING↗

Classic and Quantum Task-Based Intelligent Runtime for QIRs Running on Multiple QPUs

High-performance computing systems are rapidly evolving into heterogeneous platforms that fuse quantum accelerators with traditional classical processing units (CPUs) and graphical processing units (GPUs). This convergence calls for runtimes capable of managing both classical and quantum workloads in a unified manner. We introduce an intelligent, task-based runtime that marries the Intelligent RuntIme System (IRIS) asynchronous scheduler with a quantum programming stack through the Quantum Intermediate Representation Execution Engine (QIR-EE). Our design allows programs written in the quantum intermediate representation (QIR) to be dispatched concurrently to a variety of back-ends, including multiple quantum simulators and nascent quantum processors, enabling genuine hybrid execution on a single node. To illustrate its practicality, we partition a 4-qubit and 20-qubit circuit into three sub-circuits using quantum circuit cutting via the QCut library. Each sub-circuit is simulated independently by the QIR-EE driver within IRIS, after which a classical post-processing step merges the simulation results to recover the outcome of the original full-circuit computation. This case study demonstrates how finer task granularity can enable the parallel execution and lower the simulation burden per quantum task while preserving overall accuracy, highlighting the feasibility of our hybrid approach.

Miniskar, Narasinga Rao [ORNL] (ORCID:000000018259↗

Managing Dynamic Workflows in BEE

BEE is a powerful tool for: Managing and visualizing scientific workflows; Simplifying workflow execution on HPC and cloud platforms. BEE supports much of the CWL specification. Did not support execution of complex ”scattering” workflows. By introducing the PseudoTask: Can generate tasks to run on variable number of inputs; BEE is another step closer to supporting the entire CWL specification; BEE can now support parallelized workflows with scattering tasks.

97 MATHEMATICS AND COMPUTING↗

T-FSM: A Scalable Distributed Task-Based System for Frequent Subgraph Pattern Mining from a Big Graph

Finding frequent subgraph patterns in a big graph is an important problem with many applications such as classifying chemical compounds and building indexes to speed up graph queries. Since this problem is NP-hard, some recent parallel and distributed systems have been developed to accelerate the mining. However, they often have a huge memory cost, very long running time, suboptimal load balancing, poor scale-out capability, and possibly inaccurate results. In this article, we propose an efficient system called T-FSM for parallel mining of frequent subgraph patterns in a big graph. T-FSM supports a new anti-monotonic frequentness measure called Fraction-Score, which is more accurate than the widely used MNI measure. The execution engine of T-FSM supports both intra-machine parallelism and inter-machine parallelism. For intra-machine parallelism, T-FSM adopts a novel task-based execution model to ensure high multithreading concurrency, bounded memory consumption, and effective load balancing. For inter-machine parallelism, T-FSM ensures good scale-out performance with a lightweight pattern rebalancing approach that reduces workload skewness of pattern evaluations among machines. To avoid recomputing the contexts for migrated patterns, we design a novel context cache table to support concurrent and asynchronous requesting and caching of remote context data, which can timely evict and garbage collect used pattern contexts that are no longer needed to keep memory consumption bounded. Extensive experiments show that T-FSM is orders of magnitude faster than existing state-of-the-art parallel systems (more than 10×, 51×, 131×, 55× speedup over ScaleMine, DistGraph, Pangolin and Peregrine, respectively) and distributed systems (more than 42× and 88× over ScaleMine and DistGraph, respectively) for frequent subgraph pattern mining, and it scales out satisfactorily to 512 CPU cores on the Polaris supercomputer at Argonne National Laboratory.

97 MATHEMATICS AND COMPUTING↗