Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “asynchronous iteration”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Asynchronous Iterative Solvers for Extreme-Scale Computing

The Asynchronous Iterative Solvers for Extreme-Scale Computing (AsyncIS) project aims to explore more efficient numerical algorithms by decreasing their overhead. AsyncIS does this by replacing the outer Krylov subspace solver with an asynchronous optimized Schwarz method, thereby removing the global synchronization and bulk synchronous operations typically used in numerical codes. AsyncIS—a U.S. Department of Energy (DOE)-funded collaboration between Georgia Tech, the University of Tennessee, Knoxville, Temple University, and Sandia National Laboratories—also focuses on the development and optimization of asynchronous preconditioners (i.e., preconditioners that are generated and/or applied in an asynchronous fashion). The novel preconditioning algorithms that provide fine-grained parallelism enable preconditioned Krylov solvers to run efficiently on large-scale distributed systems and manycore accelerators like GPUs.

97 MATHEMATICS AND COMPUTING↗

Asynchronous Iterative Solvers for Extreme-Scale Computing (Final Report)

This is the final report for the project: Asynchronous Iterative Solvers for Extreme-Scale Computing. This was a collaborative project. This report only covers the activities specific to Georgia Institute of Technology. The project investigated and developed iterative solvers that operate asynchronously, thereby avoiding the high cost of synchronization that is apparent when using standard, synchronous iterative solvers at extreme levels of parallelism.

97 MATHEMATICS AND COMPUTING↗

Asynchronous Richardson iterations: theory and practice

We consider asynchronous versions of the first- and second-order Richardson methods for solving linear systems of equations. These methods depend on parameters whose values are chosen a priori. We explore the parameter values that can be proven to give convergence of the asynchronous methods. This is the first such analysis for asynchronous second-order methods. We find that for the first-order method, the optimal parameter value for the synchronous case also gives an asynchronously convergent method. For the second-order method, the parameter ranges for which we can prove asynchronous convergence do not contain the optimal parameter values for the synchronous iteration. In practice, however, the asynchronous second-order iterations may still converge using the optimal parameter values, or parameter values close to the optimal ones, despite this result. We explore this behavior with a multithreaded parallel implementation of the asynchronous methods.

97 MATHEMATICS AND COMPUTING↗

Asynchronous domain decomposition methods for nonlinear PDEs

One- and two-level parallel asynchronous methods for the numerical solution of nonlinear systems of equations, especially those arising from (nonlinear) partial differential equations, are studied. The proposed methods are based on domain decomposition techniques. Local convergence theorems are presented in several cases, with appropriate hypotheses. Computational results on a shared memory multiprocessor machine for various problems exhibiting nonlinearities are reported, illustrating the potential of these asynchronous methods, especially for heterogeneous clusters.

97 MATHEMATICS AND COMPUTING↗

Scalable Asynchronous Domain Decomposition Solvers

We discuss how parallel implementations of linear iterative solvers generally alternate between phases of data exchange and phases of local computation. Increasingly large problem sizes and more heterogeneous compute architectures make load balancing and the design of low latency network interconnects that are able to satisfy the communication requirements of linear solvers very challenging tasks. In particular, global communication patterns such as inner products become increasingly limiting at scale. We explore the use of asynchronous communication based on one-sided Message Passing Interface primitives in the context of domain decomposition solvers. In particular, a scalable asynchronous two-level Schwarz method is presented. We discuss practical issues encountered in the development of a scalable solver and show experimental results obtained on a state-of-the-art supercomputer system that illustrate the benefits of asynchronous solvers in load balanced as well as load imbalanced scenarios. Using the novel method, we can observe speedups of up to four times over its classical synchronous equivalent.

97 MATHEMATICS AND COMPUTING↗

Neutron transport methods for multiphysics heterogeneous reactor core simulation in Griffin

Griffin is a reactor physics application based on the Multiphysics Object-Oriented Simulation Environment (MOOSE). This work discloses the methods, algorithms, and implementation for simulating heterogeneous reactor dynamics models. Griffin utilizes a discontinuous finite-element method with discrete ordinates (DFEM-S ) to discretize the field variable of the multigroup neutron transport equation. Multiphysics feedback is handled using two-step tabulated cross-section methodology. Feedback quantities are evaluated using the MOOSE-MultiApp system to couple various engineering phenomena, such as heat conduction and thermal fluids. The multiphysics DFEM-S system is solved using fixed-point iteration with a fully asynchronous parallel sweeper, unstructured coarse-mesh finite difference acceleration, and a multi-timescale improved quasi-static method scheme. The implementation is applied to a multiphysics microreactor model, with two transients: one initiated by a single heat-pipe failure and another by control drum rotation. Importantly, these examples demonstrate the ability of Griffin to tractably solve the neutron transport equation considering seven independent variables and feedback.

97 MATHEMATICS AND COMPUTING↗

Error field detection and correction studies towards ITER operation

In magnetic fusion devices, error field (EF) sources, spurious magnetic field perturbations, need to be identified and corrected for safe and stable (disruption-free) tokamak operation. Within Work Package Tokamak Exploitation RT04, a series of studies have been carried out to test the portability of the novel non-disruptive method, designed and tested in DIII-D (Paz-Soldan et al 2022 Nucl. Fusion62 126007), and to perform an assessment of model-based EF control strategies towards their applicability in ITER. In this paper, the lessons learned, the physical mechanism behind the magnetic island healing, which relies on enhanced viscous torque that acts against the static electro-magnetic torque, and the main control achievements are reported, together with the first design of the asynchronous EF correction current/density controller for ITER.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Graph-based Reversible Evaluation and Tangents Library

GRETL is a C++ library for evaluation, re-evaluation and algorithmic differentiation of functional operations on an arbitrary computational graph with limited memory usage. Similar to popular machine learning frameworks in Python, like PyTorch and JAX, it tracks and stores both operations and output data as functions are evaluated. Once this composition of functions is built up, the entire chain of operations can be back propagated to compute sensitivities of the final result with respect to any number of inputs. In contrast to most machine learning applications, memory usage becomes the bottleneck for back propagation in many physics applications, especially for time-dependent PDEs. Dynamic check pointing becomes essential. An important distinguishing feature of GRETL is its ability to limit the maximum memory usage by automatically dynamic checkpointing the data output for each graph operation (see Wang, Moin, Iaccarino, 2009). During backpropagation, parts of the graph that are no longer in memory are automatically re-evaluated from upstream checkpointed states as needed for derivative sensitivity calculations (or more precisely, for vector-Jacobian products). GRETL is particularly beneficial for applications, such as coupled multi-physics, where deriving adjoint-based sensitivities and managing checkpoint memory across modules becomes onerous. Cases which can be readily handled by the GRETL library include: different time-integration algorithms per physics (e.g., coupled predictor-corrector algorithms, IMEX, etc.), sub-cycling, asynchronous integrators, state dependent timestep sizes, iterative solvers and coupling algorithms, controller algorithms, and more.

Tupek, MichaelR [Lawrence Livermore National Labor↗

Data exchange and processing synchronization in distributed systems

Systems, methods, techniques and apparatuses of asynchronous communication is distributed systems are disclosed. One exemplary embodiment is a method determining, with a plurality of agent nodes structured to communicate asynchronously in a distributed system, a first set of iterations including an iteration determined by each of the plurality of agent nodes; determining, with a first agent node of the plurality of agent nodes, a local vector clock; receiving, with the first agent node, a first iteration of the first set of iterations and a remote vector clock determined based on the first iteration; updating, with the first agent node, the local vector clock based on the received remote vector clock; and determining a first iteration of a second set of iterations based on the first set of iterations after determining all iterations of the first set of iterations have been received based on the local vector clock.

Cintuglu, Mehmet H.↗

Seven open problems in applied combinatorics

We present and discuss seven different open problems in applied combinatorics. Additionally, the application areas relevant to this compilation include quantum computing, algorithmic differentiation, topological data analysis, iterative methods, hypergraph cut algorithms, and power systems.

97 MATHEMATICS AND COMPUTING↗

Benchmark Tracking System for Performance Monitoring

Benchmarking is essential for high-performance software development, particularly for monitoring performance across code iterations. This project focused on enhancing the benchmarking process for Lamellar, an asynchronous runtime for High-Performance Computing (HPC) systems developed at Pacific Northwest National Laboratory. Prior to this work, benchmark results were difficult to track and compare across code versions, presenting significant challenges in identifying performance regressions and long-term trends. The primary objective was to establish a systematic, reproducible approach for measuring performance and detecting regressions following code commits. Our methodology involved three key components: standardizing benchmark outputs, implementing data versioning, and developing analysis tools. We standardized the benchmark output format to JSON Line records containing specific fields (execution time, hardware specifications, and environmental variables). To address data management challenges, we evaluated several options and eventually chose a git repository dedicated to benchmark data. We developed a suite of Python tools that processed benchmark results, enriched them with metadata, and facilitated search in the repository. The resulting system enables more efficient filtering and comparison of performance metrics across commit histories, hardware configurations, and benchmark variants through a unified query interface. Our implementation reduces computational overhead by first checking for existing results through configuration matching before initiating new benchmark runs, thereby conserving resources. The system has been validated by Lamellar developers. It organizes results by benchmark type and build configurations for efficient retrieval. Future developments include a planned Large Language Model interface for predicting benchmark performance, incorporating the criterion package for statistical analysis, which will enable automated detection of statistically significant performance changes, and integration with continuous integration pipelines. Despite these enhancements being reserved for future work, this project has successfully provided the Lamellar development team with a framework for maintaining consistent performance standards and identifying optimization opportunities across workloads and hardware environments.

97 MATHEMATICS AND COMPUTING↗

tomoCAM : fast model-based iterative reconstruction via GPU acceleration and non-uniform fast Fourier transforms

X-ray-based computed tomography is a well established technique for determining the three-dimensional structure of an object from its two-dimensional projections. In the past few decades, there have been significant advancements in the brightness and detector technology of tomography instruments at synchrotron sources. These advancements have led to the emergence of new observations and discoveries, with improved capabilities such as faster frame rates, larger fields of view, higher resolution and higher dimensionality. These advancements have enabled the material science community to expand the scope of tomographic measurements towards increasingly in situ and in operando measurements. In these new experiments, samples can be rapidly evolving, have complex geometries and restrictions on the field of view, limiting the number of projections that can be collected. In such cases, standard filtered back-projection often results in poor quality reconstructions. Iterative reconstruction algorithms, such as model-based iterative reconstructions (MBIR), have demonstrated considerable success in producing high-quality reconstructions under such restrictions, but typically require high-performance computing resources with hundreds of compute nodes to solve the problem in a reasonable time. Here, tomoCAM , is introduced, a new GPU-accelerated implementation of model-based iterative reconstruction that leverages non-uniform fast Fourier transforms to efficiently compute Radon and back-projection operators and asynchronous memory transfers to maximize the throughput to the GPU memory. The resulting code is significantly faster than traditional MBIR codes and delivers the reconstructive improvement offered by MBIR with affordable computing time and resources. tomoCAM has a Python front-end, allowing access from Jupyter -based frameworks, providing straightforward integration into existing workflows at synchrotron facilities.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Modifying the Asynchronous Jacobi Method for Data Corruption Resilience

Moving scientific computation from high-performance computing (HPC) and cloud computing (CC) environments to devices on the edge, i.e., physically near instruments of interest, has received tremendous interest in recent years. Such edge computing environments can operate on data in situ, offering enticing benefits over data aggregation to HPC and CC facilities that include avoiding costs of transmission, increased data privacy, and real-time data analysis. Because of the inherent unreliability of edge computing environments, new fault-tolerant approaches must be developed before the benefits of edge computing can be realized. Motivated by algorithm-based fault tolerance, a variant of the asynchronous Jacobi (ASJ) method is developed that achieves resilience to data corruption by rejecting solution approximations from neighbor devices according to a bound derived from convergence theory. Numerical results on a two-dimensional Poisson problem show that the new rejection criterion, along with a novel approximation to the shortest path length on which the criterion depends, restores convergence for the ASJ variant in the presence of certain types data corruption. Numerical results are obtained for when the singular values in the analytic bound are approximated. Additional linear systems are also explored, one with a more dense sparsity pattern and one that includes advection. All results indicate that successful resilience to data corruption depends on whether the bound tightens fast enough to reject corrupted data before the iteration evolution deviates significantly from that predicted by the convergence theory defining the bound. This observation generalizes to future work on algorithm-based fault tolerance for other asynchronous algorithms, including upcoming approaches that leverage Krylov subspaces.

97 MATHEMATICS AND COMPUTING↗

Real-Time Distribution System State Estimation with Asynchronous Measurements

We report state estimation is a fundamental task in power systems. Although distribution systems are increasingly equipped with sensing devices and smart meters, measurements are typically reported at different rates and asynchronously; these aspects pose severe strains on workhorse state estimation algorithms, which are designed to process batches of data collected in a synchronous manner from all the measurement units. In this paper, we develop a novel state estimation algorithm to continuously update the estimate of the state based on measurements received in an asynchronous manner from measurement units. The synthesis of the algorithm hinges on a proximal-point type method, implemented in an online fashion, and capable of processing measurements received sequentially from sensors. A performance analysis is presented by providing bounds on the estimation error in terms of the mean and variance that hold at each iteration and asymptotically. The scheme is also compared with a more traditional Weighted Least Squares estimator that compensates for the lack of measurement data by using, as pseudo measurements, the measurement retrieved during a certain time window. Numerical simulations on the IEEE 37-bus feeder corroborate the analytical findings.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Traveler: Navigating Task Parallel Traces for Performance Analysis

Understanding the behavior of software in execution is a key step in identifying and fixing performance issues. This is especially important in high performance computing contexts where even minor performance tweaks can translate into large savings in terms of computational resource use. To aid performance analysis, developers may collect an execution trace —a chronological log of program activity during execution. As traces represent the full history, developers can discover a wide array of possibly previously unknown performance issues, making them an important artifact for exploratory performance analysis. However, interactive trace visualization is difficult due to issues of data size and complexity of meaning. Traces represent nanosecond-level events across many parallel processes, meaning the collected data is often large and difficult to explore. The rise of asynchronous task parallel programming paradigms complicates the relation between events and their probable cause. Here, to address these challenges, we conduct a continuing design study in collaboration with high performance computing researchers. We develop diverse and hierarchical ways to navigate and represent execution trace data in support of their trace analysis tasks. Through an iterative design process, we developed Traveler , an integrated visualization platform for task parallel traces. Traveler provides multiple linked interfaces to help navigate trace data from multiple contexts. We evaluate the utility of Traveler through feedback from users and a case study, finding that integrating multiple modes of navigation in our design supported performance analysis tasks and led to the discovery of previously unknown behavior in a distributed array library.

97 MATHEMATICS AND COMPUTING↗

Scalable Pattern Matching in Metadata Graphs via Constraint Checking

Pattern matching is a fundamental tool for answering complex graph queries. Unfortunately, existing solutions have limited capabilities: They do not scale to process large graphs and/or support only a restricted set of search templates or usage scenarios. Moreover, the algorithms at the core of the existing techniques are not suitable for today’s graph processing infrastructures relying on horizontal scalability and shared-nothing clusters, as most of these algorithms are inherently sequential and difficult to parallelize. In this article we present an algorithmic pipeline that bases pattern matching on constraint checking. The key intuition is that each vertex and edge participating in a match has to meet a set of constraints implicitly specified by the search template. These constraints can be verified independently and typically are less expensive to compute than searching the full template. The pipeline we propose generates these constraints and iterates over them to eliminate all the vertices and edges that do not participate in any match, thus reducing the background graph to a subgraph that is the union of all template matches—the complete set of all vertices and edges that participate in at least one match. Additional analysis can be performed on this annotated, reduced graph, such as full match enumeration, match counting, or computing vertex/edge centrality. Furthermore, a vertex-centric formulation for constraint checking algorithms exists, and this makes it possible to harness existing high-performance, vertex-centric graph processing frameworks. This technique (i) enables highly scalable pattern matching in metadata (labeled) graphs; (ii) supports arbitrary patterns with 100% precision; (iii) enables tradeoffs between precision and time-to-solution, while always selects all vertices and edges that participate in matches, thus offering 100% recall; and (iv) supports a set of popular data analytics scenarios. We implement our approach on top of HavoqGT, an open-source asynchronous graph processing framework, and demonstrate its advantages through strong and weak scaling experiments on massive scale real-world (up to 257 billion edges) and synthetic (up to 4.4 trillion edges) labeled graphs, respectively, and at scales (1,024 nodes / 36,864 cores), orders of magnitude larger than used in the past for similar problems. This article serves two purposes: First, it synthesises the knowledge accumulated during a long-term project. Second, it presents new system features, usage scenarios, optimizations, and comparisons with related work that strengthen the confidence that pattern matching based on iterative pruning via constraint checking is an effective and scalable approach in practice. The new contributions include the following: (i) We demonstrate the ability of the constraint checking approach to efficiently support two additional search scenarios that often emerge in practice, interactive incremental search and exploratory search. (ii) We empirically compare our solution with two additional state-of-the-art systems, Arabsque and TriAD. (iii) We show the ability of our solution to accommodate a more diverse range of datasets with varying properties, e.g., scale, skewness, label distribution, and match frequency. (iv) We introduce or extend a number of system features (e.g., work aggregation, load balancing, and the ability to cap the generated traffic) and design optimizations and demonstrate their advantages with respect to improving performance and scalability. (v) We present bottleneck analysis and insights into artifacts that influence performance. (vi) We present a theoretical complexity argument that motivates the performance gains we observe.

97 MATHEMATICS AND COMPUTING↗

Non-disruptive error field identification based on magnetic island healing

Here a technique to identify intrinsic error fields (EFs) in tokamaks with minimized risk of disruption is demonstrated on the DIII-D tokamak. The method extends the conventional driven magnetic island ‘compass scan’ approach by modifying asynchronous control waveforms to enable prompt healing of the island instability. Healing of the island is achieved by reducing the imposed non-axisymmetric coil current and raising the density (here via gas fueling). The method is also shown to support multiple island threshold measurements per pulse, thus reducing the number of dedicated pulses necessary to conduct an EF identification. Non-linear modeling with the TM1 code reproduces the experimental results and approximately recovers the critical density required for island healing. Island healing is explained in the non-linear modeling by an increase in the viscous coupling between the static island and the nearby flowing plasma, thus healing the island as it accelerates into the plasma frame. Due to both simplicity and risk minimization, this technique is suitable for plasma-based EF identification in the early commissioning stages of future disruption-averse tokamaks such as ITER and SPARC.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗