Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Asynchronous iterations”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

26 records · Page 2

Performance Optimization Methods for a Memory-Bound, Unstructured-Grid CFD Application on Massively Parallel GPU Platforms

Computational performance of the FUN3D unstructured-grid computational fluid dynamics (CFD) application on massively parallel GPU environments is memory-bound and highly dependent upon efficient reads from and atomic updates to the irregular cell-, edge-, and node-based data structures. In this talk, we present recent efforts into optimizing select performance-critical kernels on NVIDIA Tesla V100 and A100 GPUs and AMD CDNA MI100 GPUs. A novel use of L2 cache residency controls and asynchronous loads into on-chip shared memory are explored on the A100 GPU for the sparse iterative solver, which is dominated by mixed-precision, sparse matrix vector multiplication. Demonstrations show that these methods improve global memory bandwidth utilization by 13.5% on the A100 GPU. Several techniques are also presented that use registers and/or shared memory to facilitate array transposition and aggregation which combine to reduce the frequency and increase the cache efficiency of floating-point atomic updates to the irregular data structures. These methods are demonstrated to improve the kernel throughput by nearly 500% on select kernels on the AMD MI100 over atomic updates directly to global memory. Overall, both V100 and A100 GPUs outperformed the MI100 GPU on kernels dominated by double-precision atomic updates; however, the techniques demonstrated here reduced the performance gap and improved the MI100 performance.

GPU CPU unstructured CFD memory↗

tomoCAM : fast model-based iterative reconstruction via GPU acceleration and non-uniform fast Fourier transforms

X-ray-based computed tomography is a well established technique for determining the three-dimensional structure of an object from its two-dimensional projections. In the past few decades, there have been significant advancements in the brightness and detector technology of tomography instruments at synchrotron sources. These advancements have led to the emergence of new observations and discoveries, with improved capabilities such as faster frame rates, larger fields of view, higher resolution and higher dimensionality. These advancements have enabled the material science community to expand the scope of tomographic measurements towards increasingly in situ and in operando measurements. In these new experiments, samples can be rapidly evolving, have complex geometries and restrictions on the field of view, limiting the number of projections that can be collected. In such cases, standard filtered back-projection often results in poor quality reconstructions. Iterative reconstruction algorithms, such as model-based iterative reconstructions (MBIR), have demonstrated considerable success in producing high-quality reconstructions under such restrictions, but typically require high-performance computing resources with hundreds of compute nodes to solve the problem in a reasonable time. Here, tomoCAM , is introduced, a new GPU-accelerated implementation of model-based iterative reconstruction that leverages non-uniform fast Fourier transforms to efficiently compute Radon and back-projection operators and asynchronous memory transfers to maximize the throughput to the GPU memory. The resulting code is significantly faster than traditional MBIR codes and delivers the reconstructive improvement offered by MBIR with affordable computing time and resources. tomoCAM has a Python front-end, allowing access from Jupyter -based frameworks, providing straightforward integration into existing workflows at synchrotron facilities.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Project Integration Architecture: Distributed Lock Management, Deadlock Detection, and Set Iteration

The migration of the Project Integration Architecture (PIA) to the distributed object environment of the Common Object Request Broker Architecture (CORBA) brings with it the nearly unavoidable requirements of multiaccessor, asynchronous operations. In order to maintain the integrity of data structures in such an environment, it is necessary to provide a locking mechanism capable of protecting the complex operations typical of the PIA architecture. This paper reports on the implementation of a locking mechanism to treat that need. Additionally, the ancillary features necessary to make the distributed lock mechanism work are discussed.

Jones, William Henry↗

Modifying the Asynchronous Jacobi Method for Data Corruption Resilience

Moving scientific computation from high-performance computing (HPC) and cloud computing (CC) environments to devices on the edge, i.e., physically near instruments of interest, has received tremendous interest in recent years. Such edge computing environments can operate on data in situ, offering enticing benefits over data aggregation to HPC and CC facilities that include avoiding costs of transmission, increased data privacy, and real-time data analysis. Because of the inherent unreliability of edge computing environments, new fault-tolerant approaches must be developed before the benefits of edge computing can be realized. Motivated by algorithm-based fault tolerance, a variant of the asynchronous Jacobi (ASJ) method is developed that achieves resilience to data corruption by rejecting solution approximations from neighbor devices according to a bound derived from convergence theory. Numerical results on a two-dimensional Poisson problem show that the new rejection criterion, along with a novel approximation to the shortest path length on which the criterion depends, restores convergence for the ASJ variant in the presence of certain types data corruption. Numerical results are obtained for when the singular values in the analytic bound are approximated. Additional linear systems are also explored, one with a more dense sparsity pattern and one that includes advection. All results indicate that successful resilience to data corruption depends on whether the bound tightens fast enough to reject corrupted data before the iteration evolution deviates significantly from that predicted by the convergence theory defining the bound. This observation generalizes to future work on algorithm-based fault tolerance for other asynchronous algorithms, including upcoming approaches that leverage Krylov subspaces.

97 MATHEMATICS AND COMPUTING↗

A Parallel Particle Swarm Optimization Algorithm Accelerated by Asynchronous Evaluations

A parallel Particle Swarm Optimization (PSO) algorithm is presented. Particle swarm optimization is a fairly recent addition to the family of non-gradient based, probabilistic search algorithms that is based on a simplified social model and is closely tied to swarming theory. Although PSO algorithms present several attractive properties to the designer, they are plagued by high computational cost as measured by elapsed time. One approach to reduce the elapsed time is to make use of coarse-grained parallelization to evaluate the design points. Previous parallel PSO algorithms were mostly implemented in a synchronous manner, where all design points within a design iteration are evaluated before the next iteration is started. This approach leads to poor parallel speedup in cases where a heterogeneous parallel environment is used and/or where the analysis time depends on the design point being analyzed. This paper introduces an asynchronous parallel PSO algorithm that greatly improves the parallel e ciency. The asynchronous algorithm is benchmarked on a cluster assembled of Apple Macintosh G5 desktop computers, using the multi-disciplinary optimization of a typical transport aircraft wing as an example.

Venter, Gerhard↗

Traveler: Navigating Task Parallel Traces for Performance Analysis

Understanding the behavior of software in execution is a key step in identifying and fixing performance issues. This is especially important in high performance computing contexts where even minor performance tweaks can translate into large savings in terms of computational resource use. To aid performance analysis, developers may collect an execution trace —a chronological log of program activity during execution. As traces represent the full history, developers can discover a wide array of possibly previously unknown performance issues, making them an important artifact for exploratory performance analysis. However, interactive trace visualization is difficult due to issues of data size and complexity of meaning. Traces represent nanosecond-level events across many parallel processes, meaning the collected data is often large and difficult to explore. The rise of asynchronous task parallel programming paradigms complicates the relation between events and their probable cause. Here, to address these challenges, we conduct a continuing design study in collaboration with high performance computing researchers. We develop diverse and hierarchical ways to navigate and represent execution trace data in support of their trace analysis tasks. Through an iterative design process, we developed Traveler , an integrated visualization platform for task parallel traces. Traveler provides multiple linked interfaces to help navigate trace data from multiple contexts. We evaluate the utility of Traveler through feedback from users and a case study, finding that integrating multiple modes of navigation in our design supported performance analysis tasks and led to the discovery of previously unknown behavior in a distributed array library.

97 MATHEMATICS AND COMPUTING↗

Real-Time Science Decisioning During High Tempo-High Intensity Mission Operations and the Role of Analogs

Introduction: NASA’s VIPER mission presents a unique operational paradigm within the history of robotic spaceflight. The proximity of the Moon to the Earth and the terrain elements (surface characteristics, light/shadow dynamics, communication links) of the lunar South Polar landing site create unprecedented operational conditions between these two planetary bodies. Apollo era lunar science and exploration included humans in situ to operate instruments and assimilate observational inputs in real-time. Previous lunar orbital missions have worked to operational timescales, e.g., decisional timelines and communication exchanges, that were weeks in length. Mars rover missions have worked to operational timescales, e.g., decisional timelines and communication exchanges between Mars and Earth, that were hours, days, and weeks in length. In the case of the VIPER mission, our operational decisioning for rover driving and instrument commanding will be compressed to minute-scale timeframes. These operational conditions directly impact the manner and speed with which the VIPER Science Team (VST) is required to synthesize and analyze data and produce timely science-driven decisions throughout surface mission operations. The VST shall provide mission enhancing scientific input to guide rover traverse planning and drill site confirmation and selection throughout surface operations. Further, the VST input will be of vital importance to the mission’s ability to maximize science return and to meet broader NASA objectives for future lunar in-situ resource utilization (ISRU)and exploration activities. The VST co-located in the Mission Science Center (MSC) will be responsive to the tactical operational cadence of the Mission Operations Center (MOC) and will provide further strategic and Long-Term Planning (LTP) guidance to the mission. The VIPER Science Operations & Integration(SO&I)team has developed an architecture that is focused on the infusion of science-decisioning into the operational framework and execution cadence of VIPER. NASA analog research has played a significant role in the construction of the VIPER science operations systems. As an example, the SO&I team has led analog missions that have focused on bringing together expertise in the sciences (natural, applied and social) and in operations in service of learning how to build and hold together interdisciplinary work environments and what tools are needed to support high tempo, high intensity integrated decisioning. These experiences have provided an essential foundation of knowledge to the VIPER team. Those analogs that specifically influenced the VIPER science operations construct were identified through a process of comparative analysis to prioritize those that offered relevance in whole or in part, and those that did not. The analog research output that provided extensibility to the VIPER science operations architecture included remote teams of humans and robots in cooperation (synchronous and asynchronous) with simulated earthbound systems, engineering and science teams, and the integrated assembly of tools that supported scientific analysis and data synthesis and provided infrastructure for the remote testing framework. Analogs which included real-time data monitoring, synthesis, visualization and access in a democratized and operationalized manner were of particular interest to the development of the VIPER MSC toolset both in terms of the technology and the processes used to develop the supporting infrastructure. We anticipate that each subsequent mission to the lunar south pole, whether with robots or humans, will be able to optimize science and exploration return by evolving strategies to infuse real-time collaborative science-decisioning. Furthermore, these efforts will result in a foundation for science operations development in support of human-robotic exploration of deep space and Mars. NASA analogs can continue to provide the opportunity to prepare, test and iterate on the operational concepts and tools that will support these ever-expanding space exploration efforts. Our presentation will include an overview of the VIPER Science Operations & Integration development process and specifics on what aspects of analog research have had a significant impact on our work systems.

D S S Lim↗

Non-disruptive error field identification based on magnetic island healing

Here a technique to identify intrinsic error fields (EFs) in tokamaks with minimized risk of disruption is demonstrated on the DIII-D tokamak. The method extends the conventional driven magnetic island ‘compass scan’ approach by modifying asynchronous control waveforms to enable prompt healing of the island instability. Healing of the island is achieved by reducing the imposed non-axisymmetric coil current and raising the density (here via gas fueling). The method is also shown to support multiple island threshold measurements per pulse, thus reducing the number of dedicated pulses necessary to conduct an EF identification. Non-linear modeling with the TM1 code reproduces the experimental results and approximately recovers the critical density required for island healing. Island healing is explained in the non-linear modeling by an increase in the viscous coupling between the static island and the nearby flowing plasma, thus healing the island as it accelerates into the plasma frame. Due to both simplicity and risk minimization, this technique is suitable for plasma-based EF identification in the early commissioning stages of future disruption-averse tokamaks such as ITER and SPARC.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗