Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel time integration”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Parallel Multi-Step/Multi-Rate Integration of Two-Time Scale Dynamic Systems

Increasing demands on the fidelity of simulations for real-time and high-fidelity simulations are stressing the capacity of modern processors. New integration techniques are required that provide maximum efficiency for systems that are parallelizable. However many current techniques make assumptions that are at odds with non-cascadable systems. A new serial multi-step/multi-rate integration algorithm for dual-timescale continuous state systems is presented which applies to these systems, and is extended to a parallel multi-step/multi-rate algorithm. The superior performance of both algorithms is demonstrated through a representative example.

dynamics↗

Space-Time Block Preconditioning for Incompressible Flow

Parallel-in-time methods have become increasingly popular in the simulation of time-dependent numerical PDEs, allowing for the efficient use of additional message passing interface processes when spatial parallelism saturates. Most methods treat the solution and parallelism in space and time separately. In contrast, all-at-once methods solve the full space-time system directly, largely treating time as simply another spatial dimension. All-at-once methods offer a number of benefits over separate treatment of space and time, most notably significantly increased parallelism and faster time to solution (when applicable). However, the development of fast, scalable all-at-once methods has largely been limited to time-dependent (advection-)diffusion problems. This paper introduces the concept of space-time block preconditioning for the all-at-once solution of incompressible flow. By extending well-known concepts of spatial block preconditioning to the space-time setting, we develop a block preconditioner whose application requires the solution of a space-time (advection-)diffusion equation in the velocity block, coupled with a pressure Schur complement approximation consisting of independent spatial solves at each time-step, and a space-time matrix-vector multiplication. The new method is tested on four classical models in incompressible flow. Finally, the results indicate perfect scalability in refinement of spatial and temporal mesh spacing, perfect scalability in nonlinear Picard iteration count when applied to a nonlinear Navier--Stokes problem, and minimal overhead in terms of number of preconditioner applications compared with sequential time-stepping.

97 MATHEMATICS AND COMPUTING↗

Asynchronous Truncated Multigrid-Reduction-in-Time

In this paper, we present the new “asynchronous truncated multigrid-reduction-in-time” (AT-MGRIT) algorithm for introducing time parallelism to the solution of discretized time-dependent problems. The new algorithm is based on the multigrid-reduction-in-time (MGRIT) approach, which, in certain settings, is equivalent to another common multilevel parallel-in-time method, Parareal. In contrast to Parareal and MGRIT that both consider a global temporal grid over the entire time interval on the coarsest level, the AT-MGRIT algorithm uses truncated local time grids on the coarsest level, each grid covering certain temporal subintervals. Further, these local grids can be solved completely in an independent way from each other, which reduces the sequential part of the algorithm and, thus, increases parallelism in the method. Here, we study the effect of using truncated local coarse grids on the convergence of the algorithm, both theoretically and numerically, and show, using challenging nonlinear problems, that the new algorithm consistently outperforms classical Parareal/MGRIT in terms of time to solution.

97 MATHEMATICS AND COMPUTING↗

LaRC computational dynamics overview

Present research centers on the development of advanced computational methods for transient simulation analyses. Aircraft, launch vehicles and space structure components are potential applications, but primary focus is presently on large space structures. There are both in-house and out-of-house activities. The in-house activity centers around the development of a multibody simulation tool for truss-like structures called LATDYN for Large Angle Transient DYNamics. Multibody analysis involves articulation of structural components as well as robotic maneuvers. These items are necessary for construction (erection or deployment) of large space structures in orbit and the carrying out of certain operations on board the space station. Thus, part of the in-house activity involves the development of methods which treat the changing mass, stiffness and constraints associated with articulating systems. The out-of-house activity involves subcycling, development of large deformation/motion beam formulation, constraint stabilization and direct time integration transient algorithms in parallel computing.

Husner, J. M.↗

PETSc/TAO Users Manual (Rev. 3.19)

This manual describes the use of the Portable, Extensible Toolkit for Scientific Computation (PETSc) and the Toolkit for Advanced Optimization (TAO) for the numerical solution of partial differential equations and related problems on high-performance computers. PETSc/TAO is a suite of data structures and routines that provide the building blocks for the implementation of large-scale application codes on parallel (and serial) computers. PETSc uses the MPI standard for all distributed memory communication. PETSc/TAO includes a large suite of parallel linear solvers, nonlinear solvers, time integrators, and opti mization that may be used in application codes written in Fortran, C, C++, and Python (via petsc4py; see Getting Started). PETSc provides many of the mechanisms needed within parallel application codes, such as parallel matrix and vector assembly routines. The library is organized hierarchically, enabling users to employ the level of abstraction that is most appropriate for a particular problem. By using techniques of object-oriented programming, PETSc provides enormous flexibility for users. PETSc is a sophisticated set of software tools; as such, for some users it initially has a much steeper learning curve than packages such as MATLAB or a simple subroutine library. In particular, for individuals without some computer science background, experience programming in C, C++, python, or Fortran and experience using a debugger such as gdb or lldb, it may require a significant amount of time to take full advantage of the features that enable efficient software use. However, the power of the PETSc design and the algorithms it incorporates may make the efficient implementation of many application codes simpler than “rolling them” yourself. For many tasks a package such as MATLAB is often the best tool; PETSc is not intended for the classes of problems for which effective MATLAB code can be written. There are several packages, built on PETSc, that may satisfy your needs without requiring directly using PETSc. We recommend reviewing these packages functionality before starting to code directly with PETSc. PETSc can be used to provide a “MPI parallel linear solver” in an otherwise sequential, or OpenMP parallel code. This approach cannot provide extremely large improvements in the application time by utilizing large numbers of MPI processes but can still improve the performance. Certainly all parts of a previously sequential code need not be parallelized but the matrix generation portion must be parallelized to expect true scalability to large numbers of MPI processes. See PCMPI for details on how to utilize the PETSc MPI linear solver server. Since PETSc is under continued development, small changes in usage and calling sequences of routines will occur. PETSc has been supported for twenty-five years; see mailing list information on our website for information on contacting support.

97 MATHEMATICS AND COMPUTING↗

The velocity correlation function in cosmic-ray diffusion theory

It is shown that Earl's (1973) eigenvalue sum for the cosmic-ray spatial diffusion coefficient parallel to the mean magnetic field is precisely equivalent to the time integral of the particle-velocity correlation function parallel to the mean field. A derivation due to Kubo (1957) is applied to cosmic-ray pitch-angle scattering, and it is proven that all nine components of the cosmic-ray diffusion tensor can be expressed as integrals over the velocity correlation function. A pitch-angle correlation function is derived, and the effect of long-wavelength turbulence on the velocity correlation function and spatial diffusion coefficients is examined. Application of the velocity-correlation method to a realistic case involving both pitch-angle scattering and appreciable fluctuation in the direction of the local field indicates that long-wavelength turbulence in the local field reduces the parallel diffusion coefficient and places an upper limit on the ratio of the perpendicular to parallel diffusion coefficients.

Forman, M. A.↗

Photochemically Induced Acousto-optics Fluid Simulations

PIAFS is a finite-difference code to solve the compressible Navier-Stokes equations with chemical heating on Cartesian grids. It models chemical reactions of air (oxygen and carbon dioxide) with ozone subject to radiation. It uses a high-order WENO spatial discretization and explicit Runge-Kutta time integration. It is capable of parallel simulations using MPI. The code is written in C/C++.

Oudin, AlbertineN [Lawrence Livermore National Lab↗

PETSc/TAO Users Manual V.3.21

This manual describes the use of the Portable, Extensible Toolkit for Scientific Computation (PETSc) and the Toolkit for Advanced Optimization (TAO) for the numerical solution of partial differential equations (PDEs) and related problems on high-performance computers. PETSc/TAO is a suite of data structures and routines that provide the building blocks for implementing large-scale application codes on parallel (and serial) computers. PETSc uses the MPI standard for all distributed memory communication. PETSc/TAO includes a large suite of parallel linear solvers, nonlinear solvers, time integrators, and optimizers that may be used in application codes written in Fortran, C, C++, and Python (via petsc4py; see Getting Started ). The library is organized hierarchically, enabling users to employ the abstraction level most appropriate for a particular problem. By using techniques of object-oriented programming, PETSc provides enormous flexibility for users.

97 MATHEMATICS AND COMPUTING↗

A time-parallel method for scalable heat transfer simulations of additive manufacturing

Here, a major challenge in simulating the thermal behavior in additive manufacturing processes is the disparate length and time scales between transport phenomena occurring in the melt pool and the component. A common simulation approach relies on spatial decomposition for parallel computing, but due to the nature of heat transfer in AM, where most of the computational expenditure is localized near the melt pool, the computational speedup from spatial parallelization saturates quickly. Therefore, additional parallelism by means of time-domain decomposition is needed to fully take advantage of high-performance computing (HPC) resources. This work introduces a time-parallel method to improve the computational scalability of additive manufacturing simulations on HPC systems, while maintaining high temporal resolution of heat transfer near the melt pool. The method, inspired by the nonlinear paraexp formalism, performs an iterative superposition of nonlinear solutions to the initial value problem, integrating the heat equation across overlapping time-parallel intervals. For a single layer of the NIST AMB2018–01 L7 benchmark problem, the method achieves a 38.51x speedup in wall-clock time with a maximum error in the global temperature solution of 0.99%. This reduces the total solution time from 196.72 min to 5.11 min on 128 nodes of the ORNL Frontier supercomputer. The tradeoff between accuracy and total wall-clock time is investigated and recommendations for time-parallel deployment for AM problems are made.

Additive manufacturing↗

TorchBraid: High-Performance Layer-Parallel Training of Deep Neural Networks with MPI and GPU Acceleration

TorchBraid is a high-performance implementation of layer-parallel training for deep neural networks (DNNs) supporting MPI-based parallelism and GPU acceleration. Layer-parallel training has been developed to overcome the serialization inherent in forward and backward propagation of DNNs that limits utilization of computational resources in the strong scaling limit. To achieve this, TorchBraid integrates the PyTorch neural network framework with the state-of-the-art XBraid time-parallel library. Furthermore, this article presents the use and performance of TorchBraid, in addition to solutions for overcoming the algorithmic challenges inherent in combining automatic differentiation with layer-parallel. Results are presented with and without GPU acceleration for the Tiny ImageNet and MNIST image classification data sets, as well as recurrent neural networks. Overall, TorchBraid enables fast training of DNNs, both in a strong and weak scaling context. In addition to the TorchBraid software, several new advances in applying layer-parallel algorithms are detailed. Integration of layer-parallel with data-parallel algorithms is presented for the first time, showing the computational advantages of the combination. Standard deep learning techniques, like batch-normalization, are developed for layer-parallel training. Finally, a new approach combining layer-parallel with spatial coarsening in order to accelerate training for 3D image classification shows roughly a 10× speedup over serial execution.

Layer-parallel↗

Fast Multigrid Reduction-in-Time for Advection via Modified Semi-Lagrangian Coarse-Grid Operators

Many iterative parallel-in-time algorithms have been shown to be highly efficient for diffusion-dominated partial differential equations (PDEs) but are inefficient or even divergent when applied to advection-dominated PDEs. We consider the application of the multigrid reduction-in-time (MGRIT) algorithm to linear advection PDEs. Here, the key to efficient time integration with this method is using a coarse-grid operator that provides a sufficiently accurate approximation to the so-called ideal coarse-grid operator. For certain classes of semi-Lagrangian discretizations, we present a novel semi-Lagrangian-based coarse-grid operator that leads to fast and scalable multilevel time integration of linear advection PDEs. The coarse-grid operator is composed of a semi-Lagrangian discretization followed by a correction term, with the correction designed so that the leading-order truncation error of the composite operator is approximately equal to that of the ideal coarse-grid operator. Parallel results show substantial speed-ups over sequential time integration for variable-wave-speed advection problems in one and two spatial dimensions, and using high-order discretizations up to order five. The proposed approach establishes the first practical method that provides small and scalable MGRIT iteration counts for advection problems.

97 MATHEMATICS AND COMPUTING↗

Development of the US3D Code for Advanced Compressible and Reacting Flow Simulations

Aerothermodynamics and hypersonic flows involve complex multi-disciplinary physics, including finite-rate gas-phase kinetics, finite-rate internal energy relaxation, gas-surface interactions with finite-rate oxidation and sublimation, transition to turbulence, large-scale unsteadiness, shock-boundary layer interactions, fluid-structure interactions, and thermal protection system ablation and thermal response. Many of the flows have a large range of length and time scales, requiring large computational grids, implicit time integration, and large solution run times. The University of Minnesota NASA US3D code was designed for the simulation of these complex, highly-coupled flows. It has many of the features of the well-established DPLR code, but uses unstructured grids and has many advanced numerical capabilities and physical models for multi-physics problems. The main capabilities of the code are described, the physical modeling approaches are discussed, the different types of numerical flux functions and time integration approaches are outlined, and the parallelization strategy is overviewed. Comparisons between US3D and the NASA DPLR code are presented, and several advanced simulations are presented to illustrate some of novel features of the code.

CFD↗

On the spectral stability of time integration algorithms for a class of constrained dynamics problems

Incomplete field formulations have recently been the subject of intense research because of their potential in coupled analysis of independently modeled substructures, adaptive refinement, domain decomposition, and parallel processing. This paper discusses the design and analysis of time-integration algorithms for these formulations and emphasizes the treatment of their inter-subdomain constraint equations. These constraints are shown to introduce a destabilizing effect in the dynamic system that can be analyzed by investigating the behavior of the time-integration algorithm at infinite and zero frequencies. Three different approaches for constructing penalty-free unconditionally stable second-order accurate solution procedures for this class of hybrid formulations are presented, discussed and illustrated with numerical examples. The theoretical results presented in this paper also apply to a large family of nonlinear multibody dynamics formulations. Some of the algorithms outlined herein are important alternatives to the popular technique consisting of transforming differential/algebraic equations into ordinary differential equations via the introduction of a stabilization term that depends on arbitrary constants and that influences the computed so1ution.

Farhat, Charbel↗

Execution environment for intelligent real-time control systems

Modern telerobot control technology requires the integration of symbolic and non-symbolic programming techniques, different models of parallel computations, and various programming paradigms. The Multigraph Architecture, which has been developed for the implementation of intelligent real-time control systems is described. The layered architecture includes specific computational models, integrated execution environment and various high-level tools. A special feature of the architecture is the tight coupling between the symbolic and non-symbolic computations. It supports not only a data interface, but also the integration of the control structures in a parallel computing environment.

Sztipanovits, Janos↗

Performance Analysis of Speculative Parallel Adaptive Local Timestepping for Conservation Laws

Stable simulation of conservation laws, such as those used to model fluid dynamics and plasma physics applications, requires the satisfaction of the so-called Courant-Friedrichs-Lewy condition. By allowing regions of the mesh to advance with different timesteps that locally satisfy this stability constraint, significant work reduction can be attained when compared to a time integration scheme using a single timestep size. However, parallelizing this algorithm presents considerable difficulty. Since the stability condition depends on the state of the system, dependencies become dynamic and potentially non-local. In this article, we present an adaptive local timestepping algorithm using an optimistic (Timewarp-based) parallel discrete event simulation. We introduce waiting heuristics to limit misspeculation and a semi-static load balancing scheme to eliminate load imbalance as parts of the mesh require finer or coarser timesteps. Last, we outline an interface for separating the physics of the specific conservation law from the temporal integration allowing for productive adoption of our proposed algorithm. We present a misspeculation study for three conservation laws, demonstrating both the productivity of the local timestepping API, for which 74% of the lines of code are reused across different conservation laws, and the robustness of the waiting heuristics—at most 1.5% of element updates are rolled back. Our performance studies demonstrate up to a 2.8× speedup versus a baseline unoptimized local timestepping approach, a 4x improvement in per-node throughput compared to an MPI parallelization of synchronous timestepping, and scalability up to 3,072 cores on NERSC’s Cori Haswell partition.

97 MATHEMATICS AND COMPUTING↗

NASA Space Launch System Operations Outlook

The National Aeronautics and Space Administration's (NASA) Space Launch System (SLS) Program, managed at the Marshall Space Flight Center (MSFC), is working with the Ground Systems Development and Operations (GSDO) Program, based at the Kennedy Space Center (KSC), to deliver a new safe, affordable, and sustainable capability for human and scientific exploration beyond Earth's orbit (BEO). Larger than the Saturn V Moon rocket, SLS will provide 10 percent more thrust at liftoff in its initial 70 metric ton (t) configuration and 20 percent more in its evolved 130-t configuration. The primary mission of the SLS rocket will be to launch astronauts to deep space destinations in the Orion Multi-Purpose Crew Vehicle (MPCV), also in development and managed by the Johnson Space Center. Several high-priority science missions also may benefit from the increased payload volume and reduced trip times offered by this powerful, versatile rocket. Reducing the life-cycle costs for NASA's space transportation flagship will maximize the exploration and scientific discovery returned from the taxpayer's investment. To that end, decisions made during development of SLS and associated systems will impact the nation's space exploration capabilities for decades. This paper will provide an update to the operations strategy presented at SpaceOps 2012. It will focus on: 1) Preparations to streamline the processing flow and infrastructure needed to produce and launch the world's largest rocket (i.e., through incorporation and modification of proven, heritage systems into the vehicle and ground systems); 2) Implementation of a lean approach to reachback support of hardware manufacturing, green-run testing, and launch site processing and activities; and 3) Partnering between the vehicle design and operations communities on state-ofthe- art predictive operations analysis techniques. An example of innovation is testing the integrated vehicle at the processing facility in parallel, rather than sequentially, saving both time and money. These themes are accomplished under the context of a new cross-program integration model that emphasizes peer-to-peer accountability and collaboration towards a common, shared goal. Utilizing the lessons learned through 50 years of human space flight experience, SLS is assigning the right number of people from appropriate backgrounds, providing them the right tools, and exercising the right processes for the job. The result will be a powerful, versatile, and capable heavy-lift, human-rated asset for the future human and scientific exploration of space.

Hefner, William Keith↗

NASA Space Launch System Operations Outlook

The National Aeronautics and Space Administration's (NASA) Space Launch System (SLS) Program, managed at the Marshall Space Flight Center (MSFC), is working with the Ground Systems Development and Operations (GSDO) Program, based at the Kennedy Space Center (KSC), to deliver a new safe, affordable, and sustainable capability for human and scientific exploration beyond Earth's orbit (BEO). Larger than the Saturn V Moon rocket, SLS will provide 10 percent more thrust at liftoff in its initial 70 metric ton (t) configuration and 20 percent more in its evolved 130-t configuration. The primary mission of the SLS rocket will be to launch astronauts to deep space destinations in the Orion Multi- Purpose Crew Vehicle (MPCV), also in development and managed by the Johnson Space Center. Several high-priority science missions also may benefit from the increased payload volume and reduced trip times offered by this powerful, versatile rocket. Reducing the lifecycle costs for NASA's space transportation flagship will maximize the exploration and scientific discovery returned from the taxpayer's investment. To that end, decisions made during development of SLS and associated systems will impact the nation's space exploration capabilities for decades. This paper will provide an update to the operations strategy presented at SpaceOps 2012. It will focus on: 1) Preparations to streamline the processing flow and infrastructure needed to produce and launch the world's largest rocket (i.e., through incorporation and modification of proven, heritage systems into the vehicle and ground systems); 2) Implementation of a lean approach to reach-back support of hardware manufacturing, green-run testing, and launch site processing and activities; and 3) Partnering between the vehicle design and operations communities on state-of-the-art predictive operations analysis techniques. An example of innovation is testing the integrated vehicle at the processing facility in parallel, rather than sequentially, saving both time and money. These themes are accomplished under the context of a new cross-program integration model that emphasizes peer-to-peer accountability and collaboration towards a common, shared goal. Utilizing the lessons learned through 50 years of human space flight experience, SLS is assigning the right number of people from appropriate backgrounds, providing them the right tools, and exercising the right processes for the job. The result will be a powerful, versatile, and capable heavy-lift, human-rated asset for the future human and scientific exploration of space.

Hefner, William Keith↗

A transient FETI methodology for large-scale parallel implicit computations in structural mechanics

Explicit codes are often used to simulate the nonlinear dynamics of large-scale structural systems, even for low frequency response, because the storage and CPU requirements entailed by the repeated factorizations traditionally found in implicit codes rapidly overwhelm the available computing resources. With the advent of parallel processing, this trend is accelerating because explicit schemes are also easier to parallelize than implicit ones. However, the time step restriction imposed by the Courant stability condition on all explicit schemes cannot yet -- and perhaps will never -- be offset by the speed of parallel hardware. Therefore, it is essential to develop efficient and robust alternatives to direct methods that are also amenable to massively parallel processing because implicit codes using unconditionally stable time-integration algorithms are computationally more efficient when simulating low-frequency dynamics. Here we present a domain decomposition method for implicit schemes that requires significantly less storage than factorization algorithms, that is several times faster than other popular direct and iterative methods, that can be easily implemented on both shared and local memory parallel processors, and that is both computationally and communication-wise efficient. The proposed transient domain decomposition method is an extension of the method of Finite Element Tearing and Interconnecting (FETI) developed by Farhat and Roux for the solution of static problems. Serial and parallel performance results on the CRAY Y-MP/8 and the iPSC-860/128 systems are reported and analyzed for realistic structural dynamics problems. These results establish the superiority of the FETI method over both the serial/parallel conjugate gradient algorithm with diagonal scaling and the serial/parallel direct method, and contrast the computational power of the iPSC-860/128 parallel processor with that of the CRAY Y-MP/8 system.

Farhat, Charbel↗