Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel time integration”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

NASA Space Launch System Operations Outlook

The National Aeronautics and Space Administration's (NASA) Space Launch System (SLS) Program, managed at the Marshall Space Flight Center (MSFC), is working with the Ground Systems Development and Operations (GSDO) Program, based at the Kennedy Space Center (KSC), to deliver a new safe, affordable, and sustainable capability for human and scientific exploration beyond Earth's orbit (BEO). Larger than the Saturn V Moon rocket, SLS will provide 10 percent more thrust at liftoff in its initial 70 metric ton (t) configuration and 20 percent more in its evolved 130-t configuration. The primary mission of the SLS rocket will be to launch astronauts to deep space destinations in the Orion Multi- Purpose Crew Vehicle (MPCV), also in development and managed by the Johnson Space Center. Several high-priority science missions also may benefit from the increased payload volume and reduced trip times offered by this powerful, versatile rocket. Reducing the lifecycle costs for NASA's space transportation flagship will maximize the exploration and scientific discovery returned from the taxpayer's investment. To that end, decisions made during development of SLS and associated systems will impact the nation's space exploration capabilities for decades. This paper will provide an update to the operations strategy presented at SpaceOps 2012. It will focus on: 1) Preparations to streamline the processing flow and infrastructure needed to produce and launch the world's largest rocket (i.e., through incorporation and modification of proven, heritage systems into the vehicle and ground systems); 2) Implementation of a lean approach to reach-back support of hardware manufacturing, green-run testing, and launch site processing and activities; and 3) Partnering between the vehicle design and operations communities on state-of-the-art predictive operations analysis techniques. An example of innovation is testing the integrated vehicle at the processing facility in parallel, rather than sequentially, saving both time and money. These themes are accomplished under the context of a new cross-program integration model that emphasizes peer-to-peer accountability and collaboration towards a common, shared goal. Utilizing the lessons learned through 50 years of human space flight experience, SLS is assigning the right number of people from appropriate backgrounds, providing them the right tools, and exercising the right processes for the job. The result will be a powerful, versatile, and capable heavy-lift, human-rated asset for the future human and scientific exploration of space.

Hefner, William Keith↗

A transient FETI methodology for large-scale parallel implicit computations in structural mechanics

Explicit codes are often used to simulate the nonlinear dynamics of large-scale structural systems, even for low frequency response, because the storage and CPU requirements entailed by the repeated factorizations traditionally found in implicit codes rapidly overwhelm the available computing resources. With the advent of parallel processing, this trend is accelerating because explicit schemes are also easier to parallelize than implicit ones. However, the time step restriction imposed by the Courant stability condition on all explicit schemes cannot yet -- and perhaps will never -- be offset by the speed of parallel hardware. Therefore, it is essential to develop efficient and robust alternatives to direct methods that are also amenable to massively parallel processing because implicit codes using unconditionally stable time-integration algorithms are computationally more efficient when simulating low-frequency dynamics. Here we present a domain decomposition method for implicit schemes that requires significantly less storage than factorization algorithms, that is several times faster than other popular direct and iterative methods, that can be easily implemented on both shared and local memory parallel processors, and that is both computationally and communication-wise efficient. The proposed transient domain decomposition method is an extension of the method of Finite Element Tearing and Interconnecting (FETI) developed by Farhat and Roux for the solution of static problems. Serial and parallel performance results on the CRAY Y-MP/8 and the iPSC-860/128 systems are reported and analyzed for realistic structural dynamics problems. These results establish the superiority of the FETI method over both the serial/parallel conjugate gradient algorithm with diagonal scaling and the serial/parallel direct method, and contrast the computational power of the iPSC-860/128 parallel processor with that of the CRAY Y-MP/8 system.

Farhat, Charbel↗

A transient FETI methodology for large-scale parallel implicit computations in structural mechanics, part 2

Explicit codes are often used to simulate the nonlinear dynamics of large-scale structural systems, even for low frequency response, because the storage and CPU requirements entailed by the repeated factorizations traditionally found in implicit codes rapidly overwhelm the available computing resources. With the advent of parallel processing, this trend is accelerating because explicit schemes are also easier to parallellize than implicit ones. However, the time step restriction imposed by the Courant stability condition on all explicit schemes cannot yet and perhaps will never be offset by the speed of parallel hardware. Therefore, it is essential to develop efficient and robust alternatives to direct methods that are also amenable to massively parallel processing because implicit codes using unconditionally stable time-integration algorithms are computationally more efficient than explicit codes when simulating low-frequency dynamics. Here we present a domain decomposition method for implicit schemes that requires significantly less storage than factorization algorithms, that is several times faster than other popular direct and iterative methods, that can be easily implemented on both shared and local memory parallel processors, and that is both computationally and communication-wise efficient. The proposed transient domain decomposition method is an extension of the method of Finite Element Tearing and Interconnecting (FETI) developed by Farhat and Roux for the solution of static problems. Serial and parallel performance results on the CRAY Y-MP/8 and the iPSC-860/128 systems are reported and analyzed for realistic structural dynamics problems. These results establish the superiority of the FETI method over both the serial/parallel conjugate gradient algorithm with diagonal scaling and the serial/parallel direct method, and contrast the computational power of the iPSC-860/128 parallel processor with that of the CRAY Y-MP/8 system.

Farhat, Charbel↗

Exponential Runge-Kutta Parareal for non-diffusive equations

Parareal is a well-known parallel-in-time algorithm that combines a coarse and fine propagator within a parallel iteration. It allows for large-scale parallelism that leads to significantly reduced computational time compared to serial time-stepping methods. However, like many parallel-in-time methods it can fail to converge when applied to non-diffusive equations such as hyperbolic systems or dispersive nonlinear wave equations. Here, this paper explores the use of exponential integrators within the Parareal iteration. Exponential integrators are particularly interesting candidates for Parareal because of their ability to resolve fast-moving waves, even at the large stepsizes used by coarse propagators. This work begins with an introduction to exponential Parareal integrators followed by several motivating numerical experiments involving the nonlinear Schrödinger equation. These experiments are then analyzed using linear analysis that approximates the stability and convergence properties of the exponential Parareal iteration on nonlinear problems. The paper concludes with two additional numerical experiments involving the dispersive Kadomtsev-Petviashvili equation and the hyperbolic Vlasov-Poisson equation. These experiments demonstrate that exponential Parareal methods offer improved time-to-solution compared to serial exponential integrators when solving certain non-diffusive equations.

97 MATHEMATICS AND COMPUTING↗

PatchworkWave: A Multipatch Infrastructure for Multiphysics/Multiscale/Multiframe/Multimethod Simulations at Arbitrary Order

We present an extension of the PatchworkMHD code [1], itself an MHD-capable extension of thePatch-workcode [2], for which several algorithms presented here were co-developed. Its purpose is to create a“multipatch” scheme compatible with numerical simulations of arbitrary equations of motion at any dis-cretization order in space and time. In thePatchworkframework, the global simulation is comprised of anarbitrary number of moving, local meshes, or “patches”, which are free to employ their own resolution, co-ordinate system/topology, physics equations, reference frame, and in our new approach, numerical method.Each local patch exchanges boundary data with a single global patch on which all other patches residethrough a client-router-server parallelization model. In generalizingPatchworkto be compatible witharbitrary order time integration,PatchworkMHDandPatchworkWavehave significantly improved theinterpatch interpolation accuracy by removing an interpolation of interpolated data feedback present in theoriginalPatchworkcode. Furthermore, we extendPatchworkto bemultimethodby allowing multiplestate vectors to be updated simultaneously, with each state vector providing its own interpatch interpolationand transformation procedures. As such, our scheme is compatible with nearly any set of hyperbolic partialdifferential equations. We demonstrate our changes through the implementation of a scalar wave toy-modelthat is evolved on arbitrary, time dependent patch configurations at 4th order accuracy.

Dennis B Bowen↗

Image sensor with high dynamic range linear output

Designs and operational methods to increase the dynamic range of image sensors and APS devices in particular by achieving more than one integration times for each pixel thereof. An APS system with more than one column-parallel signal chains for readout are described for maintaining a high frame rate in readout. Each active pixel is sampled for multiple times during a single frame readout, thus resulting in multiple integration times. The operation methods can also be used to obtain multiple integration times for each pixel with an APS design having a single column-parallel signal chain for readout. Furthermore, analog-to-digital conversion of high speed and high resolution can be implemented.

Yadid-Pecht, Orly↗

Comparative Analysis of Radial and Random Microstructures of Mesophase Pitch Carbon Fibers

Carbon fibers (CF) with radial and random microstructures are produced. Here, these fibers are subjected to identical treatment before being mechanically tested and analyzed with Weibull analysis, with the results revealing a statistically significant difference in tensile strengths of 2.23 GPa for random CF and 1.69 GPa for radial CF. Raman mapping probed the crystalline structure perpendicular to the fiber axis and found a uniform structure, while wide‐angle X‐ray diffraction showed a significant difference of 7.5 Å in the crystallites’ basal lengths parallel to the fiber. Small‐angle X‐ray scattering is completed parallel to the fiber for the first time. A cross‐section Guinier plot of the 1D azimuthal integration is generated assuming symmetric scattering, and the parallel scatterers are found to have a similar length scale to the crystallite's length, validating the testing method. Finally, transmission electron microscopy is completed on the longitudinal cross‐section of each fiber. The radial carbon fiber is found to have a core–shell structure, as evidenced further by fast Fourier transform images. Through all studies, it is shown that the structure developed during mesophase pitch spinning altered the microstructure, thus impacting the mechanical properties, confirming a direct relationship between processing, structure, and properties.

Scherschel, Alexander [Univ. of Virginia, Charlott↗

Transient Finite Element Computations on a Variable Transputer System

A parallel program to analyze transient finite element problems was written and implemented on a system of transputer processors. The program uses the explicit time integration algorithm which eliminates the need for equation solving, making it more suitable for parallel computations. An interprocessor communication scheme was developed for arbitrary two dimensional grid processor configurations. Several 3-D problems were analyzed on a system with a small number of processors.

Smolinski, Patrick J.↗

A scalable exponential-DG approach for nonlinear conservation laws: With application to Burger and Euler equations

In this work, we propose an Exponential DG framework for partial differential equations. We decompose 7 governing equations into linear and nonlinear parts to which we apply the discontinuous Galerkin 8 (DG) spatial discretization. In particular, we construct the linear part using Jacobian that effectively 9 capture stiff characteristics in the system. The former is integrated analytically, whereas the latter 10 is approximated. This approach i) is stable with a large Courant number (Cr > 1); ii) supports 11 high-order solutions both in time and space; iii) is computationally favorable compared to IMEX 12 DG methods with no preconditioner; iv) becomes comparable to explicit RKDG methods on uniform 13 mesh and beneficial on non-uniform grid for Euler equations; v) is scalable in a modern massively 14 parallel computing architecture due to its explicit nature of exponential time integrators and com15 pact communication stencil of DG method. Numerical results demonstrate the performance of our 16 proposed methods through various examples. We also discuss the stability and convergence analysis 17 for our exponential DG scheme in the context of Burgers equation.

42 ENGINEERING↗

Module-Fluidics: Building Blocks for Spatio-Temporal Microenvironment Control

Generating the desired solute concentration signal in micro-environments is vital to many applications ranging from micromixing to analyzing cellular response to a dynamic microenvironment. We propose a new modular design to generate targeted temporally varying concentration signals in microfluidic systems while minimizing perturbations to the flow field. The modularized design, here referred to as module-fluidics, similar in principle to interlocking toy bricks, is constructed from a combination of two building blocks and allows one to achieve versatility and flexibility in dynamically controlling input concentration. The building blocks are an oscillator and an integrator, and their combination enables the creation of controlled and complex concentration signals, with different user-defined time-scales. We show two basic connection patterns, in-series and in-parallel, to test the generation, integration, sampling and superposition of temporally-varying signals. All such signals can be fully characterized by analytic functions, in analogy with electric circuits, and allow one to perform design and optimization before fabrication. Such modularization offers a versatile and promising platform that allows one to create highly customizable time-dependent concentration inputs which can be targeted to the specific application of interest.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Optimizing the hit finding algorithm for liquid argon TPC neutrino detectors using parallel architectures

Neutrinos are particles that interact rarely, so identifying them requires large detectors which produce lots of data. Processing this data with the computing power available is becoming even more difficult as the detectors increase in size to reach their physics goals. Liquid argon time projection chamber (LArTPC) neutrino experiments are expected to grow in the next decade to have 100 times more wires than in currently operating experiments, and modernization of LArTPC reconstruction code, including parallelization both at data- and instruction-level, will help to mitigate this challenge. The LArTPC hit finding algorithm is used across multiple experiments through a common software framework. In this paper we discuss a parallel implementation of this algorithm. Using a standalone setup we find speedup factors of two times from vectorization and 30–100 times from multi-threading on Intel architectures. The new version has been incorporated back into the framework so that it can be used by experiments. On a serial execution, the integrated version is about 10 times faster than the previous one and, once parallelization is enabled, more speedups comparable to the standalone program are achieved.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Enhancing Scalability for FUN3D Rotorcraft Simulations with Yoga: an Overset Grid Assembler

FUN3D, an unstructured grid Navier-Stokes CFD code is capable of overset grid simu- lations, but does not have an internal method for assembling an overset grid system from a group of component grids. FUN3D currently relies on the third party codes Suggar++ and DiRTlib to perform domain assembly and provide intergrid connectivity for overset simulations. Rotorcraft simulations with moving, deforming blades require domain assembly and mesh deformation at each time step. For these simulations, the three primary drivers of computational cost for each time step are: deforming the mesh, performing domain assembly, and performing subiterations of the flow solver. FUN3D exhibits strong and weak scalability for the flow solver subiterations and the mesh deformation. However, FUN3D is currently hardwired directly to the serial version of Suggar++, which has a fixed cost for a given mesh system. Therefore, domain assembly begins to dominate the total cost of each time step as grid systems become larger. An integrated method for parallel domain assembly is presented that addresses scalability for large grid systems. Restructuring within FUN3D to accommodate integrated domain assembly is also discussed, which could enable use of the parallel Suggar++ library.

Cameron T Druyor↗

Parallel Simulation of Unsteady Turbulent Flames

Time-accurate simulation of turbulent flames in high Reynolds number flows is a challenging task since both fluid dynamics and combustion must be modeled accurately. To numerically simulate this phenomenon, very large computer resources (both time and memory) are required. Although current vector supercomputers are capable of providing adequate resources for simulations of this nature, the high cost and their limited availability, makes practical use of such machines less than satisfactory. At the same time, the explicit time integration algorithms used in unsteady flow simulations often possess a very high degree of parallelism, making them very amenable to efficient implementation on large-scale parallel computers. Under these circumstances, distributed memory parallel computers offer an excellent near-term solution for greatly increased computational speed and memory, at a cost that may render the unsteady simulations of the type discussed above more feasible and affordable.This paper discusses the study of unsteady turbulent flames using a simulation algorithm that is capable of retaining high parallel efficiency on distributed memory parallel architectures. Numerical studies are carried out using large-eddy simulation (LES). In LES, the scales larger than the grid are computed using a time- and space-accurate scheme, while the unresolved small scales are modeled using eddy viscosity based subgrid models. This is acceptable for the moment/energy closure since the small scales primarily provide a dissipative mechanism for the energy transferred from the large scales. However, for combustion to occur, the species must first undergo mixing at the small scales and then come into molecular contact. Therefore, global models cannot be used. Recently, a new model for turbulent combustion was developed, in which the combustion is modeled, within the subgrid (small-scales) using a methodology that simulates the mixing and the molecular transport and the chemical kinetics within each LES grid cell. Finite-rate kinetics can be included without any closure and this approach actually provides a means to predict the turbulent rates and the turbulent flame speed. The subgrid combustion model requires resolution of the local time scales associated with small-scale mixing, molecular diffusion and chemical kinetics and, therefore, within each grid cell, a significant amount of computations must be carried out before the large-scale (LES resolved) effects are incorporated. Therefore, this approach is uniquely suited for parallel processing and has been implemented on various systems such as: Intel Paragon, IBM SP-2, Cray T3D and SGI Power Challenge (PC) using the system independent Message Passing Interface (MPI) compiler. In this paper, timing data on these machines is reported along with some characteristic results.

Menon, Suresh↗

Development and Verification of the Charring, Ablating Thermal Protection Implicit System Simulator

The development and verification of the Charring Ablating Thermal Protection Implicit System Solver (CATPISS) is presented. This work concentrates on the derivation and verification of the stationary grid terms in the equations that govern three-dimensional heat and mass transfer for charring thermal protection systems including pyrolysis gas flow through the porous char layer. The governing equations are discretized according to the Galerkin finite element method (FEM) with first and second order fully implicit time integrators. The governing equations are fully coupled and are solved in parallel via Newton s method, while the linear system is solved via the Generalized Minimum Residual method (GMRES). Verification results from exact solutions and Method of Manufactured Solutions (MMS) are presented to show spatial and temporal orders of accuracy as well as nonlinear convergence rates.

Amar, Adam J.↗

Development and Verification of the Charring Ablating Thermal Protection Implicit System Solver

The development and verification of the Charring Ablating Thermal Protection Implicit System Solver is presented. This work concentrates on the derivation and verification of the stationary grid terms in the equations that govern three-dimensional heat and mass transfer for charring thermal protection systems including pyrolysis gas flow through the porous char layer. The governing equations are discretized according to the Galerkin finite element method with first and second order implicit time integrators. The governing equations are fully coupled and are solved in parallel via Newton's method, while the fully implicit linear system is solved with the Generalized Minimal Residual method. Verification results from exact solutions and the Method of Manufactured Solutions are presented to show spatial and temporal orders of accuracy as well as nonlinear convergence rates.

Amar, Adam J.↗

Data-Analysis System for Entry, Descent, and Landing

A report describes the Entry Descent Landing Data Analysis (EDA), which is a system of signal-processing software and computer hardware for acquiring status data conveyed by multiple-frequency-shift-keying tone signals transmitted by a spacecraft during descent to the surface of a remote planet. The design of the EDA meets the challenge of processing weak, fluctuating signals that are Doppler-shifted by amounts that are only partly predictable. The software supports both real-time and post processing. The software performs fast-Fourier-transform integration, parallel frequency tracking with prediction, and mapping of detected tones to specific events. The use of backtrack and refinement parallel-processing threads helps to minimize data gaps. The design affords flexibility to enable division of a descent track into segments, within each of which the EDA is configured optimally for processing in the face of signal conditions and uncertainties. A dynamic-lock-state feature enables the detection of signals using minimum required computing power less when signals are steadily detected, more when signals fluctuate. At present, the hardware comprises eight dual-processor personal-computer modules and a server. The hardware is modular, making it possible to increase computing power by adding computers.

Pham, Timothy↗

Computational Investigation of Oxidative Etch Pitting in FiberForm and Its Impact on Material Properties

Oxidation-driven carbon erosion does not occur uniformly but rather through the development of localized etch pits at active surface sites. These active sites form due to atomic defects on the carbon surface, making them significantly more reactive than the surrounding, non-defective areas. As a result, these sites are the first to react during ablation, leading to their removal. This process creates new defects in neighboring atoms, increasing their reactivity and causing localized carbon removal around these active sites. In this way, the highly reactive defective areas serve as nucleation points for the formation and growth of etch pits, which can have adverse effects on structural integrity of FiberForm. To better understand how these etch pits impact the material properties of carbon fiber microstructures, we have developed a new capability within the direct simulation Monte Carlo (DSMC) framework to capture the etch pit formation process. This capability, integrated into the DSMC code SPARTA (Stochastic Parallel Rarefied-gas Time-accurate Analyzer), models material removal in the presence of active sites, leading to the formation of etch pits. The current work focuses on studying the effects of these etch pits on the material properties of FiberForm, a widely used base material in thermal protection systems (TPS). The microstructure of virgin FiberForm, obtained via X-ray microtomography, is imported into SPARTA to generate the ablated geometries with etch pits. These modified microstructures are then analyzed using the Porous Microstructure Analysis (PuMA) software to compute various material properties, including elasticity, thermal conductivity, and permeability. We investigate the variation of these properties due to the complex surface topology changes caused by etch pit formation. Additionally, we compare the effects of pitting with the conventional model of shrinking fibers, traditionally used to simulate the ablation of carbon structures. Significant differences emerge between the two approaches. Consequently, this physically realistic model of material removal through etch pit formation offers improved accuracy in predicting the degradation of carbon-based TPS during oxidation. It also provides insights into other mechanisms, such as spallation, where chunks of material are removed into the flow due to etch pit growth. Ultimately, this model enhances our understanding of failure modes in these materials during ablation.

PuMA↗