Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel codes”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20

Parallel DSMC Solution of Three-Dimensional Flow Over a Finite Flat Plate

This paper describes a parallel implementation of the direct simulation Monte Carlo (DSMC) method. Runtime library support is used for scheduling and execution of communication between nodes, and domain decomposition is performed dynamically to maintain a good load balance. Performance tests are conducted using the code to evaluate various remapping and remapping-interval policies, and it is shown that a one-dimensional chain-partitioning method works best for the problems considered. The parallel code is then used to simulate the Mach 20 nitrogen flow over a finite-thickness flat plate. It is shown that the parallel algorithm produces results which compare well with experimental data. Moreover, it yields significantly faster execution times than the scalar code, as well as very good load-balance characteristics.

Nance, Robert P.↗

Improving turbo-like codes using iterative decoder analysis

The density evolution method is used to analyze the performance and optimize the structure of parallel and serial turbo codes, and generalized serial concatenations of mixtures of different outer and inner codes. Design examples are given for mixture codes.

turbo codes iterative decoding↗

Serial and Hybrid Concatenated Codes with Applications

Analytical bounds on the performance of concatenated codes on a tree structure are obtained. Analytical results are applied to examples of parallel concatenation of two codes (turbo codes), serial concatenation of two codes, hybrid concatenation of three codes, and self concatenated codes, over AWGN and fading channels.

AWGN MPSK modulations Rayleigh fading channels↗

ChatMPI: LLM-Driven MPI Code Generation for HPC Workloads

The Message Passing Interface (MPI) standard plays a crucial role in enabling scientific applications for parallel computing and is an essential component in high-performance computing (HPC). However, implementing MPI code manually—especially applying a proper domain decomposition and communication pattern—is a challenging and error-prone task. We present ChatMPI, an AI assistant for MPI parallelization of sequential C codes. In our analysis, we focus on testing six essential HPC workloads, which are based on Basic Linear Algebra Subprograms levels 1, 2, and 3 as well as sparse, stencil, and iterative operations. We analyze the process of creating ChatMPI by using the ChatHPC library. This lightweight large language model (LLM)–based infrastructure enables HPC experts to efficiently create and supervise trustworthy AI capabilities for critical HPC software tasks. We study the data required for training (fine-tuning) ChatMPI to generate parallel codes that not only use MPI syntax correctly but also apply HPC techniques to reduce memory communication and maximize performance by using proper work decomposition. With a relatively small training dataset composed of a few dozen prompts and fewer than 15 minutes of fine-tuning on one node equipped with two NVIDIA H100 GPUs, ChatMPI elevates trustworthiness for MPI code generation of current LLMs (e.g., Code Llama, ChatGPT-4o and ChatGPT 5). Additionally, we evaluate the performance of the MPI codes generated by ChatMPI in comparison with the ones generated by ChatGPT-4o and ChatGPT-5. The codes generated by ChatMPI provide up to a 4 × boost in performance by using better problem decomposition, communication patterns, and HPC techniques (e.g., communication avoiding).

Valero Lara, Pedro [ORNL] (ORCID:0000000214794310)↗

A transient FETI methodology for large-scale parallel implicit computations in structural mechanics, part 2

Explicit codes are often used to simulate the nonlinear dynamics of large-scale structural systems, even for low frequency response, because the storage and CPU requirements entailed by the repeated factorizations traditionally found in implicit codes rapidly overwhelm the available computing resources. With the advent of parallel processing, this trend is accelerating because explicit schemes are also easier to parallellize than implicit ones. However, the time step restriction imposed by the Courant stability condition on all explicit schemes cannot yet and perhaps will never be offset by the speed of parallel hardware. Therefore, it is essential to develop efficient and robust alternatives to direct methods that are also amenable to massively parallel processing because implicit codes using unconditionally stable time-integration algorithms are computationally more efficient than explicit codes when simulating low-frequency dynamics. Here we present a domain decomposition method for implicit schemes that requires significantly less storage than factorization algorithms, that is several times faster than other popular direct and iterative methods, that can be easily implemented on both shared and local memory parallel processors, and that is both computationally and communication-wise efficient. The proposed transient domain decomposition method is an extension of the method of Finite Element Tearing and Interconnecting (FETI) developed by Farhat and Roux for the solution of static problems. Serial and parallel performance results on the CRAY Y-MP/8 and the iPSC-860/128 systems are reported and analyzed for realistic structural dynamics problems. These results establish the superiority of the FETI method over both the serial/parallel conjugate gradient algorithm with diagonal scaling and the serial/parallel direct method, and contrast the computational power of the iPSC-860/128 parallel processor with that of the CRAY Y-MP/8 system.

Farhat, Charbel↗

Extension of the high-resolution thermal-hydraulics code ESCOT to hexagonal core geometries for multi-physics calculations

The extension of the capabilities of the pin-level nuclear reactor core thermal-hydraulics (T/H) code ESCOT to analyze hexagonal fueled cores and its performance are presented. ESCOT is an accurate yet fast core thermal-hydraulics solution aiming at high-fidelity and high-resolution multi-physics core analysis in the framework of massively parallel computing platforms. Its algorithm solution is based on the four-equation drift-flux model for two-phase calculations, these are numerically solved by applying the Finite Volume Method (FVM) and the Semi-Implicit Method for Pressure-Linked Equation (SIMPLE)-like algorithm in a staggered grid system. Constitutive models such as turbulent mixing, pressure drop, and vapor generation are employed to simulate key phenomena in subchannel-scale analysis. ESCOT is parallelized by a double (radial and axial) domain decomposition that enables its highly parallelized execution. The coupling of the code with the neutronics whole core solver for hexagonal geometries nTRACER is described. The newly implemented ESCOT features are validated by comparing single assembly and full core steady state nTRACER-ESCOT solutions with nTRACER standalone internal one-dimensional T/H solver results. The validation problems are based on the VVER 440 and VVER 1000 cores. ESCOT results show differences within an acceptable range with respect to the simple 1D nTRACER built-in solver. (authors)

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Parallelization of the Physical-Space Statistical Analysis System (PSAS)

Atmospheric data assimilation is a method of combining observations with model forecasts to produce a more accurate description of the atmosphere than the observations or forecast alone can provide. Data assimilation plays an increasingly important role in the study of climate and atmospheric chemistry. The NASA Data Assimilation Office (DAO) has developed the Goddard Earth Observing System Data Assimilation System (GEOS DAS) to create assimilated datasets. The core computational components of the GEOS DAS include the GEOS General Circulation Model (GCM) and the Physical-space Statistical Analysis System (PSAS). The need for timely validation of scientific enhancements to the data assimilation system poses computational demands that are best met by distributed parallel software. PSAS is implemented in Fortran 90 using object-based design principles. The analysis portions of the code solve two equations. The first of these is the "innovation" equation, which is solved on the unstructured observation grid using a preconditioned conjugate gradient (CG) method. The "analysis" equation is a transformation from the observation grid back to a structured grid, and is solved by a direct matrix-vector multiplication. Use of a factored-operator formulation reduces the computational complexity of both the CG solver and the matrix-vector multiplication, rendering the matrix-vector multiplications as a successive product of operators on a vector. Sparsity is introduced to these operators by partitioning the observations using an icosahedral decomposition scheme. PSAS builds a large (approx. 128MB) run-time database of parameters used in the calculation of these operators. Implementing a message passing parallel computing paradigm into an existing yet developing computational system as complex as PSAS is nontrivial. One of the technical challenges is balancing the requirements for computational reproducibility with the need for high performance. The problem of computational reproducibility is well known in the parallel computing community. It is a requirement that the parallel code perform calculations in a fashion that will yield identical results on different configurations of processing elements on the same platform. In some cases this problem can be solved by sacrificing performance. Meeting this requirement and still achieving high performance is very difficult. Topics to be discussed include: current PSAS design and parallelization strategy; reproducibility issues; load balance vs. database memory demands, possible solutions to these problems.

Larson, J. W.↗

SPACE: 3D parallel solvers for Vlasov-Maxwell and Vlasov-Poisson equations for relativistic plasmas with atomic transformations

A parallel, relativistic, three-dimensional particle-in-cell code SPACE has been developed for the simulation of electromagnetic fields, relativistic particle beams, and plasmas. In addition to the standard second-order Particle-in-Cell (PIC) algorithm, SPACE includes efficient novel algorithms to resolve atomic physics processes such as multi-level ionization of plasma atoms, recombination, and electron attachment to dopants in dense neutral gases. SPACE also contains a highly adaptive particle-based method, called Adaptive Particle-in-Cloud (AP-Cloud), for solving the Vlasov-Poisson problems. It eliminates the traditional Cartesian mesh of PIC and replaces it with an adaptive octree data structure. The code's algorithms, structure, capabilities, parallelization strategy, and performance have been discussed. Additionally, typical examples of SPACE applications to accelerator science and engineering problems are described.

43 PARTICLE ACCELERATORS↗

TUMME: Tsinghua University Minnesota Master Equation program

We report that TUMME is a program for assembling and solving master equations for gas-phase chemical kinetics based on chemically significant eigenmodes. TUMME has interfaces to the Gaussian, Polyrate, and/or MSTor output files that allow the master equation code to obtain the microcanonical flux coefficients needed for the coefficient matrix of the master equation. The flux coefficients for reactions with barriers can be calculated by multi-structural variational transition state theory with small-curvature tunneling (MS-VTST/SCT) or by simpler approximations to this such as conventional transition state theory without tunneling (also called RRKM theory). The flux coefficients for barrierless reactions are provided by a hard-sphere model. TUMME is written in double precision with Python 3; quadruple and octuple precision are also available for some subtasks in C++. The Python code can run in serial or parallel (MP or MPI), and the C++ code can run on a single processor or on multiple processors with OpenMP.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

UCLA parallel PIC framework

The UCLA Parallel PIC Framework (UPIC) has been developed to provide trusted components for the rapid construction of new, parallel Particle-in-Cell (PIC) codes. The Framework uses object-based ideas in Fortran95, and is designed to provide support for various kinds of PIC codes on various kinds of hardware. The focus is on student programmers. The Framework supports multiple numerical methods, different physics approximations, different numerical optimizations and implementations for different hardware. It is designed with "defensive" programming in mind, meaning that it contains many error checks and debugging helps. Above all, it is designed to hide the complexity of parallel processing. It is currently being used in a number of new Parallel PIC codes.

Norton, Charles D.↗

A Parallelized Oxidation-Driven Surface Recession Framework in DSMC Code, SPARTA

Spacecrafts rely on ablative thermal protection systems (TPS) made of composites consisting of a carbon-based reinforcement and a polymeric matrix. These materials are designed to withstand high-temperature oxidation and surface recession during re-entry into the Earth's atmosphere. However, ablation occurs due to a complex interplay of thermal, mechanical, and chemical factors, making it challenging to determine the individual impact of each on the TPS's overall degradation. In this study, we have developed an ablation model that can leverage a finite rate carbon oxidation model to predict material recession and surface states more accurately. Stochastic PArallel Rarified-gas Time-accurate Analyzer (SPARTA), a direct-simulation Monte Carlo (DSMC) code, is modified to allow oxidation-driven ablation of implicitly defined carbon surfaces. In SPARTA, implicit surfaces are generated from the grid corner point values via a marching cubes algorithm, therefore creating a new set of surface elements every time ablation is performed. The finite-rate oxidation model developed by Gopalan et. al can perform both gas-surface and pure-surface reactions and is now adapted to tally surface data on a per grid cell basis. The ablation functionality was also adjusted so once the reactions have occurred, the number of reactions leading to CO formation can be converted to corner point reduction values; therefore, carbon removal is directly proportional to surface recession. We also briefly discuss some unique challenges associated with parallelizing this dynamic surface state and geometry. Finally, we analyze the performance of this parallelized implicit chemistry model with simple 2D and 3D benchmark cases by producing surface state statistics, area changes over time, and visualization across a range of surface temperatures and processors with and without load-balancing.

DSMC↗

Parallel-vector computation for linear structural analysis and non-linear unconstrained optimization problems

Several parallel-vector computational improvements to the unconstrained optimization procedure are described which speed up the structural analysis-synthesis process. A fast parallel-vector Choleski-based equation solver, pvsolve, is incorporated into the well-known SAP-4 general-purpose finite-element code. The new code, denoted PV-SAP, is tested for static structural analysis. Initial results on a four processor CRAY 2 show that using pvsolve reduces the equation solution time by a factor of 14-16 over the original SAP-4 code. In addition, parallel-vector procedures for the Golden Block Search technique and the BFGS method are developed and tested for nonlinear unconstrained optimization. A parallel version of an iterative solver and the pvsolve direct solver are incorporated into the BFGS method. Preliminary results on nonlinear unconstrained optimization test problems, using pvsolve in the analysis, show excellent parallel-vector performance indicating that these parallel-vector algorithms can be used in a new generation of finite-element based structural design/analysis-synthesis codes.

Nguyen, D. T.↗

Characterizing the effect of hypersonic boundary layer turbulence on antenna performance: A computational approach

The degradation of antenna performance during hypersonic re-entry is a well known phenomenon that can lead to complete radio blackout. Recent additions to the Empire code establish it as a tool for the study and analysis of the problem. Coupling to the Sandia Parallel Aerodynamics and Reentry Code (SPARC) enables the electromagnetic analysis of realistic re-entry plasma profiles. The geometric flexibility afforded by both Empire and SPARC allow the consideration of arbitrary vehicle and antenna configurations. We have used this tool to study antenna performance during re-entry when the boundary layer becomes turbulent. A concise description of line-of-sight transmissions, which employs advanced statistical methods, was developed. New insights into the low altitude reflectometer readings of RAM-C2 are offered. Techniques for the reconstruction of the re-entry plasma profile from reflectometer data were explored.

42 ENGINEERING↗

Simulations of particle acceleration in parallel shocks: Direct comparison between Monte Carlo and one-dimensional hybrid codes

We have made a direct comparison between two different computer simulations of a plane, parallel, collisionless shock including particle acceleration to energies typical of those of diffuse ions observed at the earth bow shock. Despite the fact that the one-dimensional hybrid and Monte Carlo techniques employ entirely different algorithms, they give surprisingly close agreement in the overall shapes of the complete distribution functions for protons as well as heavier ions. Both methods show that energetic ions emerge smoothly from the background thermal plasma with approximately the same relative injection rate and that the fraction of the incoming plasma's energy flux that is converted into downstream enthalpy flux of the accelerated population (i.e., the acceleration efficiency) is similar in the two cases. The fraction of the downstream proton distribution made up of superthermal particles is quite large, with at least 10% of the energy flux going into protons with energies above 10 keV. In addition, an upstream precursor, produced by backstreaming energetic particles, is present in both shocks, although the Monte Carlo precursor is considerably longer than that produced in the hybrid shock. These results offer convincing evidence that, at least in these ways, the two simulations are consistent in their description of parallel shock structure and particle acceleration, and they lay the groundwork for development of shock models employing a combination of both methods.

Ellison, Donald C.↗

Efficient Routing of Quantum LDPC Codes on Programmable 2D Toric Architectures

Quantum low-density parity-check codes are promising candidates towards scalable fault-tolerant quantum computation. Among these, bivariate bicycle (BB) codes offer superior encoding rates and large code distance compared to surface codes. However, their requirement on long-range stabilizer measurements poses significant challenges for implementation on realistic hardware with limited connectivity, such as superconducting circuit platforms. In this work, we introduce a novel hardware-software co-design that leverages a programmable communication network architecture to address these limitations. Our approach utilizes a 2D toric network of oscillators as a flexible communication fabric linking qubits at each site. Such architecture significantly reduces the number of long-range couplers required from O ( n ) to O (√ n ). Dual-rail qubits, along with native gates including Swap-Wait-Swap gates and beamsplitter SWAPs, ensure that long-range two-qubit gates can be executed with high fidelity and low latency. To further enhance performance, our qubit layout and routing algorithm utilize symmetries of the codes and enable maximum parallelism for long-range two-qubit gates, maintaining a low syndrome extraction cycle duration and scalability over the code length. We perform circuit-level simulation with realistic noise modeling based on experimental hardware parameters, observing an logical error rate per logical qubit per cycle of 3.06% for [[18,4,4]] BB code, 2.6× less than the existing experimental result. These findings provide a practical roadmap and identify key technological advancements needed to achieve low-overhead fault-tolerant quantum computing at scale.

Liu, Kun [Yale Univ., New Haven, CT (United States↗

Computational strategies for three-dimensional flow simulations on distributed computer systems

An increasing amount of research activity in computational fluid dynamics has been devoted to the development of efficient algorithms for parallel computing systems. The increasing performance to price ratio of engineering workstations has led to research to development procedures for implementing a parallel computing system composed of distributed workstations. This thesis proposal outlines an ongoing research program to develop efficient strategies for performing three-dimensional flow analysis on distributed computing systems. The PVM parallel programming interface was used to modify an existing three-dimensional flow solver, the TEAM code developed by Lockheed for the Air Force, to function as a parallel flow solver on clusters of workstations. Steady flow solutions were generated for three different wing and body geometries to validate the code and evaluate code performance. The proposed research will extend the parallel code development to determine the most efficient strategies for unsteady flow simulations.

Weed, Richard Allen↗

Method of moment solutions to scattering problems in a parallel processing environment

This paper describes the implementation of a parallelized method of moments (MOM) code into an interactive workstation environment. The workstation allows interactive solid body modeling and mesh generation, MOM analysis, and the graphical display of results. After describing the parallel computing environment, the implementation and results of parallelizing a general MOM code are presented in detail.

Cwik, Tom↗