Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel codes”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Xyce Parallel Electronic Simulator Users' Guide Version 7.4

This manual describes the use of the Xyce Parallel Electronic Simulator. Xyce has been designed as a SPICE-compatible, high-performance analog circuit simulator, and has been written to support the simulation needs of the Sandia National Laboratories electrical designers. This development has focused on improving capability over the current state-of-the-art in the following areas: • Capability to solve extremely large circuit problems by supporting large-scale parallel computing platforms (up to thousands of processors). This includes support for most popular parallel and serial computers. • A differential-algebraic-equation (DAE) formulation, which better isolates the device model package from solver algorithms. This allows one to develop new types of analysis without requiring the implementation of analysis-specific device models. • Device models that are specifically tailored to meet Sandia’s needs, including some radiation-aware devices (for Sandia users only). • Object-oriented code design and implementation using modern coding practices. Xyce is a parallel code in the most general sense of the phrase — a message passing parallel implementation — which allows it to run efficiently a wide range of computing platforms. These include serial, shared-memory and distributed-memory parallel platforms. Attention has been paid to the specific nature of circuit-simulation problems to ensure that optimal parallel efficiency is achieved as the number of processors grows.

42 ENGINEERING↗

Xyce™ Parallel Electronic Simulator Users' Guide (V.7.9)

This manual describes the use of the Xyce Parallel Electronic Simulator. Xyce has been designed as a SPICE-compatible, high-performance analog circuit simulator, and has been written to support the simulation needs of the Sandia National Laboratories electrical designers. This development has focused on improving capability over the current state-of-the-art in the following areas: • Capability to solve extremely large circuit problems by supporting large-scale parallel computing platforms (up to thousands of processors). This includes support for most popular parallel and serial computers. • A differential-algebraic-equation (DAE) formulation, which better isolates the device model package from solver algorithms. This allows one to develop new types of analysis without requiring the implementation of analysis-specific device models. • Device models that are specifically tailored to meet Sandia’s needs, including some radiation-aware devices (for Sandia users only). • Object-oriented code design and implementation using modern coding practices. Xyce is a parallel code in the most general sense of the phrase — a message passing parallel implementation — which allows it to run efficiently a wide range of computing platforms. These include serial, shared-memory and distributed-memory parallel platforms. Attention has been paid to the specific nature of circuit-simulation problems to ensure that optimal parallel efficiency is achieved as the number of processors grows.

42 ENGINEERING↗

Xyce™ Parallel Electronic Simulator Users’ Guide, Version 7.10

This manual describes the use of the Xyce Parallel Electronic Simulator. Xyce has been designed as a SPICE-compatible, high-performance analog circuit simulator, and has been written to support the simulation needs of the Sandia National Laboratories electrical designers. This development has focused on improving capability over the current state-of-the-art in the following areas: • Capability to solve extremely large circuit problems by supporting large-scale parallel computing platforms (up to thousands of processors). This includes support for most popular parallel and serial computers. • A differential-algebraic-equation (DAE) formulation, which better isolates the device model package from solver algorithms. This allows one to develop new types of analysis without requiring the implementation of analysis-specific device models. • Device models that are specifically tailored to meet Sandia’s needs, including some radiation-aware devices (for Sandia users only). • Object-oriented code design and implementation using modern coding practices. Xyce is a parallel code in the most general sense of the phrase — a message passing parallel implementation — which allows it to run efficiently a wide range of computing platforms. These include serial, shared-memory and distributed-memory parallel platforms. Attention has been paid to the specific nature of circuit-simulation problems to ensure that optimal parallel efficiency is achieved as the number of processors grows.

97 MATHEMATICS AND COMPUTING↗

Xyce Parallel Electronic Simulator Users' Guide (V.7.1)

This manual describes the use of the Xyce Parallel Electronic Simulator. Xyce has been designed as a SPICE-compatible, high-performance analog circuit simulator, and has been written to support the simulation needs of the Sandia National Laboratories electrical designers. This development has focused on improving capability over the current state-of-the-art in the following areas: 1) Capability to solve extremely large circuit problems by supporting large-scale parallel computing platforms (up to thousands of processors). This includes support for most popular parallel and serial computers. 2) A differential-algebraic-equation (DAE) formulation, which better isolates the device model package from solver algorithms. This allows one to develop new types of analysis without requiring the implementation of analysis-specific device models. 3) Device models that are specifically tailored to meet Sandia's needs, including some radiation-aware devices (for Sandia users only). 4) Object-oriented code design and implementation using modern coding practices. Xyce is a parallel code in the most general sense of the phrase a message passing parallel implementation which allows it to run efficiently a wide range of computing platforms. These include serial, shared-memory and distributed-memory parallel platforms. Attention has been paid to the specific nature of circuit-simulation problems to ensure that optimal parallel efficiency is achieved as the number of processors grows.

42 ENGINEERING↗

Xyce™ Parallel Electronic Simulator Users' Guide, Version 7.5.

This manual describes the use of the Xyce Parallel Electronic Simulator. Xyce has been designed as a SPICE-compatible, high-performance analog circuit simulator, and has been written to support the simulation needs of the Sandia National Laboratories electrical designers. This development has focused on improving capability over the current state-of-the-art in the following areas: (1) Capability to solve extremely large circuit problems by supporting large-scale parallel computing platforms (up to thousands of processors). This includes support for most popular parallel and serial computers. (2) A differential-algebraic-equation (DAE) formulation, which better isolates the device model package from solver algorithms. This allows one to develop new types of analysis without requiring the implementation of analysis-specific device models. (3) Device models that are specifically tailored to meet Sandia's needs, including some radiation-aware devices (for Sandia users only). (4) Object-oriented code design and implementation using modern coding practices. Xyce is a parallel code in the most general sense of the phrase — a message passing parallel implementation — which allows it to run efficiently a wide range of computing platforms. These include serial, shared-memory and distributed-memory parallel platforms. Attention has been paid to the specific nature of circuit-simulation problems to ensure that optimal parallel efficiency is achieved as the number of processors grows.

97 MATHEMATICS AND COMPUTING↗

Xyce™ Parallel Electronic Simulator Users' Guide (V.7.6)

This manual describes the use of the Xyce™ Parallel Electronic Simulator. Xyce™ has been designed as a SPICE-compatible, high-performance analog circuit simulator, and has been written to support the simulation needs of the Sandia National Laboratories electrical designers. This development has focused on improving capability over the current state-of-the-art in the following areas: (1) Capability to solve extremely large circuit problems by supporting large-scale parallel computing platforms (up to thousands of processors). This includes support for most popular parallel and serial computers. (2) A differential-algebraic-equation (DAE) formulation, which better isolates the device model package from solver algorithms. This allows one to develop new types of analysis without requiring the implementation of analysis-specific device models. (3) Device models that are specifically tailored to meet Sandia's needs, including some radiation-aware devices (for Sandia users only). (4) Object-oriented code design and implementation using modern coding practices. Xyce™ is a parallel code in the most general sense of the phrase—a message passing parallel implementation—which allows it to run efficiently a wide range of computing platforms. These include serial, shared-memory and distributed-memory parallel platforms. Attention has been paid to the specific nature of circuit-simulation problems to ensure that optimal parallel eficiency is achieved as the number of processors grows.

97 MATHEMATICS AND COMPUTING↗

Xyce™ Parallel Electronic Simulator Users’ Guide, Version 7.8

This manual describes the use of the Xyce Parallel Electronic Simulator. Xyce has been designed as a SPICE-compatible, high-performance analog circuit simulator, and has been written to support the simulation needs of the Sandia National Laboratories electrical designers. This development has focused on improving capability over the current state-of-the-art in the following areas: (1) Capability to solve extremely large circuit problems by supporting large-scale parallel computing platforms (up to thousands of processors). This includes support for most popular parallel and serial computers. (2) A differential-algebraic-equation (DAE) formulation, which better isolates the device model package from solver algorithms. This allows one to develop new types of analysis without requiring the implementation of analysis-specific device models. (3) Device models that are specifically tailored to meet Sandia’s needs, including some radiation-aware devices (for Sandia users only). (4) Object-oriented code design and implementation using modern coding practices. Xyce is a parallel code in the most general sense of the phrase — a message passing parallel implementation — which allows it to run efficiently a wide range of computing platforms. These include serial, shared-memory and distributed-memory parallel platforms. Attention has been paid to the specific nature of circuit-simulation problems to ensure that optimal parallel efficiency is achieved as the number of processors grows.

42 ENGINEERING↗

Dynamic load balancing with enhanced shared-memory parallelism for particle-in-cell codes

Furthering our understanding of many of today’s interesting problems in plasma physics – including plasma based acceleration and magnetic reconnection with pair production due to quantum electrodynamic effects – requires large-scale kinetic simulations using particle-in-cell (PIC) codes. However, these simulations are extremely demanding, requiring that contemporary PIC codes be designed to efficiently use a new fleet of exascale computing architectures. To this end, the key issue of parallel load balance across computational nodes must be addressed. We discuss the implementation of dynamic load balancing by dividing the simulation space into many small, self-contained regions or ‘‘tiles,’’ along with shared-memory (e.g., OpenMP) parallelism both over many tiles and within single tiles. The load balancing algorithm can be used with three different topologies, including two space-filling curves. Here, we tested this implementation in the code Osiris and show low overhead and improved scalability with OpenMP thread number on simulations with both uniform load and severe load imbalance. Compared to other load-balancing techniques, our algorithm gives order-of-magnitude improvement in parallel scalability for simulations with severe load imbalance issues.

97 MATHEMATICS AND COMPUTING↗

EBS Task Force: Task 9/FEBEX Modeling Final Report: Thermo-Hydrological Modeling with PFLOTRAN

This report outlines Sandia National Laboratories modeling studies applied to Stage 1 and Stage 2 of the Full-scale Engineered Barriers Experiment in Crystalline Host Rock (FEBEX) in situ test for the SKB EBS Task Force Task 9. The FEBEX test was a full-scale test conducted over ~18 years at the Grimsel, Switzerland Underground Research Laboratory (URL) managed by NAGRA. It involved emplacing simulated waste packages, in the form of welded cylindrical heaters, inside a tunnel in crystalline granitic rock and surrounded by a bentonite barrier and cement plug. Sensors emplaced within the bentonite monitored the wetting-up, heating, and drying out of the bentonite barrier, and the large resulting data set provides an excellent opportunity for validation of multiphysics Thermal-Hydrological (TH), Thermal-Hydrologic-Chemical (THC), and Thermal-Hydrological-Mechanical (THM) modeling approaches for underground nuclear waste storage and the performance of engineered bentonite barriers. The present status of the EBS Task Force is finalizing Task 9, which follows years of modeling studies of the FEBEX test, by many notable modeling teams (Gens et al., 2009; Sanchez et al. 2010; 2012; Samper et al., 2018). These modeling studies generally use two-dimensional axisymmetric meshes, ignoring threedimensional effects, gravity and asymmetric wetting and dry out of the bentonite engineered barrier. This study investigates these effects with use of the PFLOTRAN THC code with massively parallel computational methods in modeling FEBEX Stage 1 and Stage 2 results. The PFLOTRAN numerical code is an open source, state-of-the-art, massively parallel subsurface flow and reactive transport code operating in a high-performance computing environment (Hammond et al., 2014). Section 2 describes the applied partial differential equations describing mass, momentum and energy balance used in this study, considerations derived by assuming phase equilibrium between gas and liquid phases, constitutive equations for granite, cement plug, and bentonite domains, and specific approaches for use inthe PFLOTRAN code. Section 3 describes the geometry, meshing, and model set-up. Section 4 describes modeling results, Section 5 compares modeling results to field testing data, and Section 6 gives conclusions. The Appendix provides detailed information required by the EBSTask Force for final reporting.

42 ENGINEERING↗

Parallelization and Performance Portability in Hydrodynamics Codes

With the eve of Exascale computing, performance and portability are at the forefront of all scientific codes. Adding more cores and more energy to a system is no longer a sustainable way to achieve performance, and extra effort must now be made to improve performance in all areas of code and code development. Using hydrodynamic codes as a basis, this work explores numerous techniques to achieve performance in different ways. Adaptive mesh refinement (AMR) is a necessary technique to improve memory optimization in mesh-based simulations. However it is invasive and conventionally difficult to integrate into existing applications, so we present a new branch of AMR to create a smooth transition to these optimizations, which not only improves performance, but also greatly reduces developer effort. We introduce the concept of this improvement as Phantom-Cell AMR, and assess theoretically the improvements, as well as present an application of its use. Other work included involves and investigation into an efficient data structure that ensures optimal memory layout for cache performance, with a target of making codes performant and portable across all architectures. All of the work targets both performance and portability, not just on CPU hardware, but specifically across GPU architectures. Parallel performance is key to all of the methods presented, but the research makes a great effort to improve the portability of all applications to prepare for current high performance computing systems and those on the horizon.

97 MATHEMATICS AND COMPUTING↗

Flexible and Modular Simultaneous Modeling of Flow and Reactive Transport in Rivers and Hyporheic Zones

Investigations of coupled multiphysics processes in rivers and hyporheic zones have extensively used numerical models. Most existing models use a sequential, one-way coupling between the surface and subsurface domains. Such one-way coupling potentially introduces error. To overcome this, a fully coupled model, hyporheicFoam, was developed using the open-source computational platform OpenFOAM. It captures the coupled flow and multicomponent reactive transport processes within both surface and subsurface domains and across their interface. The coupling between two domains is implemented by mapping conservative flux boundary conditions at the interface through an iterative algorithm. Reactive transport is enabled by specifying a reaction network. To start, we have implemented reaction kinetics following the double Monod-type model with inhibition. The model capability is illustrated through modeling of both conservative and reactive hyporheic flow and transport through dune bedforms. With the novel coupled model, it is now possible to quantify reactions wherein the reactants and products are constantly exchanging between domains and have feedbacks. hyporheicFoam can simulate large, three-dimensional cases owing to the computational flexibility and power offered by the code structure and parallel design of OpenFOAM.

58 GEOSCIENCES↗

Enhancing ChatPORT with CUDA-to-SYCL Kernel Translation Capability

Large Language Models (LLMs) have shown strong capabilities in general code translation. However, code translation involving parallel programming models remains largely unexplored. This work enhances the capabilities of code LLMs in CUDA-to-SYCL kernel translation with parameter-efficient fine-tuning. The resultant fine-tuned LLM, called ChatPORT, is an effort to provide high-fidelity translations from one programming model to another. We describe the preparation of datasets from heterogeneous computing benchmarks for model fine-tuning and testing, the parameter-efficient fine-tuning of 19 open-source code models ranging in size from 0.5 to 34 billion parameters and evaluate the correctness rates of the SYCL kernels by the fine-tuned models. The experimental results show that most code models fail to translate CUDA codes to SYCL correctly. However, fine-tuning these models using a small set of CUDA and SYCL kernels can enhance the capabilities of these models in kernel translation. Depending on the sizes of the models, the correctness rate ranges from 19.9% to 81.7% for a test dataset of 62 CUDA kernels.

Jin, Zheming [ORNL] (ORCID:000000027197780X)↗

Parallel Implicit Hydrodynamics for High Explosive Burn Calculations

High explosives in hostile environments will require calculational capabilities that model processes, which evolve on timescales from minutes to nanoseconds. Eventually the HE will begin to move metal. To handle this temporal evolution an implicit hydrodynamics coupled to the chemical release of the HE energy is required. In addition, the use of chemical kinetics to model the transition from, the initially, slow heating of a confined high explosive through to deflagration and on to detonation requires many computational zones to model high explosive engineering systems. This requirement means that a fully parallel implicit hydrodynamics is essential. In this paper we present the calculation of a nonlinear matrix equation for the advanced particle pressure that has been made parallel and implemented in our AMR code, BABBO. This new parallel implicit hydrodynamics has been applied to a cookoff problem, as well as, one and two dimensional shock problems. Results are presented and discussed.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

symPACK: A GPU-Capable Fan-Out Sparse Cholesky Solver

Sparse symmetric positive definite systems of equations are ubiquitous in scientific workloads and applications. Parallel sparse Cholesky factorization is the method of choice for solving such linear systems. Therefore, the development of parallel sparse Cholesky codes that can efficiently run on today’s large-scale heterogeneous distributed-memory platforms is of vital importance. Modern supercomputers offer nodes that contain a mix of CPUs and GPUs. To fully utilize the computing power of these nodes, scientific codes must be adapted to offload expensive computations to GPUs. We present symPACK, a GPU-capable parallel sparse Cholesky solver that uses one-sided communication primitives and remote procedure calls provided by the UPC++ library. We also utilize the UPC++ "memory kinds" feature to enable efficient communication of GPU-resident data. We show that on a number of large problems, symPACK outperforms comparable state-of-the-art GPU-capable Cholesky factorization codes by up to 14x on the NERSC Perlmutter supercomputer.

Bellavita, Julian↗

High-performance computing in water resources hydrodynamics

In this work, we present a vision of future water resources hydrodynamics codes that can fully utilize the strengths of modern high-performance computing. The advances to computing power, formerly driven by the improvement of central processing unit processors, now focus on parallel computing and, in particular, the use of graphics processing units (GPUs). However, this shift to a parallel framework requires refactoring the code to make efficient use of the data as well as changing even the nature of the algorithm that solves the system of equations. These concepts along with other features such as the precision for the computations, dry regions management, and input/output data are analyzed in this paper. A 2D multi-GPU flood code applied to a large-scale test case is used to corroborate our statements and ascertain the new challenges for the next-generation parallel water resources codes.

54 ENVIRONMENTAL SCIENCES↗

Performance Enhancement of APW+lo Calculations by Simplest Separation of Concerns

Full-potential linearized augmented plane wave (LAPW) and APW plus local orbital (APW+lo) codes differ widely in both their user interfaces and in capabilities for calculations and analysis beyond their common central task of all-electron solution of the Kohn–Sham equations. However, that common central task opens a possible route to performance enhancement, namely to offload the basic LAPW/APW+lo algorithms to a library optimized purely for that purpose. To explore that opportunity, we have interfaced the Exciting-Plus (“EP”) LAPW/APW+lo DFT code with the highly optimized SIRIUS multi-functional DFT package. This simplest realization of the separation of concerns approach yields substantial performance over the base EP code via additional task parallelism without significant change in the EP source code or user interface. We provide benchmarks of the interfaced code against the original EP using small bulk systems, and demonstrate performance on a spin-crossover molecule and magnetic molecule that are of size and complexity at the margins of the capability of the EP code itself.

Zhang, Long↗