gismo-cloud-deployment
Tools for executing time-consuming tasks with developer-defined custom code blocks in parallel on the AWS EKS platform.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Tools for executing time-consuming tasks with developer-defined custom code blocks in parallel on the AWS EKS platform.
A major challenge in interpreting geophysical data is how to derive consistent three-dimensional (3D) earth models of different physical properties from spatially and temporally limited measurements. Joint inversion with cross-gradient constraints is an approach to find such models by imposing structural similarities between different physical parameters. We have developed a parallel distributed-memory joint inversion code for direct-current (DC) resistivity and traveltime data using the cross-gradient constraint on unstructured mesh. The code utilizes existing E4D framework for parallel forward simulation, distributed storage and computation of the Jacobian matrix of forward operator, and parallel execution of matrix-vector multiplication during inversion. Besides, the joint inversion is solved by nonlinear conjugate gradient algorithm parallelized for DC resistivity and traveltime data. The joint inversion capability of E4D was tested using synthetic data from cross-borehole DC resistivity and traveltime data. The results indicate that the shape and size of the anomalies from the joint inversion are more reliable than those from separate inversions.
We report on beam dynamics studies of transient beam loading effects in the 10 GeV EIC electron storage ring [1]. The studies are carried out with time-dependent Vlasov-Fokker-Planck simulations performed with the parallel, particle tracking code SPACE [2], which allows to follow self-consistently the dynamics of h bunches, where h in the number of RF buckets, in arbitrary multi-bunch configurations. The specific goal of the numerical simulations is to determine stable RF cavity settings under heavy beam loading. We also study the option to operate with a passive, third-harmonic cavity (3HC) system for bunch lengthening, addressing both stability and the performance limitation due to a gap in the uniform filling pattern for ion clearing.
The goal of this milestone is to demonstrate readiness of the Sandia Parallel Aerodynamics Reentry Code (SPARC) and the Electro-Magnetic Plasma in Realistic Environments (EMPIRE) code for use on the forthcoming El Capitan exascale platform, which will enable simulation of weapons-relevant problems at unprecedented scale and fidelity.
The physical and chemical transformation during atmospheric transport of radionuclides released into the environment has the possibility of impacting consequence modeling results. Accordingly, this report identifies physical and chemical transformations that may occur following release of chemically reactive radioactive species, how those transformations may affect modeling of consequences of release to the atmosphere and identifies current capabilities – in both MACCS and other state-of-practice atmospheric transport and dispersions models– to model those transformations. It was found that the inclusion of physical and chemical transformations is currently very limited in current state-of-practice codes for atmospheric dispersion of radionuclides. State-of-practice atmospheric dispersion codes appear to be typically limited to simulating either physical-chemical transformations or radioactive transformations, but not both. A state-of-practice atmospheric dispersion code capable of performing parallel physical, chemical, and radioactive transformation was not identified. A few atmospheric dispersion codes capable of modeling physical and chemical transport of specific species such as tritium or uranium hexafluoride were identified. Consequently, there is currently no information available that clearly suggests updates to the MACCS code are needed to bring it up to state-of-practice. However, investigations concluded that the MACCS computational framework can currently accommodate multiple physical/chemical forms in one simulation. Additionally, with some major assumptions, the computational framework in MACCS can accommodate parallel physical-chemical and radioactive transformations.
For decades, Los Alamos National Laboratory has been at the forefront of neutron transport methods research and code development. One such code is PARTISN, the LANL parallel time-dependent discrete ordinate neutron transport code. In this presentation, we describe the various research efforts currently underway by the PARTISN and other code teams. Some examples of current research are a block automated mesh refinement scheme, the application of tensor trains to the discretized neutron transport equation, and GPU code porting. The block automated mesh refinement scheme uses cross section information to refine and coarsen the solution mesh to improve time to solution and reduce memory. The tensor train approach expresses discretized transport operators as tensor products of vectors and matrices to compress the size of linear systems being solved by transport codes. Rather than relying on matrix-free methods such as the transport sweep, we have access to an operator that can be inverted, reshaped, or manipulated algebraically. Finally, we describe how PARTISN is used, what problems we are looking to solve, and what the future holds for neutron transport at LANL. In addition to this, we briefly describe the various research efforts in other particle transport teams using both deterministic and Monte Carlo methods. In the presentation, we list possible opportunities for collaboration between the laboratory and faculty and students.
The Message Passing Interface (MPI) standard plays a crucial role in enabling scientific applications for parallel computing and is an essential component in high-performance computing (HPC). However, implementing MPI code manually—especially applying a proper domain decomposition and communication pattern—is a challenging and error-prone task. We present ChatMPI, an AI assistant for MPI parallelization of sequential C codes. In our analysis, we focus on testing six essential HPC workloads, which are based on Basic Linear Algebra Subprograms levels 1, 2, and 3 as well as sparse, stencil, and iterative operations. We analyze the process of creating ChatMPI by using the ChatHPC library. This lightweight large language model (LLM)–based infrastructure enables HPC experts to efficiently create and supervise trustworthy AI capabilities for critical HPC software tasks. We study the data required for training (fine-tuning) ChatMPI to generate parallel codes that not only use MPI syntax correctly but also apply HPC techniques to reduce memory communication and maximize performance by using proper work decomposition. With a relatively small training dataset composed of a few dozen prompts and fewer than 15 minutes of fine-tuning on one node equipped with two NVIDIA H100 GPUs, ChatMPI elevates trustworthiness for MPI code generation of current LLMs (e.g., Code Llama, ChatGPT-4o and ChatGPT 5). Additionally, we evaluate the performance of the MPI codes generated by ChatMPI in comparison with the ones generated by ChatGPT-4o and ChatGPT-5. The codes generated by ChatMPI provide up to a 4 × boost in performance by using better problem decomposition, communication patterns, and HPC techniques (e.g., communication avoiding).
The extension of the capabilities of the pin-level nuclear reactor core thermal-hydraulics (T/H) code ESCOT to analyze hexagonal fueled cores and its performance are presented. ESCOT is an accurate yet fast core thermal-hydraulics solution aiming at high-fidelity and high-resolution multi-physics core analysis in the framework of massively parallel computing platforms. Its algorithm solution is based on the four-equation drift-flux model for two-phase calculations, these are numerically solved by applying the Finite Volume Method (FVM) and the Semi-Implicit Method for Pressure-Linked Equation (SIMPLE)-like algorithm in a staggered grid system. Constitutive models such as turbulent mixing, pressure drop, and vapor generation are employed to simulate key phenomena in subchannel-scale analysis. ESCOT is parallelized by a double (radial and axial) domain decomposition that enables its highly parallelized execution. The coupling of the code with the neutronics whole core solver for hexagonal geometries nTRACER is described. The newly implemented ESCOT features are validated by comparing single assembly and full core steady state nTRACER-ESCOT solutions with nTRACER standalone internal one-dimensional T/H solver results. The validation problems are based on the VVER 440 and VVER 1000 cores. ESCOT results show differences within an acceptable range with respect to the simple 1D nTRACER built-in solver. (authors)
A parallel, relativistic, three-dimensional particle-in-cell code SPACE has been developed for the simulation of electromagnetic fields, relativistic particle beams, and plasmas. In addition to the standard second-order Particle-in-Cell (PIC) algorithm, SPACE includes efficient novel algorithms to resolve atomic physics processes such as multi-level ionization of plasma atoms, recombination, and electron attachment to dopants in dense neutral gases. SPACE also contains a highly adaptive particle-based method, called Adaptive Particle-in-Cloud (AP-Cloud), for solving the Vlasov-Poisson problems. It eliminates the traditional Cartesian mesh of PIC and replaces it with an adaptive octree data structure. The code's algorithms, structure, capabilities, parallelization strategy, and performance have been discussed. Additionally, typical examples of SPACE applications to accelerator science and engineering problems are described.
We report that TUMME is a program for assembling and solving master equations for gas-phase chemical kinetics based on chemically significant eigenmodes. TUMME has interfaces to the Gaussian, Polyrate, and/or MSTor output files that allow the master equation code to obtain the microcanonical flux coefficients needed for the coefficient matrix of the master equation. The flux coefficients for reactions with barriers can be calculated by multi-structural variational transition state theory with small-curvature tunneling (MS-VTST/SCT) or by simpler approximations to this such as conventional transition state theory without tunneling (also called RRKM theory). The flux coefficients for barrierless reactions are provided by a hard-sphere model. TUMME is written in double precision with Python 3; quadruple and octuple precision are also available for some subtasks in C++. The Python code can run in serial or parallel (MP or MPI), and the C++ code can run on a single processor or on multiple processors with OpenMP.
The degradation of antenna performance during hypersonic re-entry is a well known phenomenon that can lead to complete radio blackout. Recent additions to the Empire code establish it as a tool for the study and analysis of the problem. Coupling to the Sandia Parallel Aerodynamics and Reentry Code (SPARC) enables the electromagnetic analysis of realistic re-entry plasma profiles. The geometric flexibility afforded by both Empire and SPARC allow the consideration of arbitrary vehicle and antenna configurations. We have used this tool to study antenna performance during re-entry when the boundary layer becomes turbulent. A concise description of line-of-sight transmissions, which employs advanced statistical methods, was developed. New insights into the low altitude reflectometer readings of RAM-C2 are offered. Techniques for the reconstruction of the re-entry plasma profile from reflectometer data were explored.
Quantum low-density parity-check codes are promising candidates towards scalable fault-tolerant quantum computation. Among these, bivariate bicycle (BB) codes offer superior encoding rates and large code distance compared to surface codes. However, their requirement on long-range stabilizer measurements poses significant challenges for implementation on realistic hardware with limited connectivity, such as superconducting circuit platforms. In this work, we introduce a novel hardware-software co-design that leverages a programmable communication network architecture to address these limitations. Our approach utilizes a 2D toric network of oscillators as a flexible communication fabric linking qubits at each site. Such architecture significantly reduces the number of long-range couplers required from O ( n ) to O (√ n ). Dual-rail qubits, along with native gates including Swap-Wait-Swap gates and beamsplitter SWAPs, ensure that long-range two-qubit gates can be executed with high fidelity and low latency. To further enhance performance, our qubit layout and routing algorithm utilize symmetries of the codes and enable maximum parallelism for long-range two-qubit gates, maintaining a low syndrome extraction cycle duration and scalability over the code length. We perform circuit-level simulation with realistic noise modeling based on experimental hardware parameters, observing an logical error rate per logical qubit per cycle of 3.06% for [[18,4,4]] BB code, 2.6× less than the existing experimental result. These findings provide a practical roadmap and identify key technological advancements needed to achieve low-overhead fault-tolerant quantum computing at scale.
In this paper, we report reimplementation of the core algorithms of relativistic coupled cluster theory aimed at modern heterogeneous high-performance computational infrastructures. The code is designed for parallel execution on many compute nodes with optional GPU coprocessing, accomplished via the new ExaTENSOR back end. The resulting ExaCorr module is primarily intended for calculations of molecules with one or more heavy elements, as relativistic effects on the electronic structure are included from the outset. In the current work, we thereby focus on exact two-component methods and demonstrate the accuracy and performance of the software. The module can be used as a stand-alone program requiring a set of molecular orbital coefficients as the starting point, but it is also interfaced to the DIRAC program that can be used to generate these. We therefore also briefly discuss an improvement of the parallel computing aspects of the relativistic self-consistent field algorithm of the DIRAC program.
Numerical relativity is central to the investigation of astrophysical sources in the dynamical and strong-field gravity regime, such as binary black hole and neutron star coalescences. Current challenges set by gravitational-wave and multimessenger astronomy call for highly performant and scalable codes on modern massively parallel architectures. We present GR-Athena++, a general-relativistic, high-order, vertex-centered solver that extends the oct-tree, adaptive mesh refinement capabilities of the astrophysical (radiation) magnetohydrodynamics code Athena++. To simulate dynamical spacetimes, GR-Athena++ uses the Z4c evolution scheme of numerical relativity coupled to the moving puncture gauge. We demonstrate stable and accurate binary black hole merger evolutions via extensive convergence testing, cross-code validation, and verification against state-of-the-art effective-one-body waveforms. GR-Athena++ leverages the task-based parallelism paradigm of Athena++ to achieve excellent scalability. We measure strong-scaling efficiencies above 95% for up to ~1.2 × 10 4 CPUs and excellent weak scaling is shown up to ~10 5 CPUs in a production binary black hole setup with adaptive mesh refinement. GR-Athena++ thus allows for the robust simulation of compact binary coalescences and offers a viable path toward numerical relativity at exascale.
The MPACT code is designed to perform high-fidelity light-water reactor (LWR) analysis using whole-core pin-resolved neutron transport calculations on modern parallel-computing hardware. The code consists of several libraries which provide the functionality necessary to solve steady-state eigenvalue problems. Several transport capabilities are available within MPACT including both 2-D and 3-D Method of Characteristics (MOC). A three-dimensional whole core solution based on the 2D-1D solution method provides the capability for full core depletion calculations.
The study of hypersonic flows and their underlying aerothermochemical reactions is particularly important in the design and analysis of vehicles exiting and reentering Earth's atmosphere. Computational physics codes can be employed to simulate these phenomena; however, code verification of these codes is necessary to certify their credibility. To date, few approaches have been presented for verifying codes that simulate hypersonic flows, especially flows reacting in thermochemical nonequilibrium. In this work, we present our code-verification techniques for verifying the spatial accuracy and thermochemical source term in hypersonic reacting flows in thermochemical nonequilibrium. Additionally, we demonstrate the effectiveness of these techniques on the Sandia Parallel Aerodynamics and Reentry Code (SPARC).
Siera/SolidMechanics (Sierra / SM) is a Lagrangian, three-dimensional code for finite element analysis of solids and structures. It provides capabilities for explicit dynamic, implicit quasistatic and dynamic analyses. The explicit dynamics capabilities allow for the efficient and robust solution of models with extensive contact subjected to large, suddenly applied loads. For implicit problems, Sierra / SM uses a multi-level iterative solver, which enables it to effectively solve problems with large deformations, nonlinear material behavior, and contact. Sierra / SM has a versatile library of continuum and structural elements, and a large library of material models. The code is written for parallel computing environments enabling scalable solutions of extremely large problems for both implicit and explicit analyses. It is built on the SIERRA Framework, which facilitates coupling with other SIERRA mechanics codes . This document describes the functionality and input syntax for Sierra/SM.
Sierra / SolidMechanics (Sierra / SM) is a Lagrangian, three-dimensional code for finite element analysis of solids and structures. It provides capabilities for explicit dynamic, implicit quasistatic and dynamic analyses. The explicit dynamics capabilities allow for the efficient and robust solution of models with extensive contact subjected to large, suddenly applied loads. For implicit problems, Sierra / SM uses a multi-level iterative solver, which enables it to effectively solve problems with large deformations, nonlinear material behavior, and contact. Sierra / SM has a versatile library of continuum and structural elements, an d a large library of material models. The code is written for parallel computing environments enabling scalable solutions of extremely large problems for both implicit and explicit analyses. It is built on the SIERRA Framework, which facilitates coupling with other SIERRA mechanics codes . This document describes the functionality and input syntax for Sierra / SM.