Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “C CODES”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Strategic Petroleum Reserve Cavern Leaching Monitoring CY24

This report provides an analysis of the effects of raw water leaching at the U.S. SPR using the Sandia Solution Mining Code (SANSMIC). A new version of the code has been established in the past year (implemented in Python and C++ code) and was used for this annual report. Additionally, the setup, running, and post-processing of leaching modeling has now been streamlined into a single Python notebook environment for each cavern run.

58 GEOSCIENCES↗

QSpace - An open-source tensor library for Abelian and non-Abelian symmetries

This is the documentation for the tensor library QSpace (v4.0), a toolbox to exploit ‘quan tum symmetry spaces’ in tensor network states in the quantum many-body context. QSpace permits arbitrary combinations of symmetries including the abelian symmetries $\mathbb{Z}_n$ and U(1), as well as all non-abelian symmetries based on the semisimple classical Lie algebras: A n , B n , C n , and D n , or respectively, the special unitary group SU(n), the odd orthogonal group SO(2n+1), the symplectic group Sp(2n), and the even orthogonal group SO(2n). The code (C++ embedded via the MEX interface into Matlab) is available open source as of QSpace v4.0 on bitbucket under the Apache 2.0 license. QSpace is designed as a bottom-up approach for non-abelian symmetries. It starts from the defining representation and the respective Lie algebra. By explicitly comput ing and tabulating generalized Clebsch-Gordan coefficient tensors, QSpace is versatile in the type of operations that it can perform across all symmetries. At the level of an ap plication, much of the symmetry-related details are hidden within the QSpace C++ core libraries. Hence when developing tensor network algorithms with QSpace, these can be coded (nearly) as if there are no symmetries at all, despite being able to fully exploit general non-abelian symmetries.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

ChatMPI: LLM-Driven MPI Code Generation for HPC Workloads

The Message Passing Interface (MPI) standard plays a crucial role in enabling scientific applications for parallel computing and is an essential component in high-performance computing (HPC). However, implementing MPI code manually—especially applying a proper domain decomposition and communication pattern—is a challenging and error-prone task. We present ChatMPI, an AI assistant for MPI parallelization of sequential C codes. In our analysis, we focus on testing six essential HPC workloads, which are based on Basic Linear Algebra Subprograms levels 1, 2, and 3 as well as sparse, stencil, and iterative operations. We analyze the process of creating ChatMPI by using the ChatHPC library. This lightweight large language model (LLM)–based infrastructure enables HPC experts to efficiently create and supervise trustworthy AI capabilities for critical HPC software tasks. We study the data required for training (fine-tuning) ChatMPI to generate parallel codes that not only use MPI syntax correctly but also apply HPC techniques to reduce memory communication and maximize performance by using proper work decomposition. With a relatively small training dataset composed of a few dozen prompts and fewer than 15 minutes of fine-tuning on one node equipped with two NVIDIA H100 GPUs, ChatMPI elevates trustworthiness for MPI code generation of current LLMs (e.g., Code Llama, ChatGPT-4o and ChatGPT 5). Additionally, we evaluate the performance of the MPI codes generated by ChatMPI in comparison with the ones generated by ChatGPT-4o and ChatGPT-5. The codes generated by ChatMPI provide up to a 4 × boost in performance by using better problem decomposition, communication patterns, and HPC techniques (e.g., communication avoiding).

Valero Lara, Pedro [ORNL] (ORCID:0000000214794310)↗

Orca2

SAND2024-03236O Orca2 is a trajectory design and optimization code for constructing and solving problems numerically defined on homogeneous manifolds. Orca2 solves numerical trajectory design and optimization problems defined on Lie-type domains and operates primarily as a library requiring a user to write additional C++ code to leverage it. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Sparapany, Michael↗

VPIC-Kokkos

VPIC is a particle-in-cell (PIC) code that solves the Maxwell-Boltzmann system of equations and can be used to simulate a wide variety of plasmas, including those that are fully kinetic and relativistic. Utilizing the Kokkos performance-portable framework, VPIC achieves high performance on multiple CPU and GPU architectures and is adaptable to future platforms with minimal developer effort. VPIC features very powerful input decks, allowing insertion of arbitrary C++ code for custom diagnostics, boundary conditions, and additional physics models.

Luedtke, Scott↗

iharm3D: Vectorized General Relativistic Magnetohydrodynamics

iharm3D is an open-source C code for simulating black hole accretion systems in arbitrary stationary spacetimes using ideal general-relativistic magnetohydrodynamics (GRMHD). It is an implementation of the HARM (“High Accuracy Relativistic Magnetohydrodynamics”) algorithm outlined in Gammie et al. (2003) with updates as outlined in McKinney & Gammie (2004) and Noble et al. (2006). The code is most directly derived from Ryan et al. (2015) but with radiative transfer portions removed. HARM is a conservative finite-volume scheme for solving the equations of ideal GRMHD, a hyperbolic system of partial differential equations, on a logically Cartesian mesh in arbitrary coordinates.

79 ASTRONOMY AND ASTROPHYSICS↗

High-Level Synthesis of Parallel Specifications Coupling Static and Dynamic Controllers

The increased need for efficient ways to implement domain-specific accelerators is driving design methodologies towards the use of abstractions higher than the Register Transfer Level (RTL). In this scenario, High Level Synthesis (HLS) plays a significant role by enabling the automatic generation of custom hardware accelerators starting from high level descriptions (e.g., C code). Conventional HLS tools exploit parallelism mostly at the Instruction Level (ILP). They statically schedule the input specifications, and build centralized Finite State Machine (FSM) controllers. However, aggressive exploitation of ILP in many applications has diminishing returns and, usually, centralized approaches do not efficiently exploit coarser parallelism because FSMs are inherently serial. In this paper we present a HLS framework able to synthesize applications that, beside ILP, also expose Task Level Parallelism (TLP). An application can expose TLP through annotations that identify the parallel functions (i.e., tasks). To generate accelerators that efficiently execute concur- rent tasks, we need to solve several issues: devise a mechanism to support concurrent execution flows, exploit memory parallelism, and manage synchronization. To support concurrent execution flows, we introduce a novel adaptive controller. The adaptive controller is composed of a set of interacting control elements that independently manage the execution of a single operation or function call. These control elements check dependencies and resource constraints at runtime, enabling as soon as possible execution. To support parallel access to shared memories and synchronization, we introduce a novel Hierarchical Memory Interface (HMI). With respect to previous solutions, the proposed interface supports multi-ported memories and atomic memory operations, which commonly occur in parallel programming. Our framework can generate the hardware implementation of C functions by employing two different approaches, depending on its characteristics. If a function exposes TLP, then the framework generates hardware implementations based on the adaptive controller. Otherwise, the framework implements the function by exploiting a more conventional FSM approach, which is optimized for ILP exploitation. We evaluate our framework on a set of parallel applications, and show substantial performance improvements (average speedup of 4.7) with limited area over- heads (average area increase of 5.48 times).

Castellana, Vito G.↗

High-Level Synthesis of Parallel Specifications Coupling Static and Dynamic Controllers

The increased need for efficient ways to implement domain-specific accelerators is driving design methodologies towards the use of abstractions higher than the Register Transfer Level (RTL). In this scenario, High Level Synthesis (HLS) plays a significant role by enabling the automatic generation of custom hardware accelerators starting from high level descriptions (e.g., C code). Conventional HLS tools exploit parallelism mostly at the Instruction Level (ILP). They statically schedule the input specifications, and build centralized Finite State Machine (FSM) controllers. However, aggressive exploitation of ILP in many applications has diminishing returns and, usually, centralized approaches do not efficiently exploit coarser parallelism because FSMs are inherently serial. In this paper we present a HLS framework able to synthesize applications that, beside ILP, also expose Task Level Parallelism (TLP). An application can expose TLP through annotations that identify the parallel functions (i.e., tasks). To generate accelerators that efficiently execute concur- rent tasks, we need to solve several issues: devise a mechanism to support concurrent execution flows, exploit memory parallelism, and manage synchronization. To support concurrent execution flows, we introduce a novel adaptive controller. The adaptive controller is composed of a set of interacting control elements that independently manage the execution of a single operation or function call. These control elements check dependencies and resource constraints at runtime, enabling as soon as possible execution. To support parallel access to shared memories and synchronization, we introduce a novel Hierarchical Memory Interface (HMI). With respect to previous solutions, the proposed interface supports multi-ported memories and atomic memory operations, which commonly occur in parallel programming. Our framework can generate the hardware implementation of C functions by employing two different approaches, depending on its characteristics. If a function exposes TLP, then the framework generates hardware implementations based on the adaptive controller. Otherwise, the framework implements the function by exploiting a more conventional FSM approach, which is optimized for ILP exploitation. We evaluate our framework on a set of parallel applications, and show substantial performance improvements (average speedup of 4.7) with limited area over- heads (average area increase of 5.48 times).

Castellana, Vito G.↗

End-To-End Decentralized Transmission Line Protection in IBR-Dominated Weak Grids Using Interpretable Data-Driven Methods

Traditional transmission line protection relies on predictable synchronous-based fault signatures, which frequently fail under the non-standard, current-limited fault characteristics of Inverter-Based Resources (IBRs). This study investigates how to achieve secure, communication-free fault isolation in IBR-dominated weak grids without relying on opaque, computationally heavy "black-box" machine learning algorithms. To address this, we propose a novel, standalone, and inherently interpretable data-driven protection framework. Unlike centralized methods requiring multi-terminal communication, this decentralized approach relies solely on local measurements using a hierarchical linear-kernel Support Vector Machine (SVM). The methodology decomposes the protection task into four sequential stages that mimic traditional protection elements: fault detection and fault direction identification, fault type classification, zone classification, and location estimation. This multi-stage architecture allows for specialized feature engineering at each stage, combining high computational efficiency with logic traceability. The framework's end-to-end performance was validated via C-code and PSCAD/EMTDC co-simulation, utilizing a real-world utility network and an OEM black-box IBR model. The proposed relay achieves 97.2% overall accuracy and provides a reliable trip decision within a 2.5-cycle window. The results confirm 100% accuracy in fundamental fault detection, reliable zone selectivity across low to moderate fault resistances, and robust security against non-fault transients, proving its immediate viability for integration into commercial numerical relays.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Performant implementation of the atomic cluster expansion (PACE) and application to copper and silicon

The atomic cluster expansion is a general polynomial expansion of the atomic energy in multi-atom basis functions. Here we implement the atomic cluster expansion in the performant C++ code that is suitable for use in large-scale atomistic simulations. We briefly review the atomic cluster expansion and give detailed expressions for energies and forces as well as efficient algorithms for their evaluation. We demonstrate that the atomic cluster expansion as implemented in shifts a previously established Pareto front for machine learning interatomic potentials toward faster and more accurate calculations. Moreover, general purpose parameterizations are presented for copper and silicon and evaluated in detail. We show that the Cu and Si potentials significantly improve on the best available potentials for highly accurate large-scale atomistic simulations.

36 MATERIALS SCIENCE↗

Modeling of small tungsten dust grains in EAST tokamak with NDS-BOUT ++

In order to investigate the transport of small dusts as well as their evolution property along their trajectories, the NDS module is developed under the BOUT++ framework, a highly desirable C++ code package to perform parallel plasma fluid simulations with an arbitrary number of equations in three-dimensional curvilinear coordinates. Due to the severe dust ablation in fusion plasmas, the dust size would decrease from micrometer to nanometer, resulting in impurities. Small dusts in the simulations here are specified as tungsten spheres with the radii on or below the order of submicrometer. The Rayleigh limit is included in the charging process when the dust is ablated to the droplet phase. The simulation results from the NDS module show that a 200 nm radius spherical tungsten dust originated from upper divertor region of EAST Tokamak is ablated completely due to the intense heating from the incoming plasma inside the core region, well consistent with the CCD footage of EAST shot # 81459. Furthermore it is found that the magnetic field dominates the dust transport when the dust radius is below 100 nm during the ablation along the trajectory. Our simulations predict that a 10 nm radius spherical tungsten dust injected from the inner midplane is well constrained by the magnetic field, and it reaches the inner divertor target with a velocity on the order of km/s.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

3-center and 4-center 2-particle Gaussian AO integrals on modern accelerated processors

We report an implementation of the McMurchie–Davidson (MD) algorithm for 3-center and 4-center 2-particle integrals over Gaussian atomic orbitals (AOs) with low and high angular momenta l and varying degrees of contraction for graphical processing units (GPUs). This work builds upon our recent implementation of a matrix form of the MD algorithm that is efficient for GPU evaluation of 4-center 2-particle integrals over Gaussian AOs of high angular momenta (l ≥ 4) [A. Asadchev and E. F. Valeev, J. Phys. Chem. A 127, 10889–10895 (2023)]. The use of unconventional data layouts and three variants of the MD algorithm allow for the evaluation of integrals with double precision and sustained performance between 25% and 70% of the theoretical hardware peak. Performance assessment includes integrals over AOs with l ≤ 6 (a higher l is supported). Preliminary implementation of the Hartree–Fock exchange operator is presented and assessed for computations with up to a quadruple-zeta basis and more than 20 000 AOs. The corresponding C++ code is part of the experimental open-source LibintX library available at https://github.com/ValeevGroup/libintx.

Chemistry↗

Phoebe: a high-performance framework for solving phonon and electron Boltzmann transport equations

Understanding the electrical and thermal transport properties of materials is critical to the design of electronics, sensors, and energy conversion devices. Computational modeling can accurately predict material properties but, in order to be reliable, requires accurate descriptions of electron and phonon states and their interactions. While first-principles methods are capable of describing the energy spectrum of each carrier, using them to compute transport properties is still a formidable task, both computationally demanding and memory intensive, requiring integration of fine microscopic scattering details for estimation of macroscopic transport properties. To address this challenge, we present Phoebe—a newly developed software package that includes the effects of electron–phonon, phonon–phonon, boundary, and isotope scattering in computations of electrical and thermal transport properties of materials with a variety of available methods and approximations. This open source C++ code combines MPI-OpenMP hybrid parallelization with GPU acceleration and distributed memory structures to manage computational cost, allowing Phoebe to effectively take advantage of contemporary computing infrastructures. We demonstrate that Phoebe accurately and efficiently predicts a wide range of transport properties, opening avenues for accelerated computational analysis of complex crystals.

36 MATERIALS SCIENCE↗

Implementation and Synthesis of Math Library Functions

Achieving speed and accuracy for math library functions like exp, sin, and log is difficult. This is because low-level implementation languages like C do not help math library developers catch mathematical errors, build implementations incrementally, or separate high-level and low-level decision making. This ultimately puts development of such functions out of reach for all but the most experienced experts. To address this, we introduce MegaLibm, a domain-specific language for implementing, testing, and tuning math library implementations. MegaLibm is safe, modular, and tunable. Implementations in MegaLibm can automatically detect mathematical mistakes like sign flips via semantic wellformedness checks, and components like range reductions can be implemented in a modular, composable way, simplifying implementations. Once the high-level algorithm is done, tuning parameters like working precisions and evaluation schemes can be adjusted through orthogonal tuning parameters to achieve the desired speed and accuracy. MegaLibm also enables math library developers to work interactively, compiling, testing, and tuning their implementations and invoking tools like Sollya and type-directed synthesis to complete components and synthesize entire implementations. MegaLibm can express 8 state-of-the-art math library implementations with comparable speed and accuracy to the original C code, and can synthesize 5 variations and 3 from-scratch implementations with minimal guidance.

97 MATHEMATICS AND COMPUTING↗

Ume: Unstructured Mesh Explorations

Ume is an open-source collection of data structures for unstructured computational meshes and some simple algorithms that operate on them. These algorithms mimic the memory access patterns of a common class of operations found in several of the computational physics simulation codes developed at Los Alamos National Laboratory. The intent is that Ume can be used by hardware vendors to understand the memory traffic created by complex codes in a simplified environment, and to explore new means of optimization for that traffic. Ume is provided as a source-code C++ library and includes several applications that demonstrate the use of that library.

Henning, Paul↗

exawind-driver [SWR-23-10]

Exawind-driver is a C++ code that is part of the ExaWind software stack. It was designed to couple and drive hybrid-solver computational fluid dynamics (CFD) simulations where NREL's AMR-Wind (SWR-20-85) software and NREL's Nalu-Wind (SWR-20-27) CFD codes are run simultaneously and are two-way coupled via overset meshes and the TIOGA overset-mesh library.

Rood, Jonathan↗

kynema-driver [SWR-23-10]

Kynema-driver (FKA: exawind-driver) is a C++ code that is part of the kynema software stack. It is a driver for coupled Kynema-SGF/UGF simulations. It was designed to couple and drive hybrid-solver computational fluid dynamics (CFD) simulations where NLR'S kynema-sgf (SWR-20-85) software and NLR'S kynema-ugf (SWR-20-27) CFD codes are run simultaneously and are two-way coupled via overset meshes and the TIOGA overset-mesh library.

Rood, Jonathan↗

An implicit barotropic mode solver for MPAS-ocean using a modern Fortran solver interface

Here, we demonstrate use of a modern Fortran solver interface to manage solver algorithms for an implicit barotropic mode solver in the Model for Predictions Across Scales-Ocean (MPAS-O). ForTrilinos, a Fortran interface to Trilinos that contains a large collection of solver capabilities written in C++, has been implemented in MPAS-O to provide access to a suite of linear solver options. By virtue of the simplified wrapper and interface generator (SWIG) automation tool that generates modern Fortran interfaces to C++ code, we were able to implement the Fortran solver interface in MPAS-O using a familiar Fortran coding style while minimizing performance degradation. The ForTrilinos solver interface is written within MPAS-O’s time stepping modules as a subroutine in conjunction with MPAS-O code. Applied to an idealized ocean and a high-resolution realistic ocean test case, parallel performance of ForTrilinos solvers is examined. It is found that parallel scalability of the ForTrilinos solvers is highly dependent on the number of global synchronization points per solver iteration in each iterative solver algorithm. ForTrilinos solvers perform best compared to the Fortran hand-crafted (FHC) solver when the amount of work per processor is large enough. However, parallel scalability is better with the FHC solver and so when the work per core is modest FHC outperforms ForTrilinos. The intercomparison between the ForTrilinos and FHC solvers reveals that this performance hit in the ForTrilinos solver mostly comes from the global synchronization process, while suggesting that the matrix-vector multiplication process in the FHC solver needs to be optimized for better performance.

97 MATHEMATICS AND COMPUTING↗