Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “performance portable algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Evaluation of Portable Programming Models to Accelerate LArTPC Detector Simulations

The Liquid Argon Time Projection Chamber (LArTPC) technology is widely used in high energy physics experiments, including the upcoming Deep Underground Neutrino Experiment (DUNE). Accurately simulating LArTPC detector responses is essential for analysis algorithm development and physics model interpretations. Accurate LArTPC detector response simulations are computationally demanding, and can become a bottleneck in the analysis workflow. Compute devices such as General-Purpose Graphics Processing Units (GPGPUs) have the potential to substantially accelerate simulations compared to traditional CPU-only processing. The software development that requires often carries the cost of specialized code refactorization and porting to match the target hardware architecture. With the rapid evolution and increased diversity of the computer architecture landscape, it is highly desirable to have a portable solution that also maintains reasonable performance. We report our ongoing effort in evaluating Kokkos as a basis for this portable programming model using LArTPC simulations in the context of the Wire-Cell Toolkit, a C++ library for LArTPC simulations, data analysis, reconstruction and visualization.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Evaluation of Portable Programming Models to Accelerate LArTPC Detector Simulations

The Liquid Argon Time Projection Chamber (LArTPC) technology is widely used in high energy physics experiments, including the upcoming Deep Underground Neutrino Experiment (DUNE). Accurately simulating LArTPC detector responses is essential for analysis algorithm development and physics model interpretations. Accurate LArTPC detector response simulations are computationally demanding, and can become a bottleneck in the analysis workflow. Compute devices such as General-Purpose Graphics Processing Units (GPGPUs) have the potential to substantially accelerate simulations compared to traditional CPU-only processing. The software development for these compute accelerators often carries the cost of specialized code refactorization and porting to match the target hardware architecture. With the rapid evolution and increased diversity of the computer architecture landscape, it is highly desirable to have a portable solution that also maintains reasonable performance. We report our ongoing effort in evaluating Kokkos as a basis for this portable programming model using LArTPC simulations in the context of the Wire-Cell Toolkit, a C++ library for LArTPC simulations, data analysis, reconstruction and visualization.

47 OTHER INSTRUMENTATION↗

Module-OT: A Turnkey Solution for Securing Energy Systems

The Modular Security Apparatus for Managing Distributed Cryptography for Command-and-Control Messages on Operational Technology Networks (Module-OT) is a flexible and lightweight solution for grid-edge devices focusing on end-to-end security. It is a bump- in- the-wire solution acting as a secure conduit for data between devices or systems across a network. It improves the cybersecurity posture of DER systems by providing authentication, authorization, and data integrity to secure DER communications. Additionally, it performs key management, provides data security through whitelisting Internet Protocol addresses and ports, blocks unauthorized connections, controls user access, and allows serial or Ethernet connections for added flexibility. The core software is portable to various Linux-based operating systems and is developed to be customized by the developer and researcher communities. Module-OT has been validated in the lab, has been demonstrated at a 500-KW PV-plus-storage site, and has been proven ready to secure operational technology devices. Its core functionality meets current standards, including validation procedures of the NIST Cryptographic Algorithm Validation Program (CAVP) and the Federal Information Processing Standard (FIPS 140-2). Because of its capability to provide an accessible and affordable option for stepping up security across modern energy systems, Module-OT can serve as an effective technological option to standardize cybersecurity moving forward.

cryptography↗

ExaWind: Predictive Wind Energy Simulations

This presentation describes the ExaWind project and the team's progress in creating a suite of performance-portable codes designed for predictive simulations of wind farms on next-generation exascale-class supercomputers. Such simulations will require the resolution of scales spanning many orders of magnitude, from blade boundary layers to wind farm flow structures. In the U.S., the first exascale systems will be GPU accelerated, and different GPU manufacturers have been chosen for the different systems. At the heart of the ExaWind software is a hybrid-solver approach based on the codes Nalu-Wind and AMR-Wind, which are computational fluid dynamics solvers for the incompressible Navier-Stokes equations. Nalu-Wind is an unstructured-grid code used to resolve wind turbine geometry and blade boundary layers, whereas AMR-Wind is a structured-grid background solver for atmospheric turbulent flow and turbine wake propagation. The models are coupled with overset meshes and global linear systems are approximated through a loose-coupling algorithm. Results will include validation-quality high-fidelity simulations and strong/weak scaling results from the Summit supercomputer.

49 EE - Wind and Water Power Program - Wind (EE-4W↗

Data from: 'Abiotic influences on continuous conifer forest structure across a subalpine watershed'

This package archives the core data used for analysis and inference in 'Abiotic influences on continuous conifer forest structure across a subalpine watershed' (Worsham et al., 2025). All data were collected in the East River, Washington Gulch, Slate River, and Coal Creek watersheds of Colorado. In the paper, we quantified the relative influence of climate, topographic, edaphic, and geologic factors on conifer stand structure and composition, and their functional relationships, at the watershed scale. We used waveform LiDAR data to derive spatially continuous stand structure metrics. We fused these with a species-level classification map to estimate tree species abundance. We applied generalized additive and generalized boosted models to evaluate the covariability of structural and compositional metrics with abiotic variables. The package contains the essential products required for reproducing our analysis and the tables and figures reported in the publication. The products comprise four classes: (1) geospatial data, (2) tabular data used for inferential analysis, (3) tabular data describing analytical results and performance statistics, and (4) a data user guide. (1) includes discretized waveform LiDAR data, locations and attributes of individual tree crowns, sampling locations and domain boundaries, a canopy height model, and raster files of estimated forest structural and compositional metrics at 100 m grid scale. (2) includes all response and explanatory variable values applied in inferential models. Response variables include conifer forest stand density, basal area, 95th percentile height, quadratic mean diameter, and others. Explanatory variables include climatic water deficit, actual evapotranspiration, elevation, heat load, soil available water content, and others. (3) includes results of training and testing several individual tree detection (ITD) algorithms, as well as inferential modeling results. (4) is a PDF user guide for this data package, including detailed descriptions and data dictionaries for all files. The data package root contains 17 assets: 8 compressed tape archive (.tar.gz) files, 5 comma-separated values (.csv) files, 3 Geographic Tagged Image File Format (GeoTIFF) (.tif) files, and 1 Portable Document Format (.pdf) file. The compressed .tar.gz archives contain ESRI shapefiles (.shp) .tif, compressed LASer (.laz), and .csv files. The archives must first be decompressed using the widely distributed command-line software utility TAR. All other files, including constituent files within the .tar.gz archives, can be opened in the open-source R statistical computing environment. Alternatively, .csv files may also be read in any simple text editor software or Microsoft Excel. Geospatial files including .shp and .tif files can also be opened in GIS software, such as QGIS (open-source) or ESRI ArcGIS (proprietary). The .pdf Data User Guide can be read with Adobe Acrobat Reader or other compatible readers.

2018 NEON and 2025 CHESS Campaigns↗

Productive Programming of Distributed Systems with the SHAD C++ Library

High-performance computing (HPC) is often perceived as a matter of making large-scale systems (e.g., clusters) run as fast as possible, regardless the required programming effort. However, the idea of "bringing HPC to the masses" has recently emerged. Inspired by this vision, we have designed SHAD, the Scalable High-performance Algorithms and Data-structures library. SHAD is open source software, written in C++, for C++ developers. Unlike other HPC libraries for distributed systems, which rely on SPMD models, SHAD adopts a shared-memory programming abstraction, to make C++ programmers feel at home. Underneath, SHAD manages tasking and data-movements, moving the computation where data resides and taking advantage of asynchrony to tolerate network latency. At the bottom of his stack, SHAD can interface with multiple runtime systems: this not only improves developer’s productivity, by hiding the complexity of such software and of the underlying hardware, but also greatly enhance code portability. Thanks to its abstraction layers, SHAD can indeed target different systems, ranging from laptops to HPC clusters, without any need for modifying the user-level code. We have prototyped and open-sourced the implementation of (a subset of) the C++ standard library (STL) targeting multi-node HPC clusters. Our work allows plain STL-based C++ code to scale on HPC systems, with no need for rewriting the code to exploit the complex hardware. SHAD is available under Apache v2 License at https://github.com/pnnl/SHAD. In this paper we overview the design of the SHAD library, depicting its main components: runtime systems abstractions for tasking; parallel and distributed data-structures; STL-compliant interfaces and algorithms.

Castellana, Vito G.↗

A scalable matrix-free spectral element approach for unsteady PDE constrained optimization using PETSc/TAO

In this work, we provide a new approach for the efficient matrix-free application of the transpose of the Jacobian for the spectral element method for the adjoint-based solution of partial differential equation (PDE) constrained optimization. This results in optimizations of nonlinear PDEs using explicit integrators where the integration of the adjoint problem is not more expensive than the forward simulation. Solving PDE constrained optimization problems entails combining expertise from multiple areas, including simulation, computation of derivatives, and optimization. The Portable, Extensible Toolkit for Scientific computation (PETSc) together with its companion package, the Toolkit for Advanced Optimization (TAO), is an integrated numerical software library that contains an algorithmic/software stack for solving linear systems, nonlinear systems, ordinary differential equations, differential algebraic equations, and large-scale optimization problems and, as such, is an ideal tool for performing PDE-constrained optimization. This paper describes an efficient approach in which the software stack provided by PETSc/TAO can be used for large-scale nonlinear time-dependent problems. Time integration can involve a range of high-order methods, both implicit and explicit. The PDE-constrained optimization algorithm used is gradient-based and seamlessly integrated with the simulation of the physical problem.

97 MATHEMATICS AND COMPUTING↗

Development of NDE/NDT Tools for High-Volume & High-Speed Inspection of CFRP Structures in Automotive Manufacturing

Main advantages of the air-coupled ultrasound testing (ACUT) and electromagnetic testing (EMT) techniques for NDE of CFRP composites were non-contact sensing, scalability for high-speed inspection, cost-effectiveness, and non-hazardous operation. Despite these advantages, no systems that would satisfy the project requirements were commercially available. Hence, one of the major efforts of the Michigan State University (MSU) team at the initial stage of the project was to close this technological gap by developing, optimizing, and validating array sensors that would provide sufficient sensitivity, spatial coverage, and resolution for robust defect detection. Optimization of the ACUT and EMT sensor designs was performed using experimentally validated finite element models. Initial experiments using array probes were conducted on relatively flat CFRP samples. In parallel, the MSU team designed and assembled a portable platform with two robotic arms. The robots were equipped with newly designed sensors that enabled high-speed NDE of curved CFRP parts. Presently, the developed robotic platform can be used as a demo/template NDE system, which is easily adaptable to manufacturing environments and in-line NDE. The ACUT NDE system developed by the MSU team used a high-power 4-channel pulser receiver for parallel data acquisition. The array probes were designed by stacking commercially available ACUT transducers, which operated in the frequency range between 100 kHz and 500 kHz. MSU optimized the excitation procedure and developed wave focusing cones so as to reduce the crosstalk between the transducers and to provide higher pulse repletion frequency (PRF). The through-transmission (TT) and single-side access (SSA) inspection modes were successfully implemented. In the TT-ACUT, structural defects in CFRP were detected by passing ultrasonic waves through the test part. Hence, the ACUT transmitters and receivers needed to be placed on the opposite sides of the test part. In the SSA-ACUT, guided waves (GW) were excited in the test part using the transmitters and were sensed by the receivers from the same side. Multi-channel TT-ACUT and SSA-ACUT provided high-speed NDE, and were successfully validated on CFRP test samples with interlaminar delaminations and other embedded defects The EM techniques developed by the MSU team included: 1) eddy current testing (ECT), 2) capacitive imaging (CI) and hybrid dual-mode imaging. In ECT, structural damage was detected in CFRP using coils sensor arrays. In ECT, the excitation magnetic field is generated by passing an alternating current through a coil, which is placed above the test sample. The excitation field penetrates the conductive sample and induces the eddy currents in its transect. In turn, the eddy currents generate the reaction field, which affects the total field sensed by a coil. Hence, the presence of structural flaws will alter the eddy current flow and the picked-up signal. ECT is mostly sensitive to local changes of the electric conductivity of the test sample, and CFRPs are mostly conductive in the direction of carbon fibers. Hence, ECT was well suited for the detection of fiber damage/fiber irregularities. The MSU team developed printed circuit boards (PCB) with coil sensor arrays optimized for NDE of CFRP. Unlike most commercial probes designed for ECT of metallic structures, the MSU array probes were designed for operation in [1-10] MHz frequency range, which was optimal for low-conductive CFRP. Multiple sensing topologies (coil groups excitation/sensing arrangements) were implemented and successfully validated. Capacitive Imaging (CI) technique developed by MSU was complementary to ECT. In contrast to ECT, which was sensitive to local changes of the electrical conductivity, the CI was sensitive to local changes of the dielectric constant. Therefore, CI could provide information about matrix damage/matrix irregularities in CFRP. The MSU CI sensor arrays were made of multiple circular or rectangular open-plate capacitors printed on PCB. Sensors of this type are not commercially available. In addition to ECT and CI, the MSU team developed a hybrid (dual-mode) inductive/capacitive measurement technique that synergistically combined the benefits of inductive and capacitive sensing for rapid NDE of fiber reinforced polymer (FRP) composite structures. Fiber damage and fiber irregularities in FRPs were detected by configuring hybrid sensors as coil sensors. Similarly, matrix damage, matrix irregularities and interlaminar delaminations were detected by configuring hybrid sensors as capacitive sensors. ECT and CI were performed sequentially by means of electronic switching. Hence, eliminating the need for mounting two separate sensor arrays on the probe. Portable robotic platform was developed by MSU for multi-technique high-speed NDE of CFRP test parts. The platform had two 6-axis robots, which enabled inspection of curved parts in approximately a 6×6×6 ft 3 active scan area. On the software side, the MSU team integrated scripts for NDE hardware control with scripts for robot motion control. MSU also implemented automated path planning for the robots, reconstruction of part’s surfaces via stereovision, 3D rendering of inspection data, and image processing algorithms for enhanced defect detection. Automotive composite parts manufactured by Plasan Composites from Phase I were used to validate the ACUT and EMT techniques on representative testbeds. Among those parts were three X-braces for a Dodge Viper, one composite calibration plaque with known defects at known locations, and four other test sections, including sections from a front splitter, a corner section from a composite hood, and a high-pressure RTM panel made using non crimp fabric. Other test samples included CFRP and GFRP calibration plates with fiber/matrix defects fabricated at MSU/CVRC.

36 MATERIALS SCIENCE↗

Extending SEER for Extreme Heterogeneity

Heterogeneous and multi-device nodes are increasingly common in high-performance computing and data centers, yet existing programming models often lack simple, transparent, and portable support for these diverse architectures. The main contribution of this work is the development of novel SEER capabilities to address this challenge by providing a descriptive programming model that allows applications to seamlessly leverage heterogeneous nodes across various device types. SEER uses efficient memory management and can select the proper device[s] depending on the computational cost of the applications. This is completely transparent to the programmer, thereby providing a highly productive programming environment. Integrating extreme heterogeneity into the SEER library as shown with the use of NVIDIA and AMD GPUs simultaneously allows it to expand and exploit the performance possibilities. Our analysis based on the well-known Conjugate Gradient algorithm reports accelerations above 1.5 × on computationally demanding steps of such an algorithm by using both architectures simultaneously.

Teranishi, Keita [ORNL] (ORCID:0000000166472690)↗

Lifting the Garage Door on Spawn, An Open-Source BEM-Controls Engine

Spawn is the latest whole-building energy simulation engine developed by the US Department of Energy, National Labs and industry. Whereas EnergyPlus was designed as a successor to DOE-2, Spawn is not a direct successor of–nor is it intended as an imminent replacement for– EnergyPlus. Instead, Spawn reuses parts of EnergyPlus while supporting new use cases in HVAC and controls. Spawn is intended to provide several capabilities that significantly advance beyond EnergyPlus. It is intended to support the evaluation of novel HVAC and district energy systems in a more physically realistic way. Critically, it can model control in a physically realistic way, using portable specifications that can be compiled for execution on control platforms. Spawn is also intended to support co-simulation in an intrinsic way to enable integration with third-party models. This paper describes the software architecture of Spawn from model authoring to compilation and simulation. It explains how Spawn reuses the envelope and daylighting modules of EnergyPlus and couples them to HVAC and control models from the Modelica Buildings Library using the Functional Mockup Interface (FMI) standard. It presents a number of examples that: i) validate Spawn’s coupled simulation approach by comparing its results to those of EnergyPlus, ii) illustrate the Spawn methodology for modeling and simulating HVAC systems, and iii) evaluate the performance of Spawn’s Quantized State System (QSS) time integration algorithms

Wetter, Michael↗

Enabling Multireference Calculations on Multimetallic Systems with Graphic Processing Units

Modeling multimetallic systems efficiently enables faster prediction of desirable chemical properties and the design of new materials. This work describes an initial implementation for performing multireference wave function method localized active-space self-consistent field (LASSCF) calculations through the use of multiple graphics processing units (GPUs) to accelerate time-to-solution. Density fitting is leveraged to reduce memory requirements, and we demonstrate the ability to fully utilize multi-GPU compute nodes. Performance improvements of 5–10x in total application runtime were observed in LASSCF calculations for multimetallic catalyst systems up to 1200 AOs and an active space of (22e,40o) using up to four NVIDIA A100 GPUs. Furthermore, written with performance portability in mind, a comparable performance is also observed in early runs on the Aurora exascale system using Intel Max Series GPUs.

Algorithms↗

Efficient phase-space generation for hadron collider event simulation

We present a simple yet efficient algorithm for phase-space integration at hadron colliders. Individual mappings consist of a single t-channel combined with any number of s-channel decays, and are constructed using diagrammatic information. The factorial growth in the number of channels is tamed by providing an option to limit the number of s-channel topologies. We provide a publicly available, parallelized code in C++ and test its performance in typical LHC scenarios.

47 OTHER INSTRUMENTATION↗

Tensor Network Quantum Virtual Machine for Simulating Quantum Circuits at Exascale

The numerical simulation of quantum circuits is an indispensable tool for development, verification, and validation of hybrid quantum-classical algorithms intended for near-term quantum co-processors. The emergence of exascale high-performance computing (HPC) platforms presents new opportunities for pushing the boundaries of quantum circuit simulation. Here, we present a modernized version of the Tensor Network Quantum Virtual Machine (TNQVM) that serves as the quantum circuit simulation backend in the eXtreme-scale ACCelerator (XACC) framework. The new version is based on the scalable tensor network processing library ExaTN (Exascale Tensor Networks). It provides multiple configurable quantum circuit simulators that perform either an exact quantum circuit simulation via the full tensor network contraction or an approximate simulation via a suitably chosen tensor factorization scheme. Upon necessity, stochastic noise modeling from real quantum processors is incorporated into the simulations by modeling quantum channels with Kraus tensors. By combining the portable XACC quantum programming frontend and the scalable ExaTN numerical processing backend, we introduce an end-to-end virtual quantum development environment that can scale from laptops to future exascale platforms. We report initial benchmarks of our framework, which include a demonstration of the distributed execution, incorporation of quantum decoherence models, and simulation of the random quantum circuits used for the certification of quantum supremacy on Google’s Sycamore superconducting architecture.

Nguyen, Thien↗

A comparative study on deep learning models for condition monitoring of advanced reactor piping systems

Advanced nuclear reactors offer innovative applications due to their portability, reliability, resiliency, and high capacity factors. To operate them on a wider scale, reducing maintenance life-cycle costs while ensuring their integrity is essential. Autonomous operations in advanced nuclear reactors using augmented Digital Twin (DT) technology can serve as a cost-effective solution by increasing awareness about the system’s health. A key component of nuclear DT frameworks is the condition monitoring of safety systems, such as piping-equipment systems, which involves acquiring and monitoring the plant’s sensor data. Here, this research proposes a condition monitoring methodology utilizing deep learning algorithms, such as multilayer perceptions (MLP) and convolutional neural networks (CNNs), to detect degradation and its severity in nuclear piping-equipment systems. Sensor signals are processed to obtain the power spectral density and the Short-Time Fourier transform, and feature extraction methodologies are proposed to develop degradation-sensitive data repositories. The performance of MLP, one-dimensional (1D) CNN, and 2D CNN within the proposed condition monitoring framework is compared using a finite element model of a 3D piping system subjected to seismic loads as the application case study. Various approaches, such as dropout, k-Fold validation, regularization, and early stopping of training the network, are investigated to avoid overfitting the models to the input sensor data. The predictive capability and computational capacity of the deep learning algorithms are also compared to detect degradation in the Z-pipe system of the Experimental Breeder Reactor II (EBRII). The Z-pipe system is subjected to harmonic excitations that represent normal operating loads, such as pump-induced vibrations. The findings of the study indicate that the proposed artificial intelligence (AI)-driven condition monitoring framework demonstrates superior prediction accuracies with a 2D CNN, whereas the MLP exhibits higher computational efficiency.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Implementation of McMurchie–Davidson Algorithm for Gaussian AO Integrals Suited for SIMD Processors

We report an implementation of the McMurchie− Davidson evaluation scheme for 1- and 2-particle Gaussian AO integrals designed for processors with Single Instruction Multiple Data (SIMD) instruction sets. Like in our recent MD implementation for graphical processing units (GPUs) [Asadchev, A.; Valeev, E. F.. J. Chem. Phys. 2024, 160, 244109.], variable-sized batches of shellsets of integrals are evaluated at a time. By optimizing for the floating point instruction throughput rather than minimizing the number of operations, this approach achieves up to 50% of the theoretical hardware peak FP64 performance for many common SIMD-equipped platforms (AVX2, AVX512, NEON), which translates to speedups of up to 30 over the state-of-the-art one-shellset-at-a-time implementation of Obara−Saika-type schemes in Libint for a variety of primitive and contracted integrals. As with our previous work, we rely on the standard C++ programming language such as the std::simd standard library feature to be included in the 2026 ISO C++ standard without any explicit code generation to keep the code base small and portable. The implementation is part of the open source LibintX library freely available at https://github.com/ValeevGroup/libintx.

Basis sets↗

Advanced battery modeling for interfacial phenomena and optimal charging

Lithium ion batteries are one of the most promising energy storage systems for portable devices, transportation, and renewable grids. To meet the increasing requirements of these applications, higher energy density and areal capacity, long cycle life, fast charging rate and enhanced safety for lithium ion battery (LIBs) are urgently needed. To solve these challenges, the relevant physics at different length scale need to be understood. However, experimental study is time consuming and limited in small scale’s study. Modeling techniques provide us powerful tools to get a deep understanding of the relevant physics and find optimal solutions. This work focuses on studying the mechanism in advanced battery engineering techniques and developing a new charging algorithm by model-based optimization. The research topics are divided into six topics and each topic is reported as a form of journal publication. Paper Ⅰ provides a new aspect of how ALD coating improves the lithium ion diffusion at electrode particles. Paper Ⅱ explains the mechanisms by which 3D electrodes enhance battery performance and reveals guidelines for optimized 3D electrode designs by a 3D electrochemical-mechanical battery model. Paper Ⅲ investigates the electrolyte concentration impact on SEI layer growth and Li plating, especially under high charge rates. Paper Ⅳ proposes an optimized charging protocol for fast charging for reducing the charging time with minimal degradation. Paper Ⅴ reports a comprehensive degradation model for degradation estimation and life predication of energy storage system (ESS). Paper Ⅵ is a study of temperature-dependent state of charge (SOC) estimation for battery pack.

25 ENERGY STORAGE↗

PETSc/TAO Users Manual (Rev. 3.19)

This manual describes the use of the Portable, Extensible Toolkit for Scientific Computation (PETSc) and the Toolkit for Advanced Optimization (TAO) for the numerical solution of partial differential equations and related problems on high-performance computers. PETSc/TAO is a suite of data structures and routines that provide the building blocks for the implementation of large-scale application codes on parallel (and serial) computers. PETSc uses the MPI standard for all distributed memory communication. PETSc/TAO includes a large suite of parallel linear solvers, nonlinear solvers, time integrators, and opti mization that may be used in application codes written in Fortran, C, C++, and Python (via petsc4py; see Getting Started). PETSc provides many of the mechanisms needed within parallel application codes, such as parallel matrix and vector assembly routines. The library is organized hierarchically, enabling users to employ the level of abstraction that is most appropriate for a particular problem. By using techniques of object-oriented programming, PETSc provides enormous flexibility for users. PETSc is a sophisticated set of software tools; as such, for some users it initially has a much steeper learning curve than packages such as MATLAB or a simple subroutine library. In particular, for individuals without some computer science background, experience programming in C, C++, python, or Fortran and experience using a debugger such as gdb or lldb, it may require a significant amount of time to take full advantage of the features that enable efficient software use. However, the power of the PETSc design and the algorithms it incorporates may make the efficient implementation of many application codes simpler than “rolling them” yourself. For many tasks a package such as MATLAB is often the best tool; PETSc is not intended for the classes of problems for which effective MATLAB code can be written. There are several packages, built on PETSc, that may satisfy your needs without requiring directly using PETSc. We recommend reviewing these packages functionality before starting to code directly with PETSc. PETSc can be used to provide a “MPI parallel linear solver” in an otherwise sequential, or OpenMP parallel code. This approach cannot provide extremely large improvements in the application time by utilizing large numbers of MPI processes but can still improve the performance. Certainly all parts of a previously sequential code need not be parallelized but the matrix generation portion must be parallelized to expect true scalability to large numbers of MPI processes. See PCMPI for details on how to utilize the PETSc MPI linear solver server. Since PETSc is under continued development, small changes in usage and calling sequences of routines will occur. PETSc has been supported for twenty-five years; see mailing list information on our website for information on contacting support.

97 MATHEMATICS AND COMPUTING↗

A Sparse Distributed Gigascale Resolution Material Point Method

In this paper, we present a four-layer distributed simulation system and its adaptation to the Material Point Method (MPM). The system is built upon a performance portable C++ programming model targeting major High-Performance-Computing (HPC) platforms. A key ingredient of our system is a hierarchical block-tile-cell sparse grid data structure that is distributable to an arbitrary number of Message Passing Interface (MPI) ranks. We additionally propose strategies for efficient dynamic load balance optimization to maximize the efficiency of MPI tasks. Our simulation pipeline can easily switch among backend programming models, including OpenMP and CUDA, and can be effortlessly dispatched onto supercomputers and the cloud. Finally, we construct benchmark experiments and ablation studies on supercomputers and consumer workstations in a local network to evaluate the scalability and load balancing criteria. We demonstrate massively parallel, highly scalable, and gigascale resolution MPM simulations of up to 1.01 billion particles for less than 323.25 seconds per frame with 8 OpenSSH-connected workstations.

97 MATHEMATICS AND COMPUTING↗