Engineering Papers⌕ Search

Engineering topics

Bertoni, Colleen

Publications and source records attributed to Bertoni, Colleen.

The Effective Fragment Molecular Orbital Method: Achieving High Scalability and Accuracy for Large Systems

The effective fragment molecular orbital (EFMO) method has been developed to predict the total energy of a very large molecular system accurately (with respect to the underlying quantum mechanical method) and efficiently by taking advantage of the locality of strong chemical interactions and employing a two-level hierarchical parallelism. The accuracy of the EFMO method is partly attributed to the accurate and robust intermolecular interaction prediction between distant fragments, in particular, the many-body polarization and dispersion effects, which require the generation of static and dynamic polarizability tensors by solving the coupled perturbed Hartree–Fock (CPHF) and time-dependent HF (TDHF) equations, respectively. Solving the CPHF and TDHF equations is the main EFMO computational bottleneck due to the inefficient (serial) and I/O-intensive implementation of the CPHF and TDHF solvers. In this work, the efficiency and scalability of the EFMO method are significantly improved with a new CPU memory-based implementation for solving the CPHF and TDHF equations that are parallelized by either message passing interface (MPI) or hybrid MPI/OpenMP. Here, the accuracy of the EFMO method is demonstrated for both covalently bonded systems and noncovalently bound molecular clusters by systematically examining the effects of basis sets and a key distance-related cutoff parameter, R cut . R cut determines whether a fragment pair (dimer) is treated by the chosen ab initio method or calculated using the effective fragment potential (EFP) method (separated dimers). Decreasing the value of Rcut increases the number of separated (EFP) dimers, thereby decreasing the computational effort. It is demonstrated that excellent accuracy (<1 kcal/mol error per fragment) can be achieved when using a sufficiently large basis set with diffuse functions coupled with a small R cut value. With the new parallel implementation, the total EFMO wall time is substantially reduced, especially with a high number of MPI ranks. Given a sufficient workload, nearly ideal strong scaling is achieved for the CPHF and TDHF parts of the calculation. For the first time, EFMO calculations with the inclusion of long-range polarization and dispersion interactions on a hydrated mesoporous silica nanoparticle with explicit water solvent molecules (more than 15k atoms) are achieved on a massively parallel supercomputer using nearly 1000 physical nodes. In addition, EFMO calculations on the carbinolamine formation step of an amine-catalyzed aldol reaction at the nanoscale with explicit solvent effects are presented.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

The General Atomic and Molecular Electronic Structure System (GAMESS): Novel Methods on Novel Architectures

The primary focus of GAMESS over the last 5 years has been the development of new high-performance codes that are able to take effective and efficient advantage of the most advanced computer architectures, both CPU and accelerators. These efforts include employing density fitting and fragmentation methods to reduce the high scaling of well-correlated (e.g., coupled-cluster) methods as well as developing novel codes that can take optimal advantage of graphical processing units and other modern accelerators. Because accurate wave functions can be very complex, an important new functionality in GAMESS is the quasi-atomic orbital analysis, an unbiased approach to the understanding of covalent bonds embedded in the wave function. Finally, best practices for the maintenance and distribution of GAMESS are also discussed.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

HIPLZ: Enabling performance portability for exascale systems

While heterogeneous computing has emerged as a dominant trend in current and future High-Performance Computing (HPC) systems, it is also widely recognized that this shift has led to increased software complexity due to a proliferation of programming systems for different heterogeneous processors. One such example is the Heterogeneous-Compute Interface for Portability from AMD (HIP ), which is composed of a C Runtime API and C++ Kernel Language. Many HPC applications will likely use HIP on future exascale systems (e.g., Frontier and El Capitan), but HIP currently only targets AMD and NVIDIA processors. This limitation creates challenges for users who would also like to run their applications on exascale systems based on other architectures (e.g., Aurora, which is based on Intel hardware) that are currently not targeted by HIP . In this paper, we introduce the design and implementation of HIPLZ , a compiler and runtime system that uses the Intel Level Zero API to support HIP on Intel GPU architectures. We discuss the design of HIPLZ , derived from HIPCL (an implementation of HIP on top of OpenCL ), and portability issues that occur from using the Level Zero runtime as a backend. We evaluate our implementation by running several performance benchmarks and mini-apps written in HIP on Intel architectures using HIPLZ . Our results show that this approach provides competitive performance relative to Intel's OpenCL implementations on Intel Gen9 and UHD Graphics 770 GPUs, while providing good coverage of features needed by HPC applications. Overall, this approach is a promising demonstration of enabling performance portability for exascale systems.

97 MATHEMATICS AND COMPUTING↗

Porting fragmentation methods to GPUs using an OpenMP API: Offloading the resolution-of-the-identity second-order Møller–Plesset perturbation method

Here, using an OpenMP Application Programming Interface, the resolution-of-the-identity second-order Møller–Plesset perturbation (RI-MP2) method has been off-loaded onto graphical processing units (GPUs), both as a standalone method in the GAMESS electronic structure program and as an electron correlation energy component in the effective fragment molecular orbital (EFMO) framework. First, a new scheme has been proposed to maximize data digestion on GPUs that subsequently linearizes data transfer from central processing units (CPUs) to GPUs. Second, the GAMESS Fortran code has been interfaced with GPU numerical libraries (e.g., NVIDIA cuBLAS and cuSOLVER) for efficient matrix operations (e.g., matrix multiplication, matrix decomposition, and matrix inversion). The standalone GPU RI-MP2 code shows an increasing speedup of up to 7.5× using one NVIDIA V100 GPU with one IBM 42-core P9 CPU for calculations on fullerenes of increasing size from 40 to 260 carbon atoms using the 6-31G(d)/cc-pVDZ-RI basis sets. A single Summit node with six V100s can compute the RI-MP2 correlation energy of a cluster of 175 water molecules using the correlation consistent basis sets cc-pVDZ/cc-pVDZ-RI containing 4375 atomic orbitals and 14 700 auxiliary basis functions in ~0.85 h. In the EFMO framework, the GPU RI-MP2 component shows near linear scaling for a large number of V100s when computing the energy of an 1800-atom mesoporous silica nanoparticle in a bath of 4000 water molecules. The parallel efficiencies of the GPU RI-MP2 component with 2304 and 4608 V100s are 98.0% and 96.1%, respectively.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗