Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Parallel application”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Towards Generic Parallel Programming in Computer Science Education with Kokkos

Parallel patterns, views, and spaces are promising abstractions to capture the programmer's intent as well as the contextual information that can be used by an underlying runtime to efficiently map software to parallel hardware. These abstractions can be valuable in cases where an algorithm must accommodate requirements of code and performance portability across hardware architectures and vendor programming models. Kokkos is a parallel programming model for host- and accelerator architectures that relies on these abstractions and targets these requirements. It consists of a pure C++ interface, a specification, and a programming library. The programming library exposes patterns and types and maps them to an underlying abstract machine model. The abstract machine model offers a generic view of parallel hardware. While Kokkos is gaining popularity in large-scale HPC applications at some DOE laboratories, we believe that the implemented concepts are of interest to a broader audience including academia as they may contribute to a generic, vendor, and architecture-independent education of parallel programming. In this work, we give an insight into the design considerations of this programming model and list important abstractions. Further, we document best practices obtained from giving virtual classes on Kokkos and give pointers to resources that the reader may consider valuable for a lecture on generic parallel programming for students with preexisting knowledge on this matter.

Ciesko, Jan↗

An E and B gyrokinetic simulation model for kinetic Alfvén waves in tokamak plasmas

The gyrokinetic particle simulation is a powerful tool for studies of transport, nonlinear phenomenon, and energetic particle physics in tokamak plasmas. While most gyrokinetic simulations make use of the scalar and vector potentials, a new model (GK-E&B) has been developed by using the E and B field in a general form and has been implemented in simulating kinetic Alfvén waves in uniform plasma. In our work, the Chen et al. GK-E&B model has been expressed, in general, tokamak geometry using the local orthogonal coordinates and general tokamak coordinates. Its reduction for uniform plasma is verified, and the numerical results show good agreement with the original work. The theoretical dispersion relation and numerical results in the local model in screw pinch geometry are also in excellent agreement. Numerical results show excellent performance in a realistic parameter regime of burning plasmas with high values of β/(M e k$^{2}_{⊥}$ρ$^{2}_{i}$), which is a challenge for traditional methods due to the “cancellation” problem. As one application, the GK-E&B model is implemented with kinetic electrons in the local single flux surface limit. With the matched International Tokamak Physics Activity-Toroidicity-induced Alfvén Eigenmodes parameters adopted, numerical results show the capability of the GK-E&B in treating the parallel electron Landau damping for realistic tokamak plasma parameters. As another application, the global GK-E&B model has been implemented with the dominant electron contribution in the cold electron limit. Its capability in simulating the finite E || due to the finite electron mass is demonstrated.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

How does ion temperature gradient turbulence depend on magnetic geometry? Insights from data and machine learning

Magnetic geometry has a significant effect on the level of turbulent transport in fusion plasmas. Here, we model and analyse this dependence using multiple machine learning methods and a dataset of >200 000 nonlinear gyrokinetic simulations of ion-temperature-gradient turbulence in diverse non-axisymmetric geometries. The dataset is generated using a large collection of both optimised and randomly generated stellarator equilibria. At fixed gradients and other input parameters, the turbulent heat flux varies between geometries by several orders of magnitude. Trends are apparent among the configurations with particularly high or particularly low heat flux. Regression and classification techniques from machine learning are then applied to extract patterns in the dataset. Due to a symmetry of the gyrokinetic equation, the heat flux and regressions thereof should be invariant to translations of the raw features in the parallel coordinate, similar to translation invariance in computer vision applications. Multiple regression models including convolutional neural networks (CNNs) and decision trees can achieve reasonable predictive power for the heat flux in held-out test configurations, with highest accuracy for the CNNs. Using Spearman correlation, sequential feature selection and Shapley values to measure feature importance, it is consistently found that the most important geometric lever on the heat flux is the flux surface compression in regions of bad curvature. The second most important geometric feature relates to the magnitude of geodesic curvature. These two features align remarkably with surrogates that have been proposed based on theory, while the methods here allow a natural extension to more features for increased accuracy. The dataset, released with this publication, may also be used to test other proposed surrogates, and we find that many previously published proxies do correlate well with both the heat flux and stability boundary.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

GridOPTICS/GridPACK

GridPACK is a software framework consisting of a set of modules designed to simplify the development of programs that model the power grid and run on parallel, high performance computing platforms. It also contains several fully developed applications, including powerflow, dynamic simulation, state estimation, Kalman filter analysis (dynamic state estimation), contingency analysis and real time path rating. These applications can be used either standalone or as components in more complicated workflows that combine several different types of application together. The framework modules are available as a combination of libraries and software templates and consist of components for setting up and distributing power grid networks, support for modeling the behavior of individual buses and branches in the network, converting the network models to the corresponding algebraic equations, and parallel routines for manipulating and solving large algebraic systems. The framework also contains a module for distributing tasks evenly amongst computing resources, even if individual tasks vary widely in their execution times. Additional modules support input and output, basic statistical analysis of contingency based calculations, distributed data structures, as well as basic profiling and error management.

Palmer, Bruce↗

TeraChem: A graphical processing unit-accelerated electronic structure package for large-scale ab initio molecular dynamics

TeraChem was born in 2008 with the goal of providing fast on-the-fly electronic structure calculations to facilitate ab initio molecular dynamics studies of large biochemical systems such as photoswitchable proteins and multichromophoric antenna complexes. Originally developed for videogaming applications, graphics processing units (GPUs) offered a low-cost parallel computer architecture that became more accessible for general-purpose GPU computing with the release of CUDA in 2007. The evaluation of the electron repulsion integrals (ERIs) is a major bottleneck in electronic structure codes and provides an attractive target for acceleration on GPUs. Thus, highly efficient routines for evaluation of and contractions between the ERIs and density matrices were implemented in TeraChem. Here, electronic structure methods were developed and implemented to leverage these integral contraction routines, resulting in the first quantum chemistry package designed from the ground up for GPUs. This GPU acceleration makes TeraChem capable of performing large-scale ground and excited state calculations in the gas and condensed phase. Today, TeraChem's speed forms the basis for a suite of quantum chemistry applications, including optimization and dynamics of proteins, automated and interactive chemical discovery tools, and large-scale nonadiabatic dynamics simulations.

74 ATOMIC AND MOLECULAR PHYSICS↗

Fragme∩t: An Open‐Source Framework for Multiscale Quantum Chemistry Based on Fragmentation

Fragment-based quantum chemistry offers a means to circumvent the nonlinear computational scaling of conventional electronic structure calculations, by partitioning a large calculation into smaller subsystems then considering the many-body interactions between them. Variants of this approach have been used to parameterize classical force fields and machine learning potentials, applications that benefit from interoperability between quantum chemistry codes. However, there is a dearth of software that provides interoperability yet is purpose-built to handle the combinatorial complexity of fragment-based calculations. To fill this void we introduce “Fragme∩t”, an open-source software application that provides a tool for community validation of fragment-based methods, a platform for developing new approximations, and a framework for analyzing many-body interactions. Fragme∩t includes algorithms for automatic fragment generation and structure modification, and for distance- and energy-based screening of the requisite subsystems. Checkpointing, database management, and parallelization are handled internally and results are archived in a portable database. Interfaces to various quantum chemistry engines are easy to write and exist already for Q-Chem, PySCF, xTB, Orca, CP2K, MRCC, Psi4, NWChem, GAMESS, and MOPAC. Applications reported here demonstrate parallel efficiencies around 96% on more than 1000 processors but also showcase that the code can handle large-scale protein fragmentation using only workstation hardware, all with a codebase that is designed to be usable by non-experts. Fragme∩t conforms to modern software engineering best practices and is built upon well established technologies including Python, SQLite, and Ray. The source code is available under the Apache 2.0 license.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

h5bench: A unified benchmark suite for evaluating HDF5 I/O performance on pre‐exascale platforms

Summary Parallel I/O is a critical technique for moving data between compute and storage subsystems of supercomputers. With massive amounts of data produced or consumed by compute nodes, high‐performant parallel I/O is essential. I/O benchmarks play an important role in this process; however, there is a scarcity of I/O benchmarks representative of current workloads on HPC systems. Toward creating representative I/O kernels from real‐world applications, we have created h5bench , a set of I/O kernels that exercise hierarchical data format version 5 (HDF5) I/O on parallel file systems in numerous dimensions. Our focus on HDF5 is due to the parallel I/O library's heavy usage in various scientific applications running on supercomputing systems. The various tests benchmarked in the h5bench suite include I/O operations (read and write), data locality (arrays of basic data types and arrays of structures), array dimensionality (one‐dimensional arrays, two‐dimensional meshes, three‐dimensional cubes), I/O modes (synchronous and asynchronous). In this paper, we present the observed performance of h5bench executed along several of these dimensions on existing supercomputers (Cori and Summit) and pre‐exascale platforms (Perlmutter, Theta, and Polaris). h5bench measurements can be used to identify performance bottlenecks and their root causes and evaluate I/O optimizations. As the I/O patterns of h5bench are diverse and capture the I/O behaviors of various HPC applications, this study will be helpful to the broader supercomputing and I/O community.

97 MATHEMATICS AND COMPUTING↗

Impact of Module Configuration on Lithium-Ion Battery Performance and Degradation: Part I. Energy Throughput, Voltage Spread, and Current Distribution

Batteries are commonly connected in series and parallel to create modules that fulfill the power and energy requirements of specific applications. However, conclusions about battery performance and degradation under different conditions, as well as predictive models, are often derived from single cell cycling results. In this study, we evaluate the performance of six different series-parallel configurations of commercial lithium nickel manganese cobalt cells over hundreds of cycles. Each cell within the modules was individually instrumented for voltage, current, and temperature monitoring. We quantified the impact of module configuration on overall energy throughput, the voltage spread among series-connected cells, and the current heterogeneity in parallel-connected cells. This module cycling study, one of the broadest reported to date, supports systematic evaluation of the performance trade-offs, pack penalty, and safety implications of different module configurations.

25 ENERGY STORAGE↗

Demystifying asynchronous I/O Interference in HPC applications

With increasing complexity of HPC workflows, data management services need to perform expensive I/O operations asynchronously in the background, aiming to overlap the I/O with the application runtime. However, this may cause interference due to competition for resources: CPU, memory/network bandwidth. The advent of multi-core architectures has exacerbated this problem, as many I/O operations are issued concurrently, thereby competing not only with the application but also among themselves. Furthermore, the interference patterns can dynamically change as a response to variations in application behavior and I/O subsystems (e.g. multiple users sharing a parallel file system). Without a thorough understanding, I/O operations may perform suboptimally, potentially even worse than in the blocking case. To fill this gap, here we investigate the causes and consequences of interference due to asynchronous I/O on HPC systems. Specifically, we focus on multi-core CPUs and memory bandwidth, isolating the interference due to each resource. Then, we perform an in-depth study to explain the interplay and contention in a variety of resource sharing scenarios such as varying priority and number of background I/O threads and different I/O strategies: sendfile, read/write, mmap/write underlining trade-offs. The insights from this study are important both to enable guided optimizations of existing background I/O, as well as to open new opportunities to design advanced asynchronous I/O strategies.

97 MATHEMATICS AND COMPUTING↗

The organic redox transistor for neuromorphic computing

Inspired by the in-memory computing architectures of biological systems, neuromorphic computing using crossbar arrays of artificial synapses based on non-volatile memory (NVM) devices with variable conductances has emerged as a new paradigm to enable massively parallel and ultra-low power computing hardware for data centric applications. Although inference has been demonstrated successfully using crossbars based on a variety of NMV technologies, efficient learning and scaling to large arrays (>10 6 elements) remains a challenge due to the synaptic elements' non-ideal electrical characteristics which degrades ANN accuracy. A further challenge is that in the conductive state memristors draw large currents >μA resulting in significant voltage drops in the interconnect wires and increased probability of failure in scaled arrays. We suggest the organic polymer redox transistor (RT) is an alternate approach that could solve many of these challenges, enabling both inference and parallel outer product updates, as recently demonstrated by Fuller et al. An RT consists of redox-active channel and gate electrodes in contact with a liquid or solid electrolyte. lon insertion through the electrolyte controls the channel electronic conductivity, while electron transfer through an external circuit maintains overall charge neutrality. Unlike a rechargeable battery, in the RT the voltage built-up across the electrolyte is kept to a minimum (typically <100 mV) by using the same material for the gate and channel. Elimination of the voltage offset simplifies integration of the RT into programmable arrays by enabling the use of various selectors. RTs based on inorganic and organic materials have been recently demonstrated with conductance tuning occurring at potentials of just a few mV and hundreds to thousands of linearly and symmetrically programmable conductance states, enabling near ideal accuracy in neural network simulations. Introduced in the 1980's, redox transistors with metallic gate electrodes and organic channel materials, also known as organic electrochemical transistors (OECTs), have been explored for a variety of applications such as chem- and bio-sensing, neural interfaces, and low cost printed circuits. A typical channel material for OECTs is the conducting polymer poly(3,4-ethylenedioxythiophene) doped with poly(styrene sulfonate) (PEDOT:PSS). PEDOT is a p-type semiconducting polymer with mobile positively charged polarons that hop chain-to-chain.

Talin, Albert Alec↗

Asynchronous and Load-Balanced Union-Find for Distributed and Parallel Scientific Data Visualization and Analysis

We present a novel distributed union-find algorithm that features asynchronous parallelism and k-d tree based load balancing for scalable visualization and analysis of scientific data. Applications of union-find include level set extraction and critical point tracking, but distributed union-find can suffer from high synchronization costs and imbalanced workloads across parallel processes. In this study, we prove that global synchronizations in existing distributed union-find can be eliminated without changing final results, allowing overlapped communications and computations for scalable processing. We also use a k-d tree decomposition to redistribute inputs, in order to improve workload balancing. We benchmark the scalability of our algorithm with up to 1,024 processes using both synthetic and application data. Here, we demonstrate the use of our algorithm in critical point tracking and super-level set extraction with high-speed imaging experiments and fusion plasma simulations, respectively.

97 MATHEMATICS AND COMPUTING↗

hPIC2: A hardware-accelerated, hybrid particle-in-cell code for dynamic plasma-material interactions

The exascale era of high performance computing promises to bring the field of computational plasma physics ever closer to the goal of accurate multiscale modeling. Such computers will rely on hardware acceleration to offload work to dedicated components, notably general-purpose graphics processing units (GPUs). However, devices from different manufacturers require software to be written with different parallel programming models, greatly increasing the code maintenance burden of applications designed to perform on more than one such device. hPIC2 is a hybrid plasma simulation code developed with the Kokkos performance portability framework to target the architectures that will drive exascale computing for the foreseeable future. As a hybrid simulation code, hPIC2 investigates the simultaneous use of various plasma models on the same domain, at the same time. hPIC2 also optionally couples to RustBCA, which accurately models ion-material interactions using the binary collision approximation (BCA) method. In conclusion, hPIC2 therefore achieves scalable performance on a variety of computing architectures when simulating complex and diverse plasmas, particularly near plasma-material interfaces.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

LLNL/thicket

Thicket is a python-based toolkit for Exploratory Data Analysis (EDA) of parallel performance data that enables performance optimization and understanding of applications’ performance on supercomputers. It bridges the performance tool gap between being able to consider only a single instance of a simulation run (e.g., single platform, single measurement tool, or single scale) and finding actionable insights in multi-dimensional, multi-scale, multi-architecture, and multi-tool performance datasets.

Brink, Stephanie Labasan↗

Orchestrating Fault Prediction with Live Migration and Checkpointing

Checkpoint/Restart (C/R) is widely used to provide fault tolerance on High-Performance Computing (HPC) systems. However, Parallel File System (PFS) overhead and failure uncertainty cause significant application overhead. This paper develops an adaptive multi-level C/R model that incorporates a failure prediction and analysis model, which orchestrates failure prediction, checkpointing, checkpoint frequency, and proactive live migration along with the additional benefit of Burst Buffers (BB). It effectively reduces the overheads due to failures, checkpointing, and recovery. Simulation results for the Summit supercomputer yield a reduction of ~20%-86% in application overhead due to BBs, orchestrated failure prediction, and migration. We also observe a ~29% decrease in checkpoint writes to BBs, which can increase the longevity of the BB storage devices.

Behera, Subhendu↗

Non-Orthogonal Configuration Interaction for Fragments

An overview is given of the non-orthogonal configuration interaction (NOCI) implementation in the open-source program GronOR. After pointing out the specifics of NOCI and its extension to ensembles of molecules NOCI-Fragments (NOCI-F), the theoretical foundations of the approaches are described. The introduction of a common molecular orbital basis, the selection of the most contributing determinants pairs, and a massively parallel and GPU-accelerated implementation have made possible the application of NOCI to medium-sized systems as illustrated in a short overview of some of the recent applications.

de Graaf, C.↗

Efficient exascale discretizations: High-order finite element methods

Efficient exploitation of exascale architectures requires rethinking of the numerical algorithms used in many large-scale applications. These architectures favor algorithms that expose ultra fine-grain parallelism and maximize the ratio of floating point operations to energy intensive data movement. One of the few viable approaches to achieve high efficiency in the area of PDE discretizations on unstructured grids is to use matrix-free/partially assembled high-order finite element methods, since these methods can increase the accuracy and/or lower the computational time due to reduced data motion. In this paper we provide an overview of the research and development activities in the Center for Efficient Exascale Discretizations (CEED), a co-design center in the Exascale Computing Project that is focused on the development of next-generation discretization software and algorithms to enable a wide range of finite element applications to run efficiently on future hardware. CEED is a research partnership involving more than 30 computational scientists from two US national labs and five universities, including members of the Nek5000, MFEM, MAGMA and PETSc projects. We discuss the CEED co-design activities based on targeted benchmarks, miniapps and discretization libraries and our work on performance optimizations for large-scale GPU architectures. We also provide a broad overview of research and development activities in areas such as unstructured adaptive mesh refinement algorithms, matrix-free linear solvers, high-order data visualization, and list examples of collaborations with several ECP and external applications.

97 MATHEMATICS AND COMPUTING↗

On-lattice kinetic Monte Carlo approaches for modeling molecular anisotropy in resveratrol crystallization

Stilbenes are a class of organic compounds with broad-ranging pharmaceutical and agricultural applications, which are typically isolated and purified through recrystallization. We are motivated by reducing experimental waste and optimizing yield via developing predictive simulations for processing-dependent crystal morphologies. Using resveratrol as a model stilbene system, we have developed an approach for simulating crystallization with molecular resolution using on-lattice kinetic Monte Carlo. In this work, we highlight modifications to the Stochastic Parallel PARticle Kinetic Simulator (SPPARKS) software package, which were essential to this application. Key enhancements include the incorporation of non-orthogonal cell shapes and monomer anisotropy approximations using bound hard spheres. This new SPPARKS application has been applied to resveratrol with attachment energy libraries obtained from density functional theory, resulting in excellent agreement with experimental morphology prediction.

crystallization↗

Hvac: Removing I/O Bottleneck for Large-Scale Deep Learning Applications

Scientific communities are increasingly adopting deep learning (DL) models in their applications to accelerate scientific discovery processes. However, with rapid growth in the computing capabilities of HPC supercomputers, large-scale DL applications have to spend a significant portion of training time performing I/O to a parallel storage system. Previous research works have investigated optimization techniques such as prefetching and caching. Unfortunately, there exist non-trivial challenges to adopting the existing solutions on HPC supercomputers for large-scale DL training applications, which include non-performance and/or failures at extreme scale, lack of portability and generality in design, complex deployment methodology, and being limited to a specific application or dataset. To address these challenges, we propose High-Velocity AI Cache (HVAC), a distributed read-cache layer that targets and fully exploits the node-local storage or near node-local storage technology. HVAC seamlessly accelerates read I/O by aggregating node-local or near node-local storage, avoiding metadata lookups and file locking while preserving portability in the application code. We deploy and evaluate HVAC on 1,024 nodes (with over 6000 NVIDIA V100 GPUS) of the Summit supercomputer. In particular, we evaluate the scalability, efficiency, accuracy, and load distribution of HVAC compared to GPFS and XFS-on-NVMe. With four different DL applications, we observe an average 25 % performance improvement atop GPFS and 9% drop against XFS-on-NVMe, which scale linearly and are considered the performance upper bound. We envision HVAC as an important caching library for upcoming HPC supercomputers such as Frontier.

Khan, Awais↗