Sphynx: a parallel multi-GPU graph partitioner for distributed-memory systems.
Abstract not provided.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Abstract not provided.
Abstract. Previous phases of the Coupled Model Intercomparison Project (CMIP) have primarily focused on simulations driven by atmospheric concentrations of greenhouse gases (GHGs), for both idealized model experiments and climate projections of different emissions scenarios. We argue that although this approach was practical to allow parallel development of Earth system model simulations and detailed socioeconomic futures, carbon cycle uncertainty as represented by diverse, process-resolving Earth system models (ESMs) is not manifested in the scenario outcomes, thus omitting a dominant source of uncertainty in meeting the Paris Agreement. Mitigation policy is defined in terms of human activity (including emissions), with strategies varying in their timing of net-zero emissions, the balance of mitigation effort between short-lived and long-lived climate forcers, their reliance on land use strategy, and the extent and timing of carbon removals. To explore the response to these drivers, ESMs need to explicitly represent complete cycles of major GHGs, including natural processes and anthropogenic influences. Carbon removal and sequestration strategies, which rely on proposed human management of natural systems, are currently calculated in integrated assessment models (IAMs) during scenario development with only the net carbon emissions passed to the ESM. However, proper accounting of the coupled system impacts of and feedback on such interventions requires explicit process representation in ESMs to build self-consistent physical representations of their potential effectiveness and risks under climate change. We propose that CMIP7 efforts prioritize simulations driven by CO2 emissions from fossil fuel use and projected deployment of carbon dioxide removal technologies, as well as land use and management, using the process resolution allowed by state-of-the-art ESMs to resolve carbon–climate feedbacks. Post-CMIP7 ambitions should aim to incorporate modeling of non-CO2 GHGs (in particular, sources and sinks of methane and nitrous oxide) and process-based representation of carbon removal options. These developments will allow three primary benefits: (1) resources to be allocated to policy-relevant climate projections and better real-time information related to the detectability and verification of emissions reductions and their relationship to expected near-term climate impacts, (2) scenario modeling of the range of possible future climate states including Earth system processes and feedbacks that are increasingly well-represented in ESMs, and (3) optimal utilization of the strengths of ESMs in the wider context of climate modeling infrastructure (which includes simple climate models, machine learning approaches and kilometer-scale climate models).
Presented by LRST at the AIChE 2020 Annual Meeting, 11/16
High-performance computing (HPC) is often perceived as a matter of making large-scale systems (e.g., clusters) run as fast as possible, regardless the required programming effort. However, the idea of "bringing HPC to the masses" has recently emerged. Inspired by this vision, we have designed SHAD, the Scalable High-performance Algorithms and Data-structures library. SHAD is open source software, written in C++, for C++ developers. Unlike other HPC libraries for distributed systems, which rely on SPMD models, SHAD adopts a shared-memory programming abstraction, to make C++ programmers feel at home. Underneath, SHAD manages tasking and data-movements, moving the computation where data resides and taking advantage of asynchrony to tolerate network latency. At the bottom of his stack, SHAD can interface with multiple runtime systems: this not only improves developer’s productivity, by hiding the complexity of such software and of the underlying hardware, but also greatly enhance code portability. Thanks to its abstraction layers, SHAD can indeed target different systems, ranging from laptops to HPC clusters, without any need for modifying the user-level code. We have prototyped and open-sourced the implementation of (a subset of) the C++ standard library (STL) targeting multi-node HPC clusters. Our work allows plain STL-based C++ code to scale on HPC systems, with no need for rewriting the code to exploit the complex hardware. SHAD is available under Apache v2 License at https://github.com/pnnl/SHAD. In this paper we overview the design of the SHAD library, depicting its main components: runtime systems abstractions for tasking; parallel and distributed data-structures; STL-compliant interfaces and algorithms.
At the Savannah River Site (SRS), low-level liquid tank waste is mixed with cementitious materials to produce a solid, monolithic waste form referred to as saltstone. The liquid saltstone mixture is pumped into massive concrete structures referred to as saltstone disposal units (SDUs) for curing and long-term storage. SDUs are vaults made of high performance concrete forming a barrier between saltstone and the surrounding environment. Predicting the long-term performance of SDU vault concrete, as well as the saltstone material itself, is critical to understanding the potential releases of radionuclides to the surrounding environment as described in the Saltstone Disposal Facility (SDF) performance assessment (PA). The hydraulic conductivity profile for SDU vault concrete was conservatively assumed to degrade linearly from the initial saturated hydraulic conductivity of the intact concrete over a long period of time. The initial hydraulic properties of saturated SDU vault concrete have been measured, and the final hydraulic properties of degraded vault concrete are assumed to be the same as the porous backfill material that will surround the vault after closure. The assumed final condition of the vault concrete is also conservative in that it provides a greater estimate of water penetration through the vault than would generally be expected. Therefore, insights that provide more realistic but conservative estimates as to how degradation processes may change the hydraulic properties of the SDU vault concrete, the time dependent relationship of the vault concrete hydraulic properties to degradation (and thus time), and the final hydraulic properties of the vault concrete provides useful additions to the SDF PA. This paper focuses on the assumption that the saltstone vault concrete saturated hydraulic conductivity would decrease linearly as a function of degradation time; this assumption resulted from a worst-case, weighted (by relative layer thickness), arithmetic average of intact and backfill material hydraulic conductivities representing flow parallel to a layered system. This research indicates that the effective saturated hydraulic conductivity of the degraded SDU vault concrete should be based on the relative thickness-weighted geometric mean of the intact and degraded layers hydraulic conductivities. This alternative, which is supported extensively in the relevant literature and evaluated in this paper, leads to a more realistic and defensible estimate of the effective saturated hydraulic conductivity in degraded vault concrete while remaining conservative. The impact of using the more realistic estimate of the effective saturated hydraulic conductivity for the degraded concrete decreases the effective saturated hydraulic conductivity by up to several orders of magnitude for early times (i.e., when the SDU vault concrete is assumed degraded via sulfate attack and carbonation). This revised estimate of the effective saturated hydraulic conductivity has been implemented in the recently revised SDF PA. (authors)
High-fidelity simulations of complex plasma systems allow researchers to gain key insights into and understanding of these systems. To facilitate massively parallel high-fidelity plasma simulations, finite-element-based particle-in-cell capabilities are being developed within the open-source Multiphysics Object-Oriented Simulation Environment (MOOSE) based framework called Software for Advanced Large-scale Analysis of MAgnetic confinement for Numerical Design, Engineering & Research (SALAMANDER). While SALAMANDER’s primary objective is modeling edge plasmas and plasma-facing components in fusion devices, the particle-in-cell capabilities being developed are general and will support modeling low-temperature plasmas as well. Previously, collisionless magnetostatic simulation capabilities have been verified with the two-stream and Dorey-Guest-Harris instabilities, and single particle motion. Collisions were implemented using the direct simulation Monte Carlo method, and verification of this capability will be presented here several verification problems: relaxation of a randomly initialized gas to a Maxwellian distribution, Fourier heat flow, and comparison of reaction rates to both analytic calculations and those calculated using a multi-term Boltzmann solver.
This project investigated the cyber-security impacts of moving from an all analog, point-to-point, instrumentation and control (I&C) system to a digital I&C system based on Modbus and a shared communication medium. A formalism called a hybrid attack graph was expanded to support the nuclear research reactor system. The hybrid attack graph allows one to check a system for vulnerabilities, in this case cyber-security vulnerabilities, and to document the attack vectors (scenarios) causing those vulnerabilities. In parallel, a simulation of the system was developed to model both the physical reactor parameters and operations, as well as the network interconnects and communications. This simulation platform was modeled on the nuclear research reactor located at Washington State University. The simulation platform provided a sandbox to evaluate and quantify the impact of identified and proposed vulnerabilities in the system and to determine the effectiveness of countermeasures at stopping these attacks. The simulation and hybrid attack graph tools were integrated to provide a streamlined process of generating attack scenarios, playing those scenarios out in the simulation, and then analyzing the results to correlate system state to states in the hybrid attack graph. This process was used to (1) quantify the impact of attack scenarios and (2) to determine if the system moved through the hybrid attack graph as anticipated. The hybrid attack graph tool was extended and customized to produce a tool to automatically identify critical assets (CAs) and critical digital assets (CDAs) as defined by NRC Regulatory Guide 5.71. This tool was verified using the nuclear research reactor at Washington State University. Finally, a series of educational modules covering the findings of the different aspects of this research have been created.
A pair of flat parallel surfaces, each freely diffusing along the direction of their separation, will eventually come into contact. If the shapes of these surfaces also fluctuate, then contact will occur when their centers-of-mass remain separated by a nonzero distance ℓ. An example of such a situation is the motion of interfaces between two phases at conditions of thermodynamic coexistence, and in particular the annihilation of domain wall pairs under periodic boundary conditions. Here we present a general approach to calculate the probability distribution of the contact distance ℓ and determine how its most likely value ℓ* depends on the surfaces' lateral size L. Using the Edward-Wilkinson equation as a model for interfaces, we demonstrate that ℓ* scales weakly with system size, i.e., the dependence of ℓ* on L for both (1+1)- and (2+1)-dimensional interfaces is such that lim L→∞ (ℓ*/L) = 0. In particular, for (2+1)-dimensional interfaces ℓ* is an algebraic function of logL, a result that is confirmed by computer simulations of slab-shaped domains formed under periodic boundary conditions. Overall, this weak scaling implies that such domains remain topologically intact until ℓ becomes very small compared to the lateral size of the interface, contradicting expectations from equilibrium thermodynamics.
Today’s power grid is becoming more diverse and integrated with high-level distributed energy resources and smart control technologies that is creating a new set of grid management challenges in terms of large-scale, nonlinear, and non-convex problem modeling, complex and time-consuming computation, as well as difficult uncertainty handling. This project focused on solving a challenging multi-period security-constrained generation scheduling problem, which is of great importance for maximizing the social welfare of real-time dispatch, day-ahead market, as well as weekly planning of power systems. Our developed software explored parallel optimization algorithms for complex and realistic power system models, and develop fast, efficient, and robust grid optimization solutions on the high-performance computing platform that will enable increased grid economics, flexibility, resilience, as well as energy security in the United States.
Within a machine, mechanisms and motion are organized in what is known as a “kinematic arrangement,” which helps classify machines based on how they move. The most common kinematic arrangements for additive manufacturing systems are Cartesian, followed by delta, and then six-degrees-of-freedom robotic arms. However, there are a multitude of less common systems, such as the SCARA, polar robots, cable driven parallel robots, mobile platforms, and multi-agent systems. This chapter surveys these various kinematic arrangements to give a broad understanding of the mechanisms underlying motion within additive manufacturing systems. Understanding these mechanisms and their resulting motion provides a framework for discussing path planning for all scales and families of additive manufacturing.
The rapid introduction of mobile navigation aides that use real-time road network information to suggest alternate routes to drivers is making it more difficult for researchers and government transportation agencies to understand and predict the dynamics of congested transportation systems. Computer simulation is a key capability for these organizations to analyze hypothetical scenarios; however, the complexity of transportation systems makes it challenging for them to simulate very large geographical regions, such as multi-city metropolitan areas. In this article, we describe enhancements to the Mobiliti parallel traffic simulator to model dynamic rerouting behavior with the addition of vehicle controller actors and vehicle-to-controller reroute requests. The simulator is designed to support distributed-memory parallel execution using discrete event simulation and be scalable on high-performance computing platforms. We demonstrate the potential of the simulator by analyzing the impact of varying the population penetration rate of dynamic rerouting on the San Francisco Bay Area road network. Using high-performance parallel computing, we can simulate a day in the San Francisco Bay Area with 19 million vehicle trips with 50 percent dynamic rerouting penetration over a road network with 0.5 million nodes and 1 million links in less than three minutes. We present a sensitivity study on the dynamic rerouting parameters, discuss the simulator’s parallel scalability, and analyze system-level impacts of changing the dynamic rerouting penetration. Furthermore, we examine the varying effects on different functional classes and geographical regions and present a validation of the simulation results compared to real-world data.
All Cryomodule (CM) and Cryogenic Distribution System (CDS) relieving into Helium Low Pressure (LP) return header, which is connected to compressor suction so, helium can be preserved during small flow relieving event and recirculated to system. However, during worst case scenario, Helium LP header requires a parallel plate relief device to relieve excess pressure from header. To complete the CDS Warm piping header, a new design for a 10 parallel plate relief device is necessary to relieve outside of the tunnel into atmosphere.
Machine-learning potentials (MLPs) trained on data from quantum-mechanics based first-principles methods can approach the accuracy of the reference method at a fraction of the computational cost. To facilitate efficient MLP-based molecular dynamics and Monte Carlo simulations, an integration of the MLPs with sampling software is needed. Here, we develop two interfaces that link the atomic energy network (ænet) MLP package with the popular sampling packages TINKER and LAMMPS. The three packages, ænet, TINKER, and LAMMPS, are free and open-source software that enable, in combination, accurate simulations of large and complex systems with low computational cost that scales linearly with the number of atoms. Scaling tests show that the parallel efficiency of the ænet–TINKER interface is nearly optimal but is limited to shared-memory systems. The ænet–LAMMPS interface achieves excellent parallel efficiency on highly parallel distributed memory systems and benefits from the highly optimized neighbor list implemented in LAMMPS. We demonstrate the utility of the two MLP interfaces for two relevant example applications: the investigation of diffusion phenomena in liquid water and the equilibration of nanostructured amorphous battery materials.
It has long been assumed that all matter will adopt simple close-packed lattices and become metallic under pressure, in accordance with the Thomas–Fermi–Dirac (TFD) model. However, this model struggles to explain pressure-driven complex structural transitions that have been observed in elements, including sodium, challenging our conventional understanding of compressed matter. Moreover, in stark contrast to the TFD model, first-principles calculations suggest that various elements and compounds become electrides under pressure. Electrides, characterized by concentrations of charge density at interstitial regions, can be thought of as ionic compounds where electrons behave as the anions. Though ambient-pressure molecular electrides have been extensively studied via experiments and computations, high-pressure electrides (HPEs) are not well-understood. The identification and characterization of HPEs have been, to date, based purely on theory, including topological analysis of the electron density and the electron localization function. Here, we review these theoretical analysis tools and suggest guidelines that can be used to classify systems as electrides. Moreover, we describe models used to rationalize the electronic structure of HPEs, drawing parallels with ambient-pressure molecular systems, and encourage the development of experimental techniques that provide evidence for the theoretically calculated charge localization.
The quantum Hall effect is the birthplace of topological states of matter, a major theme at the forefront of condensed matter physics in the past two decades. The fractional quantum Hall (FQH) effect revolutionized our understanding of phases of electronic matter. FQH states support exotic fractionally charged excitations that obey Abelian or non-Abelian fractional statistics, which are topological excitations that result from the underlying topological order. During this project, our group discovered a previously unrecognized geometric degree of freedom of incompressible FQH states and studied that for a variety of gapped FQH states. We brought this new concept into direct contact with experiments for the first time by generalizing it to Fermi-liquid states of composite fermions. Using the newly formulated powerful infinite Density Matrix Renormalization Group method, our numerical calculations yielded a parameter free prediction that was found to be in excellent agreement with experimental findings on electron systems in semiconductor heterostructures. In parallel, we performed extensive numerical studies on different, competing phases at various Landau level filling factors, and quantum phase transitions that result from such a competition, e.g. Abelian-non-Abelian phase transitions in bilayer systems. We studied geometrical excitations dubbed “gravitons” (because of their analogy with excitations in the theory of gravitation) and ways to excite and detect them, and explored how they couple with topological excitations. In graphene-based chiral materials, we realized the ability to tune through different incompressible and compressible states in a single Landau level, and found appropriate experimental parameters for the exploration of universal Luttinger liquid behavior not obtained in semiconductor-based electron systems. We showed that topological systems had a very different response from nontopological systems to strong disorder (many-body localization) as well as periodic drives.
The next generation HPC and data centers are likely to be reconfigurable and data-centric due to the trend of hardware specialization and the emergence of data-driven applications. In this work, we propose ARENA – an asynchronous reconfigurable accelerator ring architecture as a potential scenario on how the future HPC and data centers will be like. Despite using the coarse-grained reconfigurable arrays (CGRAs) as the substrate platform, our key contribution is not only the CGRA-cluster design itself, but also the ensemble of a new architecture and programming model that enables asynchronous tasking across a cluster of reconfigurable nodes, so as to bring specialized computation to the data rather than the reverse. We presume distributed data storage without asserting any prior knowledge on the data distribution. Hardware specialization occurs at runtime when a task finds the majority of data it requires are available at the present node. In other words, we dynamically generate specialized CGRA accelerators where the data reside. The asynchronous tasking for bringing computation to data is achieved by circulating the task token, which describes the dataflow graphs to be executed for a task, among the CGRA cluster connected by a fast ring network. Evaluations on a set of HPC and data-driven applications across different domains show that ARENA can provide better parallel scalability with reduced data movement (53.9 percent). Compared with contemporary compute-centric parallel models, ARENA can bring on average 4.37× speedup. The synthesized CGRAs and their task-dispatchers only occupy 2.93mm 2 chip area under 45nm process technology and can run at 800MHz with on average 759.8mW power consumption. ARENA also supports the concurrent execution of multi-applications, offering ideal architectural support for future high-performance parallel computing and data analytics systems.
Sparse symmetric positive definite systems of equations are ubiquitous in scientific workloads and applications. Parallel sparse Cholesky factorization is the method of choice for solving such linear systems. Therefore, the development of parallel sparse Cholesky codes that can efficiently run on today’s large-scale heterogeneous distributed-memory platforms is of vital importance. Modern supercomputers offer nodes that contain a mix of CPUs and GPUs. To fully utilize the computing power of these nodes, scientific codes must be adapted to offload expensive computations to GPUs. We present symPACK, a GPU-capable parallel sparse Cholesky solver that uses one-sided communication primitives and remote procedure calls provided by the UPC++ library. We also utilize the UPC++ "memory kinds" feature to enable efficient communication of GPU-resident data. We show that on a number of large problems, symPACK outperforms comparable state-of-the-art GPU-capable Cholesky factorization codes by up to 14x on the NERSC Perlmutter supercomputer.
Summary Parallel I/O is a critical technique for moving data between compute and storage subsystems of supercomputers. With massive amounts of data produced or consumed by compute nodes, high‐performant parallel I/O is essential. I/O benchmarks play an important role in this process; however, there is a scarcity of I/O benchmarks representative of current workloads on HPC systems. Toward creating representative I/O kernels from real‐world applications, we have created h5bench , a set of I/O kernels that exercise hierarchical data format version 5 (HDF5) I/O on parallel file systems in numerous dimensions. Our focus on HDF5 is due to the parallel I/O library's heavy usage in various scientific applications running on supercomputing systems. The various tests benchmarked in the h5bench suite include I/O operations (read and write), data locality (arrays of basic data types and arrays of structures), array dimensionality (one‐dimensional arrays, two‐dimensional meshes, three‐dimensional cubes), I/O modes (synchronous and asynchronous). In this paper, we present the observed performance of h5bench executed along several of these dimensions on existing supercomputers (Cori and Summit) and pre‐exascale platforms (Perlmutter, Theta, and Polaris). h5bench measurements can be used to identify performance bottlenecks and their root causes and evaluate I/O optimizations. As the I/O patterns of h5bench are diverse and capture the I/O behaviors of various HPC applications, this study will be helpful to the broader supercomputing and I/O community.