Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Exascale applications”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

GPU Direct I/O with HDF5

Exascale HPC systems are being designed with accelerators, such as GPUs, to accelerate parts of applications. In machine learning workloads as well as large-scale simulations that use GPUs as accelerators, the CPU (or host) memory is currently used as a buffer for data transfers between GPU (or device) memory and the file system. If the CPU does not need to operate on the data, then this is sub-optimal because it wastes host memory by reserving space for duplicated data. Furthermore, this “bounce buffer” approach wastes CPU cycles spent on transferring data. A new technique, NVIDIA GPUDirect Storage (GDS), can eliminate the need to use the host memory as a bounce buffer. Thereby, it becomes possible to transfer data directly between the device memory and the file system. This direct data path shortens latency by omitting the extra copy and enables higher-bandwidth. To take full advantage of GDS in existing applications, it is necessary to provide support with existing I/O libraries, such as HDF5 and MPI-IO, which are heavily used in applications. In this paper, we describe our effort of integrating GDS with HDF5, the top I/O library at NERSC and at DOE leadership computing facilities. We design and implement this integration using a HDF5 Virtual File Driver (VFD). The GDS VFD provides a file system abstraction to the application that allows HDF5 applications to perform I/O without the need to move data between CPUs and GPUs explicitly. We compare performance of the HDF5 GDS VFD with explicit data movement approaches and demonstrate superior performance with the GDS method.

Ravi, J↗

Project DarkStar: Vision for LLNL in 2030

DarkStar was a Strategic Initiative (FY2021-FY2024) to investigate applications of Artificial Intelligence (AI) and Machine Learning (ML) to scientific problems of complex hydrodynamics, shockwave physics and energetic materials. The research focused on physics and engineering design as a process that can be tremendously accelerated through merging AI with advanced physics simulation on exascale-class platforms, and to experimentally validate this revolutionary new approach through dynamic materials campaigns. A central thread of scientific inquiry was in the application of AI to enable human understanding of how to control hydrodynamic instability (which has impacts to areas such as inertial confinement fusion) via engineering features and time-dependent sources. Motivated by an unfinished line of research started by Dr. Johnny von Neumann, AI-enabled simulation approaches were developed that allowed DarkStar researchers to uncover several ground-breaking discoveries regarding hydrodynamic instability, including how to completely suppress Richtmyer-Meshkov instability (RMI). These S&T discoveries, along with other advances, have shown the way for an entirely new approach to time-dependent problems known as inverse design – the idea that complex systems can be developed directly from a final state that is to be achieved and resolve the initial design via satisfying several constraints simultaneously via AI/ML. Through experimental campaigns conducted across a wide range of facilities in the NNSA complex (the High Explosive Application Facility at LLNL, the Dynamic Compression Sector/Advanced Photon Source at Argonne National Lab, and Special Technologies Laboratory at MSTS) the radical new AI/ML approach to engineering complex material dynamics was verified, establishing a new field of study within the realm of shock physics. As advanced manufacturing capabilities continue to develop, the great importance of inverse design as a means to apply that technology effectively for NNSA missions will feature prominently over this decade. DarkStar has positioned NNSA as a world-leader in this newly emerging cross-disciplinary area of AI methods for advanced physics simulation and pioneered multiple novel approaches that have enabled the broader scientific community. By allowing us to see past the horizon, to 2030 and beyond, DarkStar has illuminated the vast potential of AI/ML to impact a wide range of new national security missions and, consequently, multiple areas of further research have already emerged across the NNSA and DOD complex.

42 ENGINEERING↗

Exploring the Frontiers of Energy Efficiency using Power Management at System Scale

In the face of surging power demands for exascale HPC systems, this work tackles the critical challenge of understanding the impact of software-driven power management techniques like Dynamic Voltage and Frequency Scaling (DVFS) and Power Capping. These techniques have been actively developed over the past few decades. By combining insights from GPU benchmarking to understand application power profiles, we present a telemetry data-driven approach for deriving energy savings projections. This approach has been demonstrably applied to the Frontier supercomputer at scale. Our findings based on three months of telemetry data indicate that, for certain resource-constrained jobs, significant energy savings (up to 8.5%) can be achieved without compromising performance. This translates to a substantial cost reduction, equivalent to 1438 MWh of energy saved. The key contribution of this work lies in the methodology for establishing an upper limit for these best-case scenarios and its successful application. This work enables HPC professionals to optimize the power-performance trade-off within constrained power budgets, not only for the exascale era but also beyond.

Karimi, Ahmad Maroof↗

Holistic Measurement Driven Resilience: Combining Operational Fault and Failure Measurements and Fault Injection for Quantifying Fault Detection, Propagation and Impact. Final report

For HPC systems to date, application resilience to faults and failures has been accomplished by the brute- force method of checkpoint/restart, which allows an application to make forward progress in the face of system and application faults, errors, and failures independent of root cause or end result. It has remained the primary resilience mechanism because we lack a way to identify faults and anticipate consequences early enough to take meaningful mitigating action. However, checkpoint/restart implementations put a tremendous burden on system resources and on the applications themselves and is becoming less feasible at scale. Because we have not yet operated at scales at which checkpoint/restart fails to provide forward progress, despite increasing costs, vendors have had little motivation to provide the instrumentation necessary for early identification of faults and failures. However, as we move from petascale to exascale, component mean time to failure (MTTF) will render the existing techniques ineffectual and/or too expensive. Furthermore, fault recovery mechanisms such as failover and/or error correction introduce performance inconsistency. Instrumentation allowing early indication of problems and tools to enable use of such information by systems, operating systems, and applications offer an alternative, more scalable and less costly solution. In the HMDR project, we built on our experience and expertise developed and accumulated over years of research on design, monitoring, measurement, and assessment of resilient computing systems. Analysis of field data on the current and past generations of extreme-scale systems revealed several challenges that, if not addressed in increasingly larger and more complex systems, may hinder the effectiveness of future exascale computing systems. Specifically, i) file systems and interconnects in current-generation large-scale systems already operate at the margins of resiliency, including consistent performance, and may not scale to larger deployments; ii) automated, software-based failover mechanisms are frequently inadequate and can introduce wider failures, such that failures during recovery may lead to system/application failures, including system-wide outages; and iii) silent data corruption represents a critical fault mode and will require efficient detection mechanisms if next-generation applications are to take full advantage of exascale hardware. To address the above challenges, we assembled a team of world-renowned experts in resilient extreme- scale computing from the University of Illinois (Electrical and Computer Engineering, Computer Science, and NCSA), SNL, LANL, NERSC, and Cray. Our team includes representatives from centers that house many of the largest HPC resources in the world, both today and over the coming years. The team has a unique track record of research in i) system and application failure characterization based on the analysis of field data, ii) data-driven design of fault/error detection mechanisms, and iii) experimental characterization of system/application resiliency. The team includes system owners/operators who provide continuous data collection and access and ensure installation of appropriate analysis tools.

97 MATHEMATICS AND COMPUTING↗

Characterization and identification of HPC applications at leadership computing facility

High Performance Computing (HPC) is an important method for scientific discovery via large-scale simulation, data analysis, or artificial intelligence. Leadership-class supercomputers are expensive, but essential to run large HPC applications. The Petascale era of supercomputers began in 2008, with the first machines achieving performance in excess of one petaflops, and with the advent of new supercomputers in 2021 (e.g., Aurora, Frontier), the Exascale era will soon begin. However, the high theoretical computing capability (i.e., peak FLOPS) of a machine is not the only meaningful target when designing a supercomputer, as the resources demand of applications varies. A deep understanding of the characterization of applications that run on a leadership supercomputer is one of the most important ways for planning its design, development and operation. In order to improve our understanding of HPC applications, user demands and resource usage characteristics, we perform correlative analysis of various logs for different subsystems of a leadership supercomputer. This analysis reveals surprising, sometimes counter-intuitive patterns, which, in some cases, conflicts with existing assumptions, and have important implications for future system designs as well as supercomputer operations. For example, our analysis shows that while the applications spend significant time on MPI, most applications spend very little time on file I/O. Combined analysis of hardware event logs and task failure logs show that the probability of a hardware FATAL event causing task failure is low. Combined analysis of control system logs and file I/O logs reveals that pure POSIX I/O is used more widely than higher level parallel I/O. Based on holistic insights of the application gained through combined and co-analysis of multiple logs from different perspectives and general intuition, we engineer features to "fingerprint" HPC applications. We use t-SNE (a machine learning technique for dimensionality reduction) to validate the explainability of our features and finally train machine learning models to identify HPC applications or group those with similar characteristic. To the best of our knowledge, this is the first work that combines logs on file I/O, computing, and inter-node communication for insightful analysis of HPC applications in production.

Liu, Zhengchun↗

Resiliency in numerical algorithm design for extreme scale simulations

Here this work is based on the seminar titled ‘Resiliency in Numerical Algorithm Design for Extreme Scale Simulations’ held March 1–6, 2020, at Schloss Dagstuhl, that was attended by all the authors. Advanced supercomputing is characterized by very high computation speeds at the cost of involving an enormous amount of resources and costs. A typical large-scale computation running for 48 h on a system consuming 20 MW, as predicted for exascale systems, would consume a million kWh, corresponding to about 100k Euro in energy cost for executing 10 23 floating-point operations. It is clearly unacceptable to lose the whole computation if any of the several million parallel processes fails during the execution. Moreover, if a single operation suffers from a bit-flip error, should the whole computation be declared invalid? What about the notion of reproducibility itself: should this core paradigm of science be revised and refined for results that are obtained by large-scale simulation? Naive versions of conventional resilience techniques will not scale to the exascale regime: with a main memory footprint of tens of Petabytes, synchronously writing checkpoint data all the way to background storage at frequent intervals will create intolerable overheads in runtime and energy consumption. Forecasts show that the mean time between failures could be lower than the time to recover from such a checkpoint, so that large calculations at scale might not make any progress if robust alternatives are not investigated. More advanced resilience techniques must be devised. The key may lie in exploiting both advanced system features as well as specific application knowledge. Research will face two essential questions: (1) what are the reliability requirements for a particular computation and (2) how do we best design the algorithms and software to meet these requirements? While the analysis of use cases can help understand the particular reliability requirements, the construction of remedies is currently wide open. One avenue would be to refine and improve on system- or application-level checkpointing and rollback strategies in the case an error is detected. Developers might use fault notification interfaces and flexible runtime systems to respond to node failures in an application-dependent fashion. Novel numerical algorithms or more stochastic computational approaches may be required to meet accuracy requirements in the face of undetectable soft errors. These ideas constituted an essential topic of the seminar. The goal of this Dagstuhl Seminar was to bring together a diverse group of scientists with expertise in exascale computing to discuss novel ways to make applications resilient against detected and undetected faults. In particular, participants explored the role that algorithms and applications play in the holistic approach needed to tackle this challenge. This article gathers a broad range of perspectives on the role of algorithms, applications and systems in achieving resilience for extreme scale simulations. The ultimate goal is to spark novel ideas and encourage the development of concrete solutions for achieving such resilience holistically.

79 ASTRONOMY AND ASTROPHYSICS↗

Nuclear Physics Exascale Requirements Review: An Office of Science Review sponsored jointly by Advanced Scientific Computing Research and Nuclear Physics, June 15 - 17, 2016, Gaithersburg, Maryland

Imagine being able to predict — with unprecedented accuracy and precision — the structure of the proton and neutron, and the forces between them, directly from the dynamics of quarks and gluons, and then using this information in calculations of the structure and reactions of atomic nuclei and of the properties of dense neutron stars (NSs). Also imagine discovering new and exotic states of matter, and new laws of nature, by being able to collect more experimental data than we dream possible today, analyzing it in real time to feed back into an experiment, and curating the data with full tracking capabilities and with fully distributed data mining capabilities. Making this vision a reality would improve basic scientific understanding, enabling us to precisely calculate, for example, the spectrum of gravity waves emitted during NS coalescence, and would have important societal applications in nuclear energy research, stockpile stewardship, and other areas. This review presents the components and characteristics of the exascale computing ecosystems necessary to realize this vision.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Parthenon—a performance portable block-structured adaptive mesh refinement framework

On the path to exascale the landscape of computer device architectures and corresponding programming models has become much more diverse. While various low-level performance portable programming models are available, support at the application level lacks behind. To address this issue, we present the performance portable block-structured adaptive mesh refinement (AMR) framework Parthenon, derived from the well-tested and widely used Athena++ astrophysical magnetohydrodynamics code, but generalized to serve as the foundation for a variety of downstream multi-physics codes. Parthenon adopts the Kokkos programming model, and provides various levels of abstractions from multidimensional variables, to packages defining and separating components, to launching of parallel compute kernels. Parthenon allocates all data in device memory to reduce data movement, supports the logical packing of variables and mesh blocks to reduce kernel launch overhead, and employs one-sided, asynchronous MPI calls to reduce communication overhead in multi-node simulations. Using a hydrodynamics miniapp, we demonstrate weak and strong scaling on various architectures including AMD and NVIDIA GPUs, Intel and AMD x86 CPUs, IBM Power9 CPUs, as well as Fujitsu A64FX CPUs. At the largest scale on Frontier (the first TOP500 exascale machine), the miniapp reaches a total of 1.7 × 10 13 zone-cycles/s on 9216 nodes (73,728 logical GPUs) at [Formula: see text] weak scaling parallel efficiency (starting from a single node). In combination with being an open, collaborative project, this makes Parthenon an ideal framework to target exascale simulations in which the downstream developers can focus on their specific application rather than on the complexity of handling massively-parallel, device-accelerated AMR.

97 MATHEMATICS AND COMPUTING↗

Evaluating Mesoscale Convective Systems Over the US in Conventional and Multiscale Modeling Framework Configurations of E3SMv1

Organized mesoscale convective systems (MCSs) contribute a significant amount of precipitation in the Central and Eastern US during spring and summer, which impacts the availability of freshwater and flooding events. However, current global Earth system models cannot capture MCSs well and misrepresent the statistics of precipitation in the region. In this study, we investigate the representation of MCSs in three configurations of the Energy Exascale Earth System Model (E3SMv1) by tracking individual storms based on outgoing longwave radiation using a new application of TempestExtremes. Our results indicate that conventional parameterizations of convection, implemented in both low (LR; ~150 km) and high (HR; ~25 km) resolution configurations, fail to capture almost all MCS-like events, in-part because they underestimate high-level cloud ice associated with deep convection. On the other hand, the multiscale modeling framework (MMF; cloud-resolving models embedded in each grid-column of ~150 km resolution E3SMv1) configuration represents MCSs and their annual cycle better. Nevertheless, relative to observations, the E3SMv1-MMF spatial distribution of MCSs and associated precipitation is shifted eastward, and the diurnal timing is lagged. A comparison between the large-scale environment in E3SMv1-MMF and ERA5 reanalysis suggests that the biases during the summer in E3SMv1-MMF are associated with biases in low-level humidity and meridional moisture transport within the low-level jet. The fact that conventional parameterizations of convection, even with high-resolution, cannot capture MCSs over the US suggests that methods with explicit representation of kilometer-scale convective organization, such as the MMF, may be necessary for improving the simulation of these convective systems.

54 ENVIRONMENTAL SCIENCES↗

Evaluation of Flow Routing on the Unstructured Voronoi Meshes in Earth System Modeling

Flow routing is a fundamental process of Earth System Models' (ESMs) river component. Traditional flow routing models rely on Cartesian rectangular meshes, which exhibit limitations, particularly when coupled with unstructured mesh-based ocean components. They also lack the support for regionally refined models. While previous studies have highlighted the potential benefits of unstructured meshes for flow routing, their widespread application and comprehensive evaluation within ESMs remain limited. This study extends the river component of the Energy Exascale Earth System Model to unstructured Voronoi meshes. We evaluated the model's performance in simulating river discharge and water depth across three watersheds spanning the Arctic, temperate, and tropical regions. The results show that while providing several benefits, unstructured mesh-based flow routing can achieve comparable performance to structured mesh-based routing, and their difference is often less than 10%. Although the unstructured mesh-based method could address several existing limitations, this research also shows that additional improvements in the numerical method are needed to fully exploit the advantages of unstructured mesh for hydrologic and ESMs.

54 ENVIRONMENTAL SCIENCES↗

S4PST: Sustainability for Programming Systems and Tools: May Workshop Report

The US Department of Energy (DOE) Exascale Computing Project (ECP) has fostered and strengthened the use of modern software engineering practices for developing applications and libraries, and this effort has resulted in the coordinated and interoperable E4S1 and xSDK2 ecosystems. Although this approach is cost-effective, it relies on robust programming systems and tools (PST) as the underlying foundation for our HPC software. At present, our primary PST stack consists of traditional high-performance computing (HPC) languages, namely Fortran, C, C++, and the popular Python language for data analysis and AI workflows. These languages support various programming frameworks and run-time abstractions that enable parallelism and concurrency across multiple node architectures and thousands of nodes through a variety of interconnect systems. However, to accommodate users’ diverse needs, certain aspects of the HPC ecosystem are delegated to vendor-specific or third-party implementations that extend beyond a particular scientific domain. This broader scope results in a multitude of specifications and variations, which leads to a complex orchestration of many-ecosystems. Unfortunately, this complexity in the ecosystem imposes additional overhead costs on consumers during the latter stages of the development cycle. In addition to the software ecosystem challenge, the upcoming conclusion of the ECP by December 2023 has raised significant concerns within the HPC programming systems community, from both the economic and social perspectives. The ECP has implemented a management structure for software development and funding decisions across all ECP participants by following a conventional hierarchical and centralized approach. However, this structure has prompted certain considerations within the community, particularly in anticipation of the Software Sustainability initiative by the DOE’s Advanced Scientific Computing Research Program (ASCR). For the success of this new initiative, it is of utmost importance to secure consistent funding and foster close engagement with researchers and core developers of existing programming-system products. This collaboration is vital to maintaining the critical capabilities of the current software during the transition phase while proactively adapting to future technology and workforce trends. The community recognizes the significance of adapting to emerging trends and is aware of the inherent fragility of the HPC software ecosystem, particularly in relation to programming systems that cater to all users. The ability to adapt and evolve is essential to staying relevant and effectively addressing these technical, economic, and social challenges. The S4PST team, which represents one of the six ASCR Software Sustainability seedling projects, is dedicated to tackling these challenges through community-based approaches that go beyond the scope of the DOE. This involves collaboration between national laboratories with academia, non-DOE institutions, hardware and system vendors, and international partners. By fostering these partnerships, we aim to create a robust and sustainable HPC software ecosystem that can effectively meet the needs of the community. This new community effort, driven by the eight DOE labs, will take on the responsibility of guiding funding decisions for programming-systems development and maintenance with transparency and consistency across all decisions. Additionally, the team will offer common technical services to the programming systems community, irrespective of their funding situations, and facilitate community-wide incubation to proactively nurture the software ecosystem. By actively engaging with stakeholders and employing a collaborative approach, we can collectively shape the future of programming systems and ensure a robust and thriving HPC software landscape. On May 11–12, 2023, the S4PST team conducted its inaugural kick-off workshop at the Innovative Computing Laboratory (ICL) in the University of Tennessee, Knoxville, hosted by Hartwig Anzt. The workshop encompassed various sessions dedicated to presentations and discussions, with the aim of comprehending the team members’ perspectives on the vision of software sustainability. Additionally, the workshop aimed to identify the technical, economic, and social requirements for sustaining the programming-systems community in the field of HPC. This report provides a summary of the S4PST effort by highlighting five major thrust areas discussed during the workshop: (i) community, (ii) technical support, (iii) training and diversity, (iv) verification, validation and correctness, and (v) emerging technologies. It also encompasses an overview of the presentations and discussions held throughout the event, our views and potential synergies with other seedling efforts, along with the outcomes and key takeaways from our initial discussions.

97 MATHEMATICS AND COMPUTING↗

Understanding processes that control dust spatial distributions with global climate models and satellite observations

Dust aerosol is important in modulating the climate system at local and global scales, yet its spatiotemporal distributions simulated by global climate models (GCMs) are highly uncertain. In this study, we evaluate the spatiotemporal variations of dust extinction profiles and dust optical depth (DOD) simulated by the Community Earth System Model version 1 (CESM1) and version 2 (CESM2), the Energy Exascale Earth System Model version 1 (E3SMv1), and the Modern-Era Retrospective analysis for Research and Applications version 2 (MERRA-2) against satellite retrievals from Cloud-Aerosol Lidar with Orthogonal Polarization (CALIOP), Moderate Resolution Imaging Spectroradiometer (MODIS), and Multi-angle Imaging SpectroRadiometer (MISR). We find that CESM1, CESM2, and E3SMv1 underestimate dust transport to remote regions. E3SMv1 performs better than CESM1 and CESM2 in simulating dust transport and the northern hemispheric DOD due to its higher mass fraction of fine dust. CESM2 performs the worst in the Northern Hemisphere due to its lower dust emission than in the other two models but has a better dust simulation over the Southern Ocean due to the overestimation of dust emission in the Southern Hemisphere. DOD from MERRA-2 agrees well with CALIOP DOD in remote regions due to its higher mass fraction of fine dust and the assimilation of aerosol optical depth. The large disagreements in the dust extinction profiles and DOD among CALIOP, MODIS, and MISR retrievals make the model evaluation of dust spatial distributions challenging. Our study indicates the importance of representing dust emission, dry/wet deposition, and size distribution in GCMs in correctly simulating dust spatiotemporal distributions.

54 ENVIRONMENTAL SCIENCES↗

Lifting and Dropping VMs to Dynamically Transition Between Time- and Space-sharing for Large-Scale HPC Systems

As HPC environments increasingly integrate with edge based systems, system architectures will need to handle a broader class of workloads and scheduling requirements. One result of this shift will be the need to simultaneously support bulk-synchronous parallel (BSP) and on-demand service based applications on the same infrastructure. This in turn will require that future resource management approaches utilize both space-shared as well as time-shared resource scheduling strategies. In this work we introduce the concept of "VM-lifting'' (and its inverse "VM-Dropping'') which allows dynamically switching an HPC workload between space-shared and time-shared scheduling regimes. Our work targets co-kernel based HPC system software environments, in which multiple specialized OS kernels execute natively on dedicated physical resource partitions inside a single compute node. With VM-lifting, a native co-kernel can be migrated at runtime to and from locally hosted Virtual Machine Environments due to changing scheduling requirements of the node. This allows an HPC node to be dynamically (re-)configured as either a time-shared Infrastructure-as-a-Service (IaaS) resource or a dedicated space shared resource based on the current workload demands. We have implemented this approach in the context of the Hobbes Exascale System Software stack and have demonstrated that a node can be reconfigured with minimal impact on the running applications.

Gordon, Nick↗

Preparing MPICH for exascale

The advent of exascale supercomputers heralds a new era of scientific discovery, yet it introduces significant architectural challenges that must be overcome for MPI applications to fully exploit its potential. Among these challenges is the adoption of heterogeneous architectures, particularly the integration of GPUs to accelerate computation. Additionally, the complexity of multithreaded programming models has also become a critical factor in achieving performance at scale. The efficient utilization of hardware acceleration for communication, provided by modern NICs, is also essential for achieving low latency and high throughput communication in such complex systems. In response to these challenges, the MPICH library, a high-performance and widely used Message Passing Interface (MPI) implementation, has undergone significant enhancements. Here, this paper presents four major contributions that prepare MPICH for the exascale transition. First, we describe a lightweight communication stack that leverages the advanced features of modern NICs to maximize hardware acceleration. Second, our work showcases a highly scalable multithreaded communication model that addresses the complexities of concurrent environments. Third, we introduce GPU-aware communication capabilities that optimize data movement in GPU-integrated systems. Finally, we present a new datatype engine aimed at accelerating the use of MPI derived datatypes on GPUs. These improvements in the MPICH library not only address the immediate needs of exascale computing architectures but also set a foundation for exploiting future innovations in high-performance computing. By embracing these new designs and approaches, MPICH-derived libraries from HPE Cray and Intel were able to achieve real exascale performance on OLCF Frontier and ALCF Aurora respectively.

Guo, Yanfei [Argonne National Laboratory (ANL), Ar↗

Singleton Sieving: Overcoming the Memory/Speed Trade-Off in Exascale k-mer Analysis

Traditional filter data structures, such as Bloom filters, do not offer necessary features that modern high-performance data analytics applications need in order to efficiently perform complex data analysis tasks. For example, MetaHipMer, a de novo metagenome assembler, can use filters to weed out singleton k-mers and reduce memory usage by 30%-70%. However, the filter needs the ability to associate values with k-mers in order to perform the analysis in a single communication pass. Bloom filters do not support value associations and cause the application to perform an extra communication pass, thereby increasing the run time. Therefore, MetaHipMer faces a trade off between memory and speed due to the limited capabilities of traditional filters. In this paper, we overcome the memory and speed trade off in MetaHipMer by integrating a GPU-based feature-rich filter, the Two-Choice filter (TCF), in the MetaHipMer pipeline. The TCF uses key-value association to approximately store k-mers with extensions. This allows MetaHipMer to perform k-mer analysis on the GPUs in a single communication pass. Our empirical analysis shows a 50% reduction in memory usage in k-mer analysis on each node in MetaHipMer without any effect on the overall run time or assembly quality. The memory reduction in turn results in a 43% reduction in the number of nodes required to assemble datasets and enables MetaHipMer to scale to much larger datasets.

McCoy, Hunter↗

2022 Operational Assessment Report - Argonne Leadership Computing Facility

This Operational Assessment Report describes how the Argonne Leadership Computing Facility (ALCF) met or exceeded every one of its goals for the calendar year (CY) 2022. In CY 2022, ALCF operated Theta, an Intel-based Cray XC40 system (11.7 petaflops) augmented with 24 NVIDIA DGX A100-based nodes (3.9 petaflops), that supports diverse workloads, integrating data analytics with artificial intelligence (AI) training and learning in a single platform; and Polaris, a 44-petaflop AMD and NVIDIA-based HPE Apollo 6500 Gen10+ system that provides a powerful new platform to prepare applications and workloads for Aurora, Argonne National Laboratory’s (Argonne’s) upcoming Intel-Hewlett Packard Enterprise (HPE) exascale supercomputer.

97 MATHEMATICS AND COMPUTING↗

Uncertainty Quantification and Sensitivity Analysis of Low-Dimensional Manifold via Co-Kurtosis PCA in Combustion Modeling

For multi-scale multi-physics applications e.g., the turbulent combustion code Pele, robust and accurate dimensionality reduction is crucial to solving problems at exascale and beyond. A recently developed technique, Co-Kurtosis based Principal Component Analysis (CoK-PCA) which leverages principal vectors of co-kurtosis, is a promising alternative to traditional PCA for complex chemical systems. To improve the effectiveness of this approach, we employ Artificial Neural Networks for reconstructing thermo-chemical scalars, species production rates, and overall heat release rates corresponding to the full state space. Our focus is on bolstering confidence in this deep learning based non-linear reconstruction through Uncertainty Quantification (UQ) and Sensitivity Analysis (SA). UQ involves quantifying uncertainties in inputs and outputs, while SA identifies influential inputs. One of the noteworthy challenges is the computational expense inherent in both endeavors. To address this, we employ the Monte Carlo methods to effectively quantify and propagate uncertainties in our reduced spaces while managing computational demands. Our research carries profound implications not only for the realm of combustion modeling but also for a broader audience in UQ. By showcasing the reliability and robustness of CoK-PCA in dimensionality reduction and deep learning predictions, we empower researchers and decision-makers to navigate complex combustion systems with greater confidence.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

DIMPLES: Distributed Influence Maximization for Pandemic pLanning on Exascale Systems

We study exascale parallel algorithms for the selection of intervention or monitoring strategies in massive realistic socio-technical networks through scalable Influence Maximization (InfMax) algorithms. We employ novel techniques to enable efficient scaling on up to 8k nodes of OLCF Frontier, with 65k AMD GPUs and 458k AMD CPU cores. Current state-of-the-art InfMax tools are limited to networks with only a few million actors (vertices) and a few hundred million interactions (edges). By overcoming these limitations, we show that our approach is capable of processing a realistic social contact network of the United States with 285 million nodes and about 8 billion edges. This two orders-of-magnitude improvement over the previous state-of-the-art is obtained by leveraging algorithmic advancements for the InfMax problem and designing several problem-specific approaches to overlap communication with computation, improve GPU efficiency, and lower the application’s memory requirements. We evaluate strong scaling for computing 10k most influential seeds using up to 8k nodes of an exascale system, and weak scaling from 128 to 8k system nodes for seed sets ranging from 625 to 40k seeds. We achieve the fastest-known runtime of 25 minutes while performing 48 million diffusion simulations totaling 2.31 petabytes to identify 40k influential seeds using 8k nodes, and take 5.75 minutes to identify 10k seeds while using 4k nodes.

Minutoli, Marco [Pacific Northwest National Labora↗