Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “extreme-scale computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Domain-decomposition nonlinear manifold reduced order model

This software combines nonlinear-manifold reduced order models (NM-ROMs) with domain decomposition (DD) techniques. NM-ROMs, which utilize a shallow, sparse autoencoder trained with full order model (FOM) snapshot data, approximate the FOM state on a nonlinear manifold. These models offer advantages over linear-subspace ROMs (LS-ROMs) particularly in scenarios with slowly decaying Kolmogorov n-width. However, the training of NM-ROMs involves a number of parameters that scale with the size of the FOM, and storing high-dimensional FOM snapshots can significantly increase the cost of ROM training for extreme-scale problems. To mitigate these costs, the software employs DD to partition the FOM into smaller subdomains, computes NM-ROMs for each, and then integrates these to form a global NM-ROM. This strategy offers multiple benefits: it enables parallel training of subdomain NM-ROMs, reduces the number of parameters needed, decreases the dimensional requirements of subdomain FOM training data, and allows for customization to the unique characteristics of each FOM subdomain. The use of a shallow, sparse autoencoder architecture in each subdomain NM-ROM facilitates the application of hyper-reduction (HR), simplifying the nonlinear complexities and enhancing computational speed. This software marks the inaugural application of NM-ROM combined with HR to a DD problem. It features an algebraic DD reformulation of the FOM, training of NM-ROMs with HR for each subdomain, and employs a sequential quadratic programming (SQP) solver for the evaluation of the coupled global NMROM. The effectiveness of the DD NM-ROM with HR is numerically demonstrated on the 2D steady-state Burgers' equation, showing an order of magnitude improvement in accuracy over the DD LS-ROM with HR.

Diaz, AlejandroN↗

GPU-enabled extreme-scale turbulence simulations: Fourier pseudo-spectral algorithms at the exascale using OpenMP offloading

Fourier pseudo-spectral methods for nonlinear partial differential equations are of wide interest in many areas of advanced computational science, including direct numerical simulation of three-dimensional (3-D) turbulence governed by the Navier-Stokes equations in fluid dynamics. This paper presents a new capability for simulating turbulence at a new record resolution up to 35 trillion grid points, on the world's first exascale computer, Frontier, comprising AMD MI250x GPUs with HPE's Slingshot interconnect and operated by the US Department of Energy's Oak Ridge Leadership Computing Facility (OLCF). Key programming strategies designed to take maximum advantage of the machine architecture involve performing almost all computations on the GPU which has the same memory capacity as the CPU, performing all-to-all communication among sets of parallel processes directly on the GPU, and targeting GPUs efficiently using OpenMP offloading for intensive number-crunching including 1-D Fast Fourier Transforms (FFT) performed using AMD ROCm library calls. With 99% of computing power on Frontier being on the GPU, leaving the CPU idle leads to a net performance gain via avoiding the overhead of data movement between host and device except when needed for some I/O purposes. Memory footprint including the size of communication buffers for MPI_ALLTOALL is managed carefully to maximize the largest problem size possible for a given node count. Detailed performance data including separate contributions from different categories of operations to the elapsed wall time per step are reported for five grid resolutions, from 2048 3 on a single node to 32768 3 on 4096 or 8192 nodes out of 9408 on the system. Both 1D and 2D domain decompositions which divide a 3D periodic domain into slabs and pencils respectively are implemented. The present code suite (labeled by the acronym GESTS, GPUs for Extreme Scale Turbulence Simulations) achieves a figure of merit (in grid points per second) exceeding goals set in the Center for Accelerated Application Readiness (CAAR) program for Frontier. The performance attained is highly favorable in both weak scaling and strong scaling, with notable departures only for 2048 3 where communication is entirely intra-node, and for 32768 3 , where a challenge due to small message sizes does arise. Communication performance is addressed further using a lightweight test code that performs all-to-all communication in a manner matching the full turbulence simulation code. Performance at large problem sizes is affected by both small message size due to high node counts as well as dragonfly network topology features on the machine, but is consistent with official expectations of sustained performance on Frontier. Overall, although not perfect, the scalability achieved at the extreme problem size of 32768 3 (and up to 8192 nodes — which corresponds to hardware rated at just under 1 exaflop/sec of theoretical peak computational performance) is arguably better than the scalability observed using prior state-of-the-art algorithms on Frontier's predecessor machine (Summit) at OLCF. New science results for the study of intermittency in turbulence enabled by this code and its extensions are to be reported separately in the near future.

3D fast Fourier transform↗

A Cast of Thousands: How the IDEAS Productivity Project Has Advanced Software Productivity and Sustainability

Computational and data-enabled science and engineering are revolutionizing advances throughout science and society, at all scales of computing. For example, teams in the U.S. Department of Energy’s Exascale Computing Project have been tackling new frontiers in modeling, simulation, and analysis by exploiting unprecedented exascale computing capabilities—building an advanced software ecosystem that supports next-generation applications and addresses disruptive changes in computer architectures. However, concerns are growing about the productivity of the developers of scientific software. Members of the Interoperable Design of Extreme-scale Application Software project serve as catalysts to address these challenges through fostering software communities, incubating and curating methodologies and resources, and disseminating knowledge to advance developer productivity and software sustainability. This article discusses how these synergistic activities are advancing scientific discovery—mitigating technical risks by building a firmer foundation for reproducible, sustainable science at all scales of computing, from laptops to clusters to exascale and beyond.

97 MATHEMATICS AND COMPUTING↗

MrHyDE v.1.0

SAND2024-01324O MrHyDE, which stands for Multi-resolution Hybridized Differential Equations, is a general-purpose C++ package for the solution of coupled multiphysics and multiscale systems on massively parallel computing systems. MrHyDE is designed to enable moving beyond forward simulation for multiscale applications which includes optimization, control, uncertainty quantification, and stochastic inversion. The framework provides interfaces to several packages within the Trilinos framework and leverages automatic differentiation to enable adjoint capabilities for large-scale, gradient-based optimization. MrHyDE provides automated multiscale capabilities through a subgrid model interface and multiscale Dirichlet-to-Neumann maps. For extreme-scale applications, MrHyDE provides in situ data-compression algorithms to reduce memory requirements while maintaining performance. MrHyDE is a general-purpose, computational framework for the solution of multiscale and multiphysics applications. It uses a combination of structure-preserving, physics-compatible discretizations, fully implicit methods, multi-resolution schemes, or fully explicit methods. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

SciDAC↗

Science & Technology Review: The Road to Exascale Computing

At Lawrence Livermore National Laboratory, we focus on science and technology research to ensure our nation’s security. We also apply that expertise to solve other important national problems in energy, bioscience, and the environment. Science & Technology Review is published eight times a year to communicate, to a broad audience, the Laboratory’s scientific and technological accomplishments in fulfilling its primary missions. The publication’s goal is to help readers understand these accomplishments and appreciate their value to the individual citizen, the nation, and the world. The Department of Energy’s Exascale Computing Project (ECP) and Lawrence Livermore’s RADIUSS (Rapid Application Development via an Institutional Universal Software Stack) initiative benefit from strategically developed software tools. The front cover shows a simulation of advection under twisting rotation that uses high-order finite elements from Livermore’s Modular Finite Element Methods (MFEM) software library and GLVis visualization tool. On the back cover, the logo (also created with GLVis) for the MFEM project illustrates the curved mesh and sub-element resolution used in high-order simulations. MFEM and GLVis are key components of the ECP’s co-design Center for Efficient Exascale Discretizations (CEED) and RADIUSS. MFEM is also part of ECP’s Extreme-Scale Scientific Software Development Kit (xSDK).

97 MATHEMATICS AND COMPUTING↗

A fast and accurate domain decomposition nonlinear manifold reduced order model

Here, this paper integrates nonlinear-manifold reduced order models (NM-ROMs) with domain decomposition (DD). NM ROMs approximate the full order model (FOM) state in a nonlinear-manifold by training a shallow, sparse autoencoder using FOM snapshot data. These NM-ROMs can be advantageous over linear-subspace ROMs (LS-ROMs) for problems with slowly decaying Kolmogorov n-width. However, the number of NM-ROM parameters that need to be trained scales with the size of the FOM. Moreover, for “extreme-scale” problems, the storage of high-dimensional FOM snapshots alone can make ROM training expensive. To alleviate the training cost, this paper applies DD to the FOM, computes NM-ROMs on each subdomain, and couples them to obtain a global NM-ROM. This approach has several advantages: Subdomain NM-ROMs can be trained in parallel, involve fewer parameters to be trained than global NM-ROMs, require smaller subdomain FOM dimensional training data, and can be tailored to subdomain specific features of the FOM. The shallow, sparse architecture of the autoencoder used in each subdomain NM-ROM allows application of hyper-reduction (HR), reducing the complexity caused by nonlinearity and yielding computational speedup of the NM-ROM. This paper provides the first application of NM-ROM (with HR) to a DD problem. In particular, this paper details an algebraic DD reformulation of the FOM, training a NM-ROM with HR for each sub domain, and a sequential quadratic programming (SQP) solver to evaluate the coupled global NM-ROM. Theoretical convergence results for the SQP method and a priori and a posteriori error estimates for the DD NM-ROM with HR are provided. The proposed DD NM-ROM with HR approach is numerically compared to a DD LS-ROM with HR on the 2D steady-state Burgers’ equation, showing an order of magnitude improvement in accuracy of the proposed DD NM-ROM over the DD LS-ROM.

97 MATHEMATICS AND COMPUTING↗

Tensor Network Quantum Virtual Machine for Simulating Quantum Circuits at Exascale

The numerical simulation of quantum circuits is an indispensable tool for development, verification, and validation of hybrid quantum-classical algorithms intended for near-term quantum co-processors. The emergence of exascale high-performance computing (HPC) platforms presents new opportunities for pushing the boundaries of quantum circuit simulation. Here, we present a modernized version of the Tensor Network Quantum Virtual Machine (TNQVM) that serves as the quantum circuit simulation backend in the eXtreme-scale ACCelerator (XACC) framework. The new version is based on the scalable tensor network processing library ExaTN (Exascale Tensor Networks). It provides multiple configurable quantum circuit simulators that perform either an exact quantum circuit simulation via the full tensor network contraction or an approximate simulation via a suitably chosen tensor factorization scheme. Upon necessity, stochastic noise modeling from real quantum processors is incorporated into the simulations by modeling quantum channels with Kraus tensors. By combining the portable XACC quantum programming frontend and the scalable ExaTN numerical processing backend, we introduce an end-to-end virtual quantum development environment that can scale from laptops to future exascale platforms. We report initial benchmarks of our framework, which include a demonstration of the distributed execution, incorporation of quantum decoherence models, and simulation of the random quantum circuits used for the certification of quantum supremacy on Google’s Sycamore superconducting architecture.

Nguyen, Thien↗

BCSR on GPU: A Way Forward Extreme-scale Graph Processing on Accelerator-enabled Frontier Supercomputer

Handling large graphs in a distributed environment requires effective partitioning across processors and efficient management of local partitions. In 2D partitioning, local graphs often become too sparse, making memory-efficient data structures crucial. Using the Compressed Sparse Row (CSR) format wastes space, especially for > 83% of vertices with empty edges for the sparse graphs. This study explores bit-CSR (BCSR), a modified CSR representation, on GPUs to reduce memory usage in graph computations. We achieved 16.67% memory savings on a sparse rmat dataset with 268 million vertices and 357 million edges, without performance degradation, supported by both theoretical and experimental storage savings of 33%. However, we observed a 1.7× slowdown in degree lookup times due to bitwise operations on AMD CPUs. This analysis highlights the potential of BCSR on GPUs for improving Graph500 benchmark performance on GPU-accelerated systems, such as the Frontier supercomputer.

Sattar, Naw Safrin↗

Summarizing the interoperabilities between xSDK members

Rapid, efficient production of high-quality, sustainable extreme-scale scientific applications is best accomplished using a rich ecosystem of state-of-the art reusable libraries, tools, lightweight frameworks, and defined software methodologies, developed by a community of scientists who are striving to identify, adapt, and adopt best practices in software engineering. The vision of the xSDK is to provide infrastructure for and interoperability of a collection of related and complementary software elements — developed by diverse, independent teams throughout the high-performance computing (HPC) community — that provide the building blocks, tools, models, processes, and related artifacts for rapid and efficient development of high-quality applications. A primary component, and challenge, of the xSDK is to improve interoperability among software libraries and domain components. This document summarizes the current and planned interoperabilities of the twenty-three xSDK member packages as of March 2021. Additionally, the status of the xSDK example codes that demonstrate and test the interoperabilities within the xSDK is provided.

97 MATHEMATICS AND COMPUTING↗

Rapid Optimization of Total Variation with Applications in Imaging, Additive Manufacturing, and Qualification

Total Variation optimization penalizes the gradient of a control variable or state. While this work focuses on image processing in particular, it has also found applications in inverse problems and topology optimization. In image processing, the goal is to maintain faithfulness to the original image while denoising and/or deblurring. Additionally, bilevel optimization over the spatially varying regularization weights can illuminate interfaces such as damage regions and other anomalies. We will address two fundamental challenges with TV-optimization: (i) the typical slow convergence of existing TV-optimization methods, and (ii) the selection of spatially varying TV parameters to promote interface detection. Additionally, we will apply such techniques to image data collected in additive manufacturing. In said context, stochasticity in build events induces flaws in the manufactured piece, compromising the integrity of said part. There is a critical need for in-situ monitoring to spot anomalies once they form, and in this setting we apply our total variation and hyperparameter solvers. We will develop a customized algorithm based on for extreme-scale TV-optimization that achieves super-linear or quadratic-convergence, a critical property for real-time, image-by-image analysis. A worst-case outcome is a preprocessing step that enhances image quality in-situ, specifically for out-of-focus and noisy images.

36 MATERIALS SCIENCE↗

Aerodynamic Rotor Design for a 25 MW Offshore Downwind Turbine

Continuously increasing offshore wind turbine scales require rotor designs that maximize power and performance. Downwind rotors offer advantages in lower mass due to reduced potential for tower strike, and is especially true at large scales, e.g., for a 25 MW turbine. In this study, three 25 MW downwind rotors, each with different prescribed lift coefficient distributions were designed (chord, geometry, and twist) and compared to maximize power production at unprecedented scales and Reynolds numbers, including a new approach to optimize rotor tilt and coning based on aeroelastic effects. To achieve this objective the design process was focused on achieving high power coefficients, while maximizing swept area and minimizing blade mass. Maximizing swept area was achieved by prescribing pre-cone and shaft tilt angles to ensure the aeroelastic orientation when the blades point upwards was nearly vertical at nearly rated conditions. Maximizing the power coefficient was achieved by prescribing axial induction factor and lift coefficient distributions which were then used as inputs for an inverse rotor design tool. The resulting rotors were then simulated to compare performance and subsequently optimized for minimum rotor mass. To achieve these goals, a high Reynolds number design space was developed using computational predictions as well as new empirical correlations for flatback airfoil drag and maximum lift. Within this design space, three rotors of small, medium and large chords were considered for clean airfoil conditions (effects of premature transition were also considered but did not significantly modify the design space). The results indicated that the medium chord design provided the best performance, producing the highest power in Region 2 from simulations while resulting in the lowest rotor mass, both of which support minimum LCOE. The methodology developed herein can be used for the design of other extreme-scale (upwind and downwind) turbines.

downwind rotors↗

ExaFEL: extreme-scale real-time data processing for X-ray free electron laser science

ExaFEL is an HPC-capable X-ray Free Electron Laser (XFEL) data analysis software suite for both Serial Femtosecond Crystallography (SFX) and Single Particle Imaging (SPI) developed in collaboration with the Linac Coherent Lightsource (LCLS), Lawrence Berkeley National Laboratory (LBNL) and Los Alamos National Laboratory. ExaFEL supports real-time data analysis via a cross-facility workflow spanning LCLS and HPC centers such as NERSC and OLCF. Our work therefore constitutes initial path-finding for the US Department of Energy's (DOE) Integrated Research Infrastructure (IRI) program. We present the ExaFEL team's 7 years of experience in developing real-time XFEL data analysis software for the DOE's exascale supercomputers. We present our experiences and lessons learned with the Perlmutter and Frontier supercomputers. Furthermore we outline essential data center services (and the implications for institutional policy) required for real-time data analysis. Finally we summarize our software and performance engineering approaches and our experiences with NERSC's Perlmutter and OLCF's Frontier systems. This work is intended to be a practical blueprint for similar efforts in integrating exascale compute resources into other cross-facility workflows.

59 BASIC BIOLOGICAL SCIENCES↗

Challenges of and Opportunities for a Large Diverse Software Team

A large software team consisting of members with different expertise, skillsets, personalities, ethnicities, and involving collaboration on a large and complex software product presents many technical and cultural challenges, but also provides unique opportunities. In this article, we discuss the essential issues we faced when successfully transforming a collection of various independently developed software libraries into one large integrated product: the eXtreme-scale scientific Software Development Kit (xSDK). Furthermore, we argue it is just as important to pay attention to cultural challenges, such as establishment of reliable communication channels that considers, among others, differences in personalities and backgrounds as well as overcoming geographical separation and time-zone distribution when collaborating, as technical challenges. Finally, we discuss opportunities stemming from participating in a large diverse software team, such as increased internal expertise, variety of skillsets, broadened connections to external experts, and access to a larger pool of ideas or solutions.

97 MATHEMATICS AND COMPUTING↗

RISE: Reducing I/O Contention in Staging-based Extreme-Scale In-situ Workflows

While in-situ workflow formulations have addressed some of the data-related challenges associated with extreme-scale scientific workflows, these workflows involve complex interactions and different modes of data exchange. In the context of increasing system complexity, such workflows present significant resource management challenges, requiring complex cost-performance tradeoffs. This paper presents RISE, an intelligent staging-based data management middleware, which builds on the DataSpaces framework and performs intelligent scheduling of data management operations to reduce I/O contention. In RISE, data are always written immediately to local buffers to reduce the effect of the transfer impact upon application performance. RISE identifies applications’ data access patterns and moves data towards data consumers only when the network is expected to be idle, reducing the impact of asynchronous background data movement upon critical data read/write requests. Here, we experimentally demonstrate that RISE can take advantage of staging nodes to offload data during writes without degrading application data movement performance.

97 MATHEMATICS AND COMPUTING↗

Parked aeroelastic field rotor response for a 20% scaled demonstrator of a 13‐MW downwind turbine

Abstract Aeroelastic parked testing of a unique downwind two‐bladed subscale rotor was completed to characterize the response of an extreme‐scale 13‐MW turbine in high‐wind parked conditions. A 20% geometric scaling was used resulting in scaled 20‐m‐long blades, whose structural and stiffness properties were designed using aeroelastic scaling to replicate the nondimensional structural aeroelastic deflections and dynamics that would occur for a lightweight, downwind 13‐MW rotor. The subscale rotor was mounted and field tested on the two‐bladed Controls Advanced Research Turbine (CART2) at the National Renewable Energy Laboratory's Flatiron Campus (NREL FC). The parked testing of these highly flexible blades included both pitch‐to‐run and pitch‐to‐feather configurations with the blades in the horizontal braked orientation. The collected experimental data includes the unsteady flapwise root bending moments and tip deflections as a function of inflow wind conditions. The bending moments are based on strain gauges located in the root section, whereas the tip deflections are captured by a video camera on the hub of the turbine pointed toward the tip of the blade. The experimental results are compared against computational predictions generated by FAST, a wind turbine simulation software, for the subscale and full‐scale models with consistent unsteady wind fields. FAST reasonably predicted the bending moments and deflections of the experimental data in terms of both the mean and standard deviations. These results demonstrate the efficacy of the first such aeroelastically scaled turbine test and demonstrate that a highly flexible lightweight downwind coned rotor can be designed to withstand extreme loads in parked conditions.

17 WIND ENERGY↗

Constructing a new predictive scaling formula for ITER's divertor heat-load width informed by a simulation-anchored machine learning

Understanding and predicting divertor heat-load width λq is a critically important problem for an easier and more robust operation of ITER with high fusion gain. Previous predictive simulation data for λ q using the extreme-scale edge gyrokinetic code XGC1 [S. Ku et al., Phys. Plasmas 25, 056107 (2018)] in the electrostatic limit under attached divertor plasma conditions in three major US tokamaks [C. S. Chang et al., Nucl. Fusion 57, 116023 (2017)] reproduced the Eich and Goldston attached-divertor formula results [formula #14 in T. Eich et al., Nucl. Fusion 53, 093031 (2013) and R. J. Goldston, Nucl. Fusion 52, 013009 (2012)] and furthermore predicted over six times wider λ q than the maximal Eich and Goldston formula predictions on a full-power (Q = 10) scenario ITER plasma. After adding data from further predictive simulations on a highest current JET and highest-current Alcator C-Mod, a machine learning program is used to identify a new scaling formula for λ q as a simple modification to the Eich formula #14, which reproduces the Eich scaling formula for the present tokamaks and which embraces the wide λ q XGC for the full-current Q = 10 ITER plasma. Additionally, the new formula is then successfully tested on three more ITER plasmas: two corresponding to long burning scenarios with Q = 5 and one at low plasma current to be explored in the initial phases of ITER operation. The new physics that gives rise to the wider λ q XGC is identified to be the weakly collisional, trapped-electron-mode turbulence across the magnetic separatrix, which is known to be an efficient transporter of the electron heat and mass. Electromagnetic turbulence and high-collisionality effects on the new formula are the next study topics for XGC1.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗