Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “extreme-scale”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Optimizing Data Movement for GPU-Based In-Situ Workflow Using GPUDirect RDMA

The extreme-scale computing landscape is increasingly dominated by GPU-accelerated systems. At the same time, in-situ workflows that employ memory-to-memory inter-application data exchanges have emerged as an effective approach for leveraging these extreme-scale systems. In the case of GPUs, GPUDirect RDMA enables third-party devices, such as network interface cards, to access GPU memory directly and has been adopted for intra-application communications across GPUs. In this paper, we present an interoperable framework for GPU-based in-situ workflows that optimizes data movement using GPUDirect RDMA. Specifically, we analyze the characteristics of the possible data movement pathways between GPUs from an in-situ workflow perspective, and design a strategy that maximizes throughput. Furthermore, we implement this approach as an extension of the DataSpaces data staging service, and experimentally evaluate its performance and scalability on a current leadership GPU cluster. The performance results show that the proposed design reduces data-movement time by up to 53% and 40% for the sender and receiver, respectively, and maintains excellent scalability for up to 256 GPUs.

Zhang, Bo↗

Collaborative: in situ visual analytics technologies for extreme scale combustion simulations

This project aims to drastically enhance the usability of in situ analysis and visualization for extreme-scale scientific simulations. Current exascale computing capabilities promise to offer greater predictive ability of simulations and to further push the frontiers of science and technology. However, to validate the simulation output at extreme scale, examine the modeled phenomena, and discover previously unknowns from the output data, the output must be reduced or transformed in situ as it is being generated during the simulation such that the amount of data to examine and store is kept to a minimum. Such in situ approaches allow us to process and analyze the data and any embedded geometry to an extent that would be prohibitively expensive, if not impossible, to perform as a post hoc task. While in situ processing has been demonstrated to be a feasible and promising approach, its full potential has not yet been leveraged. In this project, we have developed comprehensive enhancements to in situ technology based on probability distributions in data. Our research focuses on jointly developing new ways of interacting with massive statistical samples while creatively utilizing new state-of-the-art computational resources to push the boundaries of in situ exploration. Moreover, we have developed new time-dependent techniques to enable previously unattainable capabilities in areas such as intelligent simulation steering and precise feature identification. We have experimentally studied our design and implementation at NERSC and OLCF, and are able to leverage existing in situ infrastructures whenever possible. While the exemplar in this project is combustion, many other fields for which turbulent transport is important, e.g., fusion, climate, astrophysics among others, encounter similar issues as simulations scale up to the exascale. This project shows its potential to generate high impact on DOE missions since the resulting technology promises to improve scientists’ ability to rapidly and correctly interpret and tune extreme-scale simulations, leading to new scientific understanding and advancements.

97 MATHEMATICS AND COMPUTING↗

Aerodynamic Rotor Design for a 25 MW Offshore Downwind Turbine

Continuously increasing offshore wind turbine scales require rotor designs that maximize power and performance. Downwind rotors offer advantages in lower mass due to reduced potential for tower strike, and is especially true at large scales, e.g., for a 25 MW turbine. In this study, three 25 MW downwind rotors, each with different prescribed lift coefficient distributions were designed (chord, geometry, and twist) and compared to maximize power production at unprecedented scales and Reynolds numbers, including a new approach to optimize rotor tilt and coning based on aeroelastic effects. To achieve this objective the design process was focused on achieving high power coefficients, while maximizing swept area and minimizing blade mass. Maximizing swept area was achieved by prescribing pre-cone and shaft tilt angles to ensure the aeroelastic orientation when the blades point upwards was nearly vertical at nearly rated conditions. Maximizing the power coefficient was achieved by prescribing axial induction factor and lift coefficient distributions which were then used as inputs for an inverse rotor design tool. The resulting rotors were then simulated to compare performance and subsequently optimized for minimum rotor mass. To achieve these goals, a high Reynolds number design space was developed using computational predictions as well as new empirical correlations for flatback airfoil drag and maximum lift. Within this design space, three rotors of small, medium and large chords were considered for clean airfoil conditions (effects of premature transition were also considered but did not significantly modify the design space). The results indicated that the medium chord design provided the best performance, producing the highest power in Region 2 from simulations while resulting in the lowest rotor mass, both of which support minimum LCOE. The methodology developed herein can be used for the design of other extreme-scale (upwind and downwind) turbines.

downwind rotors↗

A fast and accurate domain decomposition nonlinear manifold reduced order model

Here, this paper integrates nonlinear-manifold reduced order models (NM-ROMs) with domain decomposition (DD). NM ROMs approximate the full order model (FOM) state in a nonlinear-manifold by training a shallow, sparse autoencoder using FOM snapshot data. These NM-ROMs can be advantageous over linear-subspace ROMs (LS-ROMs) for problems with slowly decaying Kolmogorov n-width. However, the number of NM-ROM parameters that need to be trained scales with the size of the FOM. Moreover, for “extreme-scale” problems, the storage of high-dimensional FOM snapshots alone can make ROM training expensive. To alleviate the training cost, this paper applies DD to the FOM, computes NM-ROMs on each subdomain, and couples them to obtain a global NM-ROM. This approach has several advantages: Subdomain NM-ROMs can be trained in parallel, involve fewer parameters to be trained than global NM-ROMs, require smaller subdomain FOM dimensional training data, and can be tailored to subdomain specific features of the FOM. The shallow, sparse architecture of the autoencoder used in each subdomain NM-ROM allows application of hyper-reduction (HR), reducing the complexity caused by nonlinearity and yielding computational speedup of the NM-ROM. This paper provides the first application of NM-ROM (with HR) to a DD problem. In particular, this paper details an algebraic DD reformulation of the FOM, training a NM-ROM with HR for each sub domain, and a sequential quadratic programming (SQP) solver to evaluate the coupled global NM-ROM. Theoretical convergence results for the SQP method and a priori and a posteriori error estimates for the DD NM-ROM with HR are provided. The proposed DD NM-ROM with HR approach is numerically compared to a DD LS-ROM with HR on the 2D steady-state Burgers’ equation, showing an order of magnitude improvement in accuracy of the proposed DD NM-ROM over the DD LS-ROM.

97 MATHEMATICS AND COMPUTING↗

Multifacets of lossy compression for scientific data in the Joint-Laboratory of Extreme Scale Computing

The Joint Laboratory on Extreme-Scale Computing (JLESC) was initiated at the same time lossy compression for scientific data became an important topic for the scientific communities. The teams involved in the JLESC played and are still playing an important role in developing the research, techniques, methods, and technologies making lossy compression for scientific data a key tool for scientists and engineers. Here, in this paper, we present the evolution of lossy compression for scientific data from 2015, describing the situation before the JLESC started, the evolution of this discipline in the past 8 years (until 2023) through the prism of the JLESC collaborations on this topic and some of the remaining open research questions.

Compression for AI↗

Solution-phase sample-averaged single-particle spectroscopy of quantum emitters with femtosecond resolution

Here, the development of many optical quantum technologies depends on the availability of solid-state single quantum emitters with near-perfect optical coherence. However, a standing issue that limits systematic improvement is the significant sample heterogeneity and lack of mechanistic understanding of microscopic energy flow at the single emitter level and ultrafast timescales. Here we develop solution-phase single-particle pump-probe spectroscopy with photon correlation detection that captures sample-averaged dynamics in single molecules and/or defect states with unprecedented clarity at femtosecond resolution. We apply this technique to single quantum emitters in two-dimensional hexagonal boron nitride, which suffers from significant heterogeneity and low quantum efficiency. From millisecond to nanosecond timescales, the translation diffusion, metastable-state-related bunching shoulders, rotational dynamics, and antibunching features are disentangled by their distinct photon-correlation timescales, which collectively quantify the normalized two-photon emission quantum yield. Leveraging its femtosecond resolution, spectral selectivity and ultralow noise (two orders of magnitude improvement over solid-state methods), we visualize electron-phonon coupling in the time domain at the single defect level, and discover the acceleration of polaronic formation driven by multi-electron excitation. Corroborated with results from a theoretical polaron model, we show how this translates to sample-averaged photon fidelity characterization of cascaded emission efficiency and optical decoherence time. Our work provides a framework for ultrafast spectroscopy in single emitters, molecules, or defects prone to photoluminescence intermittency and heterogeneity, opening new avenues of extreme-scale characterization and synthetic improvements for quantum information applications.

36 MATERIALS SCIENCE↗

A Cast of Thousands: How the IDEAS Productivity Project Has Advanced Software Productivity and Sustainability

Computational and data-enabled science and engineering are revolutionizing advances throughout science and society, at all scales of computing. For example, teams in the U.S. Department of Energy’s Exascale Computing Project have been tackling new frontiers in modeling, simulation, and analysis by exploiting unprecedented exascale computing capabilities—building an advanced software ecosystem that supports next-generation applications and addresses disruptive changes in computer architectures. However, concerns are growing about the productivity of the developers of scientific software. Members of the Interoperable Design of Extreme-scale Application Software project serve as catalysts to address these challenges through fostering software communities, incubating and curating methodologies and resources, and disseminating knowledge to advance developer productivity and software sustainability. This article discusses how these synergistic activities are advancing scientific discovery—mitigating technical risks by building a firmer foundation for reproducible, sustainable science at all scales of computing, from laptops to clusters to exascale and beyond.

97 MATHEMATICS AND COMPUTING↗

DDStore: Distributed Data Store for Scalable Training of Graph Neural Networks on Large Atomistic Modeling Datasets

Graph neural networks (GNNs) are a class of Deep Learning models used in designing atomistic materials for effective screening of large chemical spaces. To ensure robust prediction, GNN models must be trained on large volumes of atomistic data on leadership class supercomputers. Even with the advent of modern architectures that consist of multiple storage layers that include node-local NVMe devices in addition to device memory for caching large datasets, extreme-scale model training faces I/O challenges at scale.We present DDStore, an in-memory distributed data store designed for GNN training on large-scale graph data. DDStore provides a hierarchical, distributed, data caching technique that combines data chunking, replication, low-latency random access, and high throughput communication. DDStore achieves near-linear scaling for training a GNN model using up to 1000 GPUs on the Summit and Perlmutter supercomputers, and reaches up to a 6.15x reduction in GNN training time compared to state-of-the-art methodologies.

Choi, Jong Youl↗

MrHyDE v.1.0

SAND2024-01324O MrHyDE, which stands for Multi-resolution Hybridized Differential Equations, is a general-purpose C++ package for the solution of coupled multiphysics and multiscale systems on massively parallel computing systems. MrHyDE is designed to enable moving beyond forward simulation for multiscale applications which includes optimization, control, uncertainty quantification, and stochastic inversion. The framework provides interfaces to several packages within the Trilinos framework and leverages automatic differentiation to enable adjoint capabilities for large-scale, gradient-based optimization. MrHyDE provides automated multiscale capabilities through a subgrid model interface and multiscale Dirichlet-to-Neumann maps. For extreme-scale applications, MrHyDE provides in situ data-compression algorithms to reduce memory requirements while maintaining performance. MrHyDE is a general-purpose, computational framework for the solution of multiscale and multiphysics applications. It uses a combination of structure-preserving, physics-compatible discretizations, fully implicit methods, multi-resolution schemes, or fully explicit methods. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

SciDAC↗

Domain-decomposition nonlinear manifold reduced order model

This software combines nonlinear-manifold reduced order models (NM-ROMs) with domain decomposition (DD) techniques. NM-ROMs, which utilize a shallow, sparse autoencoder trained with full order model (FOM) snapshot data, approximate the FOM state on a nonlinear manifold. These models offer advantages over linear-subspace ROMs (LS-ROMs) particularly in scenarios with slowly decaying Kolmogorov n-width. However, the training of NM-ROMs involves a number of parameters that scale with the size of the FOM, and storing high-dimensional FOM snapshots can significantly increase the cost of ROM training for extreme-scale problems. To mitigate these costs, the software employs DD to partition the FOM into smaller subdomains, computes NM-ROMs for each, and then integrates these to form a global NM-ROM. This strategy offers multiple benefits: it enables parallel training of subdomain NM-ROMs, reduces the number of parameters needed, decreases the dimensional requirements of subdomain FOM training data, and allows for customization to the unique characteristics of each FOM subdomain. The use of a shallow, sparse autoencoder architecture in each subdomain NM-ROM facilitates the application of hyper-reduction (HR), simplifying the nonlinear complexities and enhancing computational speed. This software marks the inaugural application of NM-ROM combined with HR to a DD problem. It features an algebraic DD reformulation of the FOM, training of NM-ROMs with HR for each subdomain, and employs a sequential quadratic programming (SQP) solver for the evaluation of the coupled global NMROM. The effectiveness of the DD NM-ROM with HR is numerically demonstrated on the 2D steady-state Burgers' equation, showing an order of magnitude improvement in accuracy over the DD LS-ROM with HR.

Diaz, AlejandroN↗

Globus service enhancements for exascale applications and facilities

Many extreme-scale applications require the movement of large quantities of data to, from, and among leadership computing facilities, as well as other scientific facilities and the home institutions of facility users. These applications, particularly when leadership computing facilities are involved, can touch upon edge cases (e.g., terabyte files) that had not been a focus of previous Globus optimization work, which had emphasized rather the movement of many smaller (megabyte to gigabyte) files. We report here on how automated client-driven chunking can be used to accelerate both the movement of large files and the integrity checking operations that have proven to be essential for large data transfers. In conclusion, we present detailed performance studies that provide insights into the benefits of these modifications in a range of file transfer scenarios.

97 MATHEMATICS AND COMPUTING↗

libEnsemble: A complete Python toolkit for dynamic ensembles of calculations

Almost all science and engineering applications eventually stop scaling: their runtime no longer decreases as available computational resources increase. Therefore, many applications will struggle to efficiently use emerging extreme-scale high-performance, parallel, and distributed systems. libEnsemble is a complete Python toolkit and workflow system for intelligently driving ensembles of experiments or simulations at massive scales. It enables and encourages multidisciplinary design, decision, and inference studies portably running on laptops, clusters, and supercomputers.

97 MATHEMATICS AND COMPUTING↗

Integrated Research Infrastructure Architecture Blueprint Activity (Final Report 2023)

The complexity of scientific pursuits is increasing rapidly with aspects that require dynamic integration of experiment, observation, theory, modeling, simulation, visualization, machine learning (ML), artificial intelligence (AI), and analysis. Research projects across the Department of Energy (DOE) are increasingly data and compute intensive. Innovative research teams are accelerating the pace of discovery by using high-performance computational and data tools in their research workflows and leveraging multiple research infrastructures. Additionally, several recent high-level U.S. government reports underscore the necessity of a new advanced computing ecosystem for international competitiveness and national security. International competitors are moving forward with major research infrastructure integration efforts that seek to capture a competitive advantage in the global innovation race. Owing to its unparalleled constellation of world-class experimental and observational facilities and high-performance and extreme-scale computational, data, and networking infrastructure, DOE is positioned to be a global leader in this new era of integrated science. However, this new integration paradigm will demand continuing evolution to ensure the U.S. remains a global leader in research and innovation. The DOE Office of Science (SC) has seized on the strategic importance of integration and has adopted a vision for Integrated Research Infrastructure (IRI): To empower researchers to meld DOE’s world-class research tools, infrastructure, and user facilities seamlessly and securely in novel ways to radically accelerate discovery and innovation. To respond to the evolving computational requirements of research and the competitive international innovation landscape, experimental facilities could be connected with high performance computing resources for near real-time analysis, and resources should be provided for merging enormous and diverse data for AI/ML techniques and analysis.

97 MATHEMATICS AND COMPUTING↗

A Contextually-Aware Sensitivity Analysis to Guide the Design of Randomized Least Squares Solvers in Applications

Our work on the DOE-sponsored project “A Contextually-Aware Sensitivity Analysis to Guide the Design of Randomized Least Squares Solvers in Applications,” was an effort to address critical challenges in nu merical computing and its applications to optimization. The increasing demand for robust and scalable solutions to large-scale linear algebra problems has highlighted the limitations of traditional approaches, particularly in heterogeneous and extreme-scale computing environments. Randomized Numerical Linear Algebra (RandNLA) offers a promising framework to address these challenges, and this proposal builds on this foundation by introducing innovations in sensitivity analysis and computational adaptability.

97 MATHEMATICS AND COMPUTING↗

Rapid Optimization of Total Variation with Applications in Imaging, Additive Manufacturing, and Qualification

Total Variation optimization penalizes the gradient of a control variable or state. While this work focuses on image processing in particular, it has also found applications in inverse problems and topology optimization. In image processing, the goal is to maintain faithfulness to the original image while denoising and/or deblurring. Additionally, bilevel optimization over the spatially varying regularization weights can illuminate interfaces such as damage regions and other anomalies. We will address two fundamental challenges with TV-optimization: (i) the typical slow convergence of existing TV-optimization methods, and (ii) the selection of spatially varying TV parameters to promote interface detection. Additionally, we will apply such techniques to image data collected in additive manufacturing. In said context, stochasticity in build events induces flaws in the manufactured piece, compromising the integrity of said part. There is a critical need for in-situ monitoring to spot anomalies once they form, and in this setting we apply our total variation and hyperparameter solvers. We will develop a customized algorithm based on for extreme-scale TV-optimization that achieves super-linear or quadratic-convergence, a critical property for real-time, image-by-image analysis. A worst-case outcome is a preprocessing step that enhances image quality in-situ, specifically for out-of-focus and noisy images.

36 MATERIALS SCIENCE↗

Simulation Center for Runaway Electron Avoidance and Mitigation (SCREAM SciDAC) (Technical Final Report)

Runaway electrons can severely damage the plasma facing components on ITER during a major disruption and pose a major risk for tokamak fusion. It has been recognized that an adequate disruption mitigation system (DMS) is essential for the safe operation of ITER. The United States is responsible for the design and implementation of the disruption mitigation system on ITER, and in July 2016 the Simulation Center for Runaway Electron Avoidance and Mitigation (SCREAM) was launched by DOE, in a joint Fusion Energy Sciences (FES) and Advanced Scientific Computing Research (ASCR) collaboration. SCREAM was a comprehensive theory and simulation SciDAC center that provided physics guidance in the avoidance and mitigation of runaway electrons, and in tandem with domestic and international experiments, helped establish the qualitative and quantitative bases for safe operational scenarios and viable mitigation techniques. The SCREAM center assembled a national team of experts in runaway electron physics, tokamak disruptions, magnetohydrodynamic (MHD) simulation, and advanced algorithms and computing. The team combined advanced simulation and analysis capability facilitated by direct participation of ASCR SciDAC institutes with theoretical models and code development by FES scientists to focus on the runaway risk for ITER and tokamaks in general. The research scope was focussed on integrated simulations of kinetic runaway electrons, including MHD and fluid models of impurity transport, within a research plan guided by theory. The specific research tasks were (1) establish the fundamental physics of runaway generation, saturation, and dynamical evolution in a tokamak; (2) examine the critical path toward runaway avoidance; and (3) investigate the viability and effectiveness of the leading candidate schemes for runaway mitigation. In all three areas, members of the team carried out scoping studies that established the readiness for rapid and critical advances, especially in the deployment and further development of large-to extreme-scale simulation tools. Our multi-pronged computational approach included (1) relativistic Fokker-Planck solvers with discretization in phase space, (2) self-consistent particle-in-cell techniques, (3) particle-based Monte-Carlo, and (4) MHD-particle hybrid simulations. Cross-check between these different methods provided an additional means for verification and further bolstered the fidelity of our physics prediction. Validation against experimental results brings confidence to the predictive capability for ITER and frequently leads to new ideas for understanding and mitigating the thermal quench driven runaway electron phenomenon.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Tachyon: Intelligent Multi-Scale Modeling of Distributed Resilient Infrastructure and Workflows for Data Intensive HEP Analyses

The DOE High Energy Physics (HEP) program in Neutrino and Collider science drives data-intensive science and simulation on extreme-scale platforms. Modeling and optimizing the complex distributed components from experimental to leadership computing facilities are essential for HEP workflows to achieve required response times and resilience under various conditions. Tachyon proposes a framework for scalable modeling, simulation, and validation of key performance characteristics for the distributed infrastructure between FNAL and ALCF, along with associated HEP workflows.

Carothers, Chris [Rensselaer Poly.]↗

Tachyon: Intelligent Multi-Scale Modeling of Distributed Resilient Infrastructure and Workflows for Data Intensive HEP Analyses

The DOE High Energy Physics (HEP) program in Neutrino and Collider science drives data-intensive science and simulation on extreme-scale platforms. Modeling and optimizing the complex distributed components from experimental to leadership computing facilities are essential for HEP workflows to achieve required response times and resilience under various conditions. Tachyon proposes a framework for scalable modeling, simulation, and validation of key performance characteristics for the distributed infrastructure between FNAL and ALCF, along with associated HEP workflows.

Carothers, Chris [Rensselaer Poly.]↗