Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “extreme-scale computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

81 records · Page 5

Challenges of and Opportunities for a Large Diverse Software Team

A large software team consisting of members with different expertise, skillsets, personalities, ethnicities, and involving collaboration on a large and complex software product presents many technical and cultural challenges, but also provides unique opportunities. In this article, we discuss the essential issues we faced when successfully transforming a collection of various independently developed software libraries into one large integrated product: the eXtreme-scale scientific Software Development Kit (xSDK). Furthermore, we argue it is just as important to pay attention to cultural challenges, such as establishment of reliable communication channels that considers, among others, differences in personalities and backgrounds as well as overcoming geographical separation and time-zone distribution when collaborating, as technical challenges. Finally, we discuss opportunities stemming from participating in a large diverse software team, such as increased internal expertise, variety of skillsets, broadened connections to external experts, and access to a larger pool of ideas or solutions.

97 MATHEMATICS AND COMPUTING↗

Modeling Impacts of Resilience Architectures for Extreme-Scale Storage Systems

The Scientific Modeling and Simulation of Storage project focused on improving our ability to store information reliably. The project, done jointly at the University of Wisconsin-Madison, Los Alamos National Laboratories, and Sandia National Laboratories, had the main goal of transforming the art of reliable storage system construction into a more precise science. Along those lines, the project developed numerous tools to expose problems in existing systems; these tools, when applied to the state of the art, uncovered numerous flaws in existing approaches, showing that they both are error prone and in some cases contain serious design flaws. The project also led to the creation of new approaches in reliable system construction. One such approach shows how a storage system can adapt its behavior on the fly to deliver higher performance when there are few underlying failures, but then act more conservatively when more failures are occurring to ensure the safe storage of user data. Another approach develops a storage system consistency model which both delivers high performance but high levels of reliability, a significant advance over previous approaches. Overall, the tools developed within the project advanced the science of reliable storage design, and the new approaches that resulted advanced the science of storage system construction.

97 MATHEMATICS AND COMPUTING↗

RISE: Reducing I/O Contention in Staging-based Extreme-Scale In-situ Workflows

While in-situ workflow formulations have addressed some of the data-related challenges associated with extreme-scale scientific workflows, these workflows involve complex interactions and different modes of data exchange. In the context of increasing system complexity, such workflows present significant resource management challenges, requiring complex cost-performance tradeoffs. This paper presents RISE, an intelligent staging-based data management middleware, which builds on the DataSpaces framework and performs intelligent scheduling of data management operations to reduce I/O contention. In RISE, data are always written immediately to local buffers to reduce the effect of the transfer impact upon application performance. RISE identifies applications’ data access patterns and moves data towards data consumers only when the network is expected to be idle, reducing the impact of asynchronous background data movement upon critical data read/write requests. Here, we experimentally demonstrate that RISE can take advantage of staging nodes to offload data during writes without degrading application data movement performance.

97 MATHEMATICS AND COMPUTING↗

Parked aeroelastic field rotor response for a 20% scaled demonstrator of a 13‐MW downwind turbine

Abstract Aeroelastic parked testing of a unique downwind two‐bladed subscale rotor was completed to characterize the response of an extreme‐scale 13‐MW turbine in high‐wind parked conditions. A 20% geometric scaling was used resulting in scaled 20‐m‐long blades, whose structural and stiffness properties were designed using aeroelastic scaling to replicate the nondimensional structural aeroelastic deflections and dynamics that would occur for a lightweight, downwind 13‐MW rotor. The subscale rotor was mounted and field tested on the two‐bladed Controls Advanced Research Turbine (CART2) at the National Renewable Energy Laboratory's Flatiron Campus (NREL FC). The parked testing of these highly flexible blades included both pitch‐to‐run and pitch‐to‐feather configurations with the blades in the horizontal braked orientation. The collected experimental data includes the unsteady flapwise root bending moments and tip deflections as a function of inflow wind conditions. The bending moments are based on strain gauges located in the root section, whereas the tip deflections are captured by a video camera on the hub of the turbine pointed toward the tip of the blade. The experimental results are compared against computational predictions generated by FAST, a wind turbine simulation software, for the subscale and full‐scale models with consistent unsteady wind fields. FAST reasonably predicted the bending moments and deflections of the experimental data in terms of both the mean and standard deviations. These results demonstrate the efficacy of the first such aeroelastically scaled turbine test and demonstrate that a highly flexible lightweight downwind coned rotor can be designed to withstand extreme loads in parked conditions.

17 WIND ENERGY↗

Constructing a new predictive scaling formula for ITER's divertor heat-load width informed by a simulation-anchored machine learning

Understanding and predicting divertor heat-load width λq is a critically important problem for an easier and more robust operation of ITER with high fusion gain. Previous predictive simulation data for λ q using the extreme-scale edge gyrokinetic code XGC1 [S. Ku et al., Phys. Plasmas 25, 056107 (2018)] in the electrostatic limit under attached divertor plasma conditions in three major US tokamaks [C. S. Chang et al., Nucl. Fusion 57, 116023 (2017)] reproduced the Eich and Goldston attached-divertor formula results [formula #14 in T. Eich et al., Nucl. Fusion 53, 093031 (2013) and R. J. Goldston, Nucl. Fusion 52, 013009 (2012)] and furthermore predicted over six times wider λ q than the maximal Eich and Goldston formula predictions on a full-power (Q = 10) scenario ITER plasma. After adding data from further predictive simulations on a highest current JET and highest-current Alcator C-Mod, a machine learning program is used to identify a new scaling formula for λ q as a simple modification to the Eich formula #14, which reproduces the Eich scaling formula for the present tokamaks and which embraces the wide λ q XGC for the full-current Q = 10 ITER plasma. Additionally, the new formula is then successfully tested on three more ITER plasmas: two corresponding to long burning scenarios with Q = 5 and one at low plasma current to be explored in the initial phases of ITER operation. The new physics that gives rise to the wider λ q XGC is identified to be the weakly collisional, trapped-electron-mode turbulence across the magnetic separatrix, which is known to be an efficient transporter of the electron heat and mass. Electromagnetic turbulence and high-collisionality effects on the new formula are the next study topics for XGC1.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Running Ensemble Workflows at Extreme Scale: Lessons Learned and Path Forward

The ever-increasing volumes of scientific data combined with sophisticated techniques for extracting information from them have led to the increasing popularity of ensemble workflows which are a collection of runs of individual workflows. A traditional approach followed by scientists to run ensembles is to rely on simple scripts to execute different runs and manage resources. This approach is not scalable and is error-prone, thereby motivating the development of workflow management systems that specialize in executing ensembles on HPC clusters. However, when the size of both the ensemble and the target system reach extreme scales, existing workflow management systems face new challenges that hamper their efficient execution. In this paper, we describe our experience scaling an ensemble workflow from the computational biology domain from the early design stages to the execution at extreme scale on Summit, a leadership class supercomputer at the Oak Ridge National Laboratory. We discuss challenges that arise when scaling ensembles to several million runs on thousands of HPC nodes. We identify challenges with composition of the ensemble itself, its execution at large scale, post-processing of the generated data, and scalability of the file system. Based on the experience acquired, we develop a generic vision of the capabilities and abstractions to add to existing workflow management systems to enable the execution of ensemble workflows at extreme scales. We believe that the understanding of these fundamental challenges will help application teams along with workflow system developers with designing the next generation of infrastructure for composing and executing extreme-scale ensemble workflows.

Mehta, Kshitij↗

Program Verification for Extreme-Scale Applications (Final Scientific/Technical Report)

The project seeks to develop tools and techniques to help software developers verify the correctness and accuracy of their code. The target domain is scientific software of the kind widely used and developed in the Department of Energy research community, with a particular focus on "extreme scale" programs - those that are expected to involve possibly millions of parallel threads of execution. The report covers the University of Delaware contribution to the collaborative project. The project had a number of successful outcomes, especially regarding the development and extension of the CIVL software verification framework. CIVL is a verification tool for C or Fortran programs that use MPI, OpenMP, CUDA, and/or Pthreads for parallelization. CIVL went through vast improvements and extensions, and was successfully applied to a number of challenging codes. It found a subtle bug in the Devito PDE framework. It was able to verify the functional correctness of a conjugate gradient solver using a novel probabilistic technique with vanishingly small chance of error. CIVL was used very successfully in verification competitions, and the CIVL solutions to the competition challenges were presented and published. CIVL was also used successfully by other researchers on a computational chemistry kernel.

97 MATHEMATICS AND COMPUTING↗

Extending the Publish/Subscribe Abstraction for High-Performance I/O and Data Management at Extreme Scale

The Adaptable I/O System (ADIOS) represents the culmination of substantial investment in Scientific Data Management, and it has demonstrated success for several important extreme-scale science cases. However, looking towards the exascale and beyond, we see the development of yet more stringent data management requirements that require new abstractions. Therefore, there is an opportunity to attempt to connect the traditional realms of HPC I/O optimization with the Database / Data Management community. As such, in this paper we offer some specific examples from our ongoing work in managing data structures, services, and performance at the extreme scale for scientific computing. Using the publish/subscribe model afforded by ADIOS, we demonstrate a set of services that connect data format, metadata, queries, data reduction, and high-performance delivery. The resulting publish/subscribe framework facilitates connection to on-line workflow systems to enable the dynamic capabilities that will be required for exascale science.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗