Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “HPC Access”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

170 records · Page 10

Experiences Readying Applications for Exascale

The advent of Exascale computing invites an assessment of existing best practices for developing application readiness on the world's largest supercomputers. This work details observations from the last four years in preparing scientific applications to run on the Oak Ridge Leadership Computing Facility's (OLCF) Frontier system. This paper addresses a range of topics in software including programmability, tuning, and portability considerations that are key to moving applications from existing systems to future installations. A set of representative workloads provides case studies for general system and software testing. We evaluate the use of early access systems for development across several generations of hardware. Finally, we discuss how best practices were identified and disseminated to the community through a wide range of activities including user-guides and trainings. We conclude with recommendations for ensuring application readiness on future leadership computing systems.

exascale↗

Reusability First: Toward FAIR Workflows

The FAIR principles of open science (Findable, Accessible, Interoperable, and Reusable) have had transformative effects on modern large-scale computational science. In particular, they have encouraged more open access to and use of data, an important consideration as collaboration among teams of researchers accelerates and the use of workflows by those teams to solve problems increases. How best to apply the FAIR principles to workflows themselves, and software more generally, is not yet well understood. We argue that the software engineering concept of technical debt management provides a useful guide for application of those principles to workflows, and in particular that it implies reusability should be considered as ‘first among equals’. Moreover, our approach recognizes a continuum of reusability where we can make explicit and selectable the tradeoffs required in workflows for both their users and developers.To this end, we propose a new abstraction approach for reusable workflows, with demonstrations for both synthetic workloads and real-world computational biology workflows. Through application of novel systems and tools that are based on this abstraction, these experimental workflows are refactored to rightsize the granularity of workflow components to efficiently fill the gap between end-user simplicity and general customizability. Our work makes it easier to selectively reason about and automate the connections between trade-offs across user and developer concerns when exposing degrees of freedom for reuse. Additionally, by exposing fine-grained reusability abstractions we enable performance optimizations, as we demonstrate on both institutional-scale and leadership-class HPC resources.

Wolf, Matthew↗

Improved Distributed-memory Triangle Counting by Exploiting the Graph Structure

Graphs are ubiquitous in modeling complex systems and representing interactions between entities to uncover structural information of the domain. Traditionally, graph analytics workloads are challenging to efficiently scale (both strong and weak cases) on distributed memory due to the irregular memory-access driven nature (with little or no computations) of the methods. The structure of graphs and their relative distribution over the processing elements poses another level of complexity, making it difficult to attain sustainable scalability across platforms. In this paper, we discuss enhancements to TriC, a distributed-memory implementation of graph triangle counting using Message Passing Interface (MPI), which was featured in the 2020 Graph Challenge competition. We have made some incremental enhancements to TriC, primarily adopting a user-defined buffering strategy to overcome the startup problem for large graphs (by fixing the memory for intermediate data), and experimenting with probabilistic data structures such as bloom filter to improve the query response time for assessing edge existence, at the expense of increasing the overall false positive rate. These adjustments have led to a modest improvements in most cases, as compared to the previous version.

Graph Analytics, HPC↗

Pavilion 2 Feature Additions [Slides]

Pavilion 2 is a critical component of system testing for LANL HPC systems. Pavilion 2 is system independent, which allows for the creation of a large suite of generalized tests than can be applied to multiple systems. These additions to Pavilion 2 improve our ability to test systems and diagnose issues. With such a large suite, running the proper tests when diagnosing and fixing a system can be difficult, as well as locating old tests for reference. Test configurations may contain several permutations of a test, not all of which may apply to a given system. The ability to run sets of test permutations lets the tester focus on specific parts of the test without having to edit the test configurations or run extraneous tests. The addition of the filter argument improves the ability to find test results based on test attributes, including past tests. Pavilion 2's system independence is largely based on various layers of configuration files. Configuration files for the host, test, and modes exist to allow generalized testing. The operating system configuration layer can provide helpful OS defaults to a test. Survey, a collection and reporting tool, can help us gauge the performance of a system as well as diagnose issues. The addition of a Survey mode in Pavilion 2 makes running tests with Survey very straightforward. The Survey results are combined with Pavilion 2's test results to make it easy to access.

97 MATHEMATICS AND COMPUTING↗

Performance and Portability of a Linear Solver Across Emerging Architectures

A linear solver algorithm used by a large-scale unstructured-grid computational fluid dynamics application is examined for a broad range of familiar and emerging architectures. Efficient implementation of a linear solver is challenging on recent CPUs offering vector architectures. Vector loads and stores are essential to effectively utilize available memory bandwidth on CPUs, and maintaining performance across different CPUs can be difficult in the face of varying vector lengths offered by each. A similar challenge occurs on GPU architectures, where it is essential to have coalesced memory accesses to utilize memory bandwidth effectively. In this work, we demonstrate that restructuring a computation, and possibly data layout, with regard to architecture is essential to achieve optimal performance by establishing a performance benchmark for each target architecture in a low level language such as vector intrinsics or CUDA. In doing so, we demonstrate how a linear solver kernel can be mapped to Intel® Xeon™ and Xeon Phi™, Marvell® ThunderX2®, NEC® SX-Aurora™ TSUBASA Vector Engine, and NVIDIA® and AMD® GPUs. We further demonstrate that the required code restructuring can be achieved in higher level programming environments such as OpenACC, OCCA, and Intel® OneAPI™/SYCL, and that each generally results in optimal performance on the target architecture. Relative performance metrics for all implementations are shown, and subjective ratings for ease of implementation and optimization are suggested.

Programming models↗

NASA Earth Exchange – Current overview of climate and wildfire-oriented works

NASA Earth Exchange (NEX) combines state-of-the-art supercomputing, Earth system modeling, and NASA remote sensing data feeds to deliver a work environment for exploring and analyzing petabyte-scale datasets covering large regions, continents, or the globe. As an accessible platform, NEX can accelerate fundamental research, develop new applications, and reduce overall project costs by providing the research community with data, software, and high-end computing power. Two research thrusts for the NEX community are developing and distributing the NEX-GDDP-CMIP6 downscaled climate dataset and the community-driven wildfire research. NEX-GDDP-CMIP6 dataset is comprised of global downscaled climate scenarios derived from the General Circulation Model (GCM) runs conducted under the Coupled Model Intercomparison Project Phase 6 (CMIP6) and across two of the four “Tier 1” greenhouse gas emissions scenarios known as Shared Socioeconomic Pathways (SSPs). The wildfire research thrust leverages geostationary and low earth orbit remote sensing platforms and advanced modeling (WRF). We present a high-level overview of the NEX community, discuss relevant datasets and show several visualization techniques of data used in the wildfire research activity.

NEX↗

MOSIQS: Persistent Memory Object Storage With Metadata Indexing and Querying for Scientific Computing

Scientific applications often require high-bandwidth shared storage to perform joint simulations and collaborative data analytics. Shared memory pools provide a chance to satisfy such needs. Recently, a high-speed network such as Gen-Z utilizing persistent memory (PM) offers an opportunity to create a shared memory pool connected to compute nodes. However, there are several challenges to use scientific applications on the shared memory pool directly such as scalability, failure-atomicity, and lack of scientific metadata-based search and query. In this paper, we propose MOSIQS, a persistent memory object storage framework with metadata indexing and querying for scientific computing. We design MOSIQS based on the key idea that memory objects on PM pool can live beyond the application lifetime and can become the sharing currency for applications and scientists. MOSIQS provides an aggregate memory pool atop an array of persistent memory devices to store and access memory objects to accelerate scientific computing. MOSIQS uses a lightweight persistent memory key-value store to manage the metadata of memory objects, which enables memory object sharing. To facilitate metadata search and query over millions of memory objects resident on memory pool, we introduce Group Split and Merge (GSM), a novel persistent index data structure designed primarily for scientific datasets. GSM splits and merges dynamically to minimize the query search space and maintains low query processing time while overcoming the index storage overhead. MOSIQS is implemented on top of PMDK. We evaluate the proposed approach on many-core server with an array of real PM devices. Experimental results show that MOSIQS gains a 100% write performance improvement and executes multi-attribute queries efficiently with 2.7× less index storage overhead offering significant potential to speed up scientific computing applications.

97 MATHEMATICS AND COMPUTING↗

Cloud Computing Methods for Near Rectilinear Halo Orbit Trajectory Design

Complicated mission design problems require innovative computational solutions. As spacecraft depart from a proposed Gateway in a Near Rectilinear Halo Orbit (NRHO), recontact analysis is required to avoid risk of collision and ensure safe operations. Escape dynamics from NRHOs are governed by multiple gravitational bodies, yielding a trajectory design space that is exhaustively large. This paper summarizes the recontact analysis for departure from the NRHO and describes how the Deep Space Trajectory Explorer (DSTE) trajectory design software incorporates high performance cloud computing to compute and visualize the orbit design space. Recent focus on exploration missions to cislunar space has kindled accelerated interest in multibody orbit solutions. Trajectory analysis in the presence of multiple gravity fields is complex, and innovative computational tools are needed to simplify complicated design spaces, to generate large quantities of data quickly, and to visualize the output for user accessibility. The Gateway mission is a prime example. The Gateway1 is proposed as a human outpost in deep space. The current baseline orbit for the Gateway is a Near Rectilinear Halo Orbit (NRHO) near the Moon.2 The NRHO exists in a regime that experiences the gravitational effects of the Earth and the Moon simultaneously, complicating orbit analysis. The mission design process benefits greatly from updated computational tools for multibody missions like the Gateway. As an example, consider the problem of assessing the risk of collision in an NRHO. As a staging location to missions to the lunar surface and beyond the Earth-Moon system, the Gateway will experience spacecraft and other objects regularly arriving and departing. Departing objects potentially include spent logistics modules, visiting crew vehicles, debris objects, wastewater particles, and cubesats. Each departure is governed by the dynamics of the Gateway orbit and the surrounding dynamical environment. Over time, any unmaintained object in such an orbit eventually departs due to the small instabilities associated with the NRHOs. A separation maneuver speeds the departure from the NRHO, but the effects of the maneuver on the spacecraft behavior depend on the location, magnitude, and direction of the burn. Escape dynamics from the NRHO with regard to these maneuver options open up an enormous potential trajectory design space where subtle changes in input can produce dramatically large changes in the results. Any departing object must avoid recontacting the Gateway as it leaves the lunar vicinity, and a recontact analysis thus involves a significant number of computations and extensive output data. To explore the dynamics of this extensive design space, the Deep Space Trajectory Explorer3 (DSTE) trajectory design software incorporates new High Performance Computing (HPC) services and novel interactive visualizations. This paper details the HPC and cloud infrastructure techniques that are implemented in the DSTE, applying the new capabilities to analysis of recontact risk with the Gateway in NRHO. NEAR RECTILINEAR HALO ORBITS The Gateway is planned to fly in a lunar NRHO as its baseline orbit. The NRHO families of orbits are subsets of the larger halo families, which originate from planar orbits near the L1 and L2 libration points; the Earth-Moon L2 halo family appears in Figure 1. Each halo orbit is perfectly periodic in the Circular Restricted 3-Body Problem (CR3BP) and becomes a quasi-periodic orbit in a higher fidelity ephemeris force model. The NRHOs are defined as those members of the halo family with bounded stability properties;2 they pass near the Moon at perilune and are nearly polar. Families exist with apolunes located both above the lunar north pole and above the lunar south pole; the Gateway is planned to reside in a southern L2 NRHO in a 9:2 resonance with the lunar synodic period. The 9:2 NRHO is characterized by a period of about 6.5 days, a perilune radius of about 3,500 km, and an apolune radius of about 71,000 km; it is strongly affected by the gravity of both the Earth and the Moon simultaneously. This NRHO offers extended communications with assets on the south pole of the Moon,4 as well as low-cost orbit maintenance and attitude control,5 favorable eclipse avoidance properties,6 and inexpensive transfers from Earth and to other destinations.5,7 The NRHO portion of the southern L2 halo family is highlighted in black in Figure 1, and the 9:2 NRHO appears in blue.

Phillips, Sean M.↗