Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “concurrent computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Synchronous and Concurrent Multidomain Computing Method for Cloud Computing Platforms

We present a numerical method for synchronous and concurrent solution of transient elastodynamics problem where the computational domain is divided into subdomains that may reside on separate computational platforms. Here, this work employs the variational multiscale discontinuous Galerkin (VMDG) method to develop interdomain transmission conditions for transient problems. The fine-scale modeling concept leads to variationally consistent coupling terms at the common interfaces. The method admits a large class of time discretization schemes, and decoupling of the solution for each subdomain is achieved by selecting any explicit algorithm. Numerical tests with a manufactured solution problem show optimal convergence rates. The energy history in a free vibration problem is in agreement with that of the solution from a monolithic computational domain.

97 MATHEMATICS AND COMPUTING↗

SYMBIOSYS: A Methodology for Performance Analysis of Composable HPC Data Services

Microservices are a powerful new way of building, customizing, and deploying distributed services owing to their flexibility and maintainability. Several large-scale distributed platforms have emerged to serve the growing needs of data-centric workloads and services in commercial computing. Concurrently, high-performance computing (HPC) systems and software are rapidly evolving to meet the demands of diversified applications and heterogeneity. The interplay of hardware factors, software configuration parameters, and the flexibility offered with a microservice architecture makes it nontrivial to estimate the optimal service instantiation for a given application workload. Further, this problem is exacerbated when considering that these services operate in a dynamic and heterogeneous HPC environment. An optimally integrated service can be vastly more performant than a haphazardly integrated one. Existing performance tools for HPC either fail to understand the request-response model of communication inherent to microservices or they operate within a narrow scope, limiting the insight that can be gleaned from employing them in isolation. We propose a methodology for integrated performance analysis of HPC microservices frameworks and applications called SYMBIOSYS. We describe its design and implementation within the context of the Mochi framework. This integration is achieved by combining distributed callpath profiling and tracing with a performance data exchange strategy that collects fine-grained, low-level metrics from the RPC communication library and network layers. The result is a portable, low-overhead performance analysis setup that provides a holistic profile of the dependencies among microservices and how they interact with the Mochi RPC software stack. Using HEPnOS, a production-quality Mochi data service, we demonstrate the low-overhead operation of SYMBIOSYS at scale and use it to identify the root causes of poorly performing service configurations.

microservices↗

PythonFOAM: In-situ data analyses with OpenFOAM and Python

Here, we outline the development of a general-purpose Python-based data analysis tool for OpenFOAM. Our implementation relies on the construction of OpenFOAM applications that have bindings to data analysis libraries in Python. Double precision data in OpenFOAM is cast to a NumPy array using the NumPy C-API and Python modules may then be used for arbitrary data analysis and manipulation on flow-field information. We highlight how the proposed wrapper may be used for an in-situ online singular value decomposition (SVD) implemented in Python and accessed from the OpenFOAM solver PimpleFOAM. Here, 'in-situ' refers to a programming paradigm that allows for a concurrent computation of the data analysis on the same computational resources utilized for the partial differential equation solver. In addition, to demonstrate parallel deployments, we deploy a distributed SVD, which collects snapshot data across the ranks of a distributed simulation to compute the global left singular vectors. Crucially, both OpenFOAM and Python share the same message passing interface (MPI) communicator for this deployment which allows Python objects and functions to exchange NumPy arrays across ranks. Subsequently, we provide scaling assessments of this distributed SVD on multiple nodes of Intel Broadwell and KNL architectures for canonical test cases such as the large eddy simulations of a backward facing step and a channel flow at friction Reynolds number of 395. Finally, we demonstrate the deployment of a deep neural network for compressing the flow-field information using an autoencoder to demonstrate an ability to use state-of-the-art machine learning tools in the Python ecosystem.

97 MATHEMATICS AND COMPUTING↗

NekRS, a GPU-accelerated spectral element Navier–Stokes solver

The development of NekRS, a GPU-oriented thermal-fluids simulation code based on the spectral element method (SEM) is described. For performance portability, the code is based on the open concurrent compute abstraction and leverages scalable developments in the SEM code Nek5000 and in libParanumal, which is a library of high-performance kernels for high-order discretizations and PDE-based miniapps. Critical performance sections of the Navier–Stokes time advancement are addressed. Performance results on several platforms are presented here, including scaling to 27,648 V100s on OLCF Summit, for calculations of up to 60B gridpoints.

97 MATHEMATICS AND COMPUTING↗

Structural Simluation Toolkit (SST) v.12.0

The Structural Simulation Toolkit (SST) was developed to explore innovations in highly concurrent computing systems where the instruction set architecture (ISA), micro-architecture, and memory interact with the programming model and communications system. The package provides a fully modular design for extensive exploration of an individual system parameter as well as a parallel simulation environment based on message passing interface (MPI) which enable a high level of performance as well as the ability to look at large systems. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Rodrigues, ArunF.↗

Enabling Scalable and Extensible Memory-mapped Datastores in Userspace

Exascale workloads are expected to incorporate data-intensive processing in close coordination with traditional physics simulations. These emerging scientific, data-analytics and machine learning applications need to access a wide variety of datastores in flat files and structured databases. Programmer productivity is greatly enhanced by mapping datastores into the application process's virtual memory space to provide a unified “in-memory” interface. Currently, memory mapping is provided by system software primarily designed for generality and reliability. However, scalability at high concurrency is a formidable challenge on exascale systems. Also, there is a need for extensibility to support new datastores potentially requiring HPC data transfer services. In this article, we present UMap , a scalable and extensible userspace service for memory-mapping datastores. Furthermore, through decoupled queue management, concurrency aware adaptation, and dynamic load balancing, UMap enables application performance to scale even at high concurrency. We evaluate UMap in data-intensive applications, including sorting, graph traversal, database operations, and metagenomic analytics. Our results show that UMap as a userspace service outperforms an optimized kernel-based service across a wide range of intra-node concurrency by 1.22-1.9 × . We performed two case studies to demonstrate UMap 's extensibility. First, a new datastore residing in remote memory is incorporated into UMap as an application-specific plugin. Second, we present a persistent memory allocator Metall built atop UMap for unified storage/memory.

97 MATHEMATICS AND COMPUTING↗

GeantV: Results from the Prototype of Concurrent Vector Particle Transport Simulation in HEP

Full detector simulation was among the largest CPU consumers in all CERN experiment software stacks for the first two runs of the Large Hadron Collider. In the early 2010s, it was projected that simulation demands would scale linearly with increasing luminosity, with only partial compensation from increasing computing resources. The extension of fast simulation approaches to cover more use cases that represent a larger fraction of the simulation budget is only part of the solution, because of intrinsic precision limitations. The remainder corresponds to speeding up the simulation software by several factors, which is not achievable by just applying simple optimizations to the current code base. In this context, the GeantV R&D project was launched, aiming to redesign the legacy particle transport code in order to benefit from features of fine-grained parallelism, including vectorization and increased locality of both instruction and data. This paper provides an extensive presentation of the results and achievements of this R&D project, as well as the conclusions and lessons learned from the beta version prototype.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

IRIS-MEMFLOW: Data Flow-Enabled Portable Memory Orchestration in IRIS Runtime for Diverse Heterogeneity

Task-based programming models and execution paradigms provide a means to decompose a computation by expressing it as a graph in which each node represents a specific computation operating on memory objects and the edges define the dependencies in the execution flow. In this execution model, independent nodes in the graph can be executed concurrently in different computing devices, making it suitable for heterogeneous systems in which computing devices with different architectures coexist. However, careful memory orchestration across heterogeneous devices is needed because copies of the same memory object may reside in multiple devices during execution. Manually ensuring such an orchestration is quite challenging. Not only must an application developer guard against race conditions, but they must also optimize data movement between the host and devices because unnecessary data movement significantly impacts performance. To mitigate these challenges, we enhance the IRIS heterogeneous runtime and introduce IRIS-MEMFLOW–a data flow–enabled portable memory abstraction for seamlessly orchestrating memory in diverse heterogeneous computing environments. By using data-flow analysis, IRIS-MEMFLOW guards against race conditions while multiple heterogeneous devices access memory objects. IRIS-MEMFLOW also optimizes data movement between the host and devices without manual intervention. As a result, IRIS provides improved programming productivity, performance, and portability for multidevice heterogeneous executions in high-performance computing and cloud systems that run diverse architectures from different vendors. The efficacy of IRIS-MEMFLOW is evaluated through experiments that show its capability in terms of programming productivity, multidevice heterogeneity, portability, and low overhead versus the state of the art.

Monil, M. A. H. [ORNL] (ORCID:0000000334194037)↗

Strain Fluctuations Unlock Ferroelectricity in Wurtzites

Ferroelectrics are of practical interest for non-volatile data storage due to their reorientable, crystallographically defined polarization. Yet efforts to integrate conventional ferroelectrics into ultrathin memories have been frustrated by film-thickness limitations, which impede polarization reversal under low applied voltage. Wurtzite materials, including magnesium-substituted zinc oxide (Zn,Mg)O, have been shown to exhibit scalable ferroelectricity as thin films. In this work, the origins of ferroelectricity in (Zn,Mg)O are explained, showing that large strain fluctuations emerge locally in (Zn,Mg)O and can reduce local barriers to ferroelectric switching by more than 40%. Concurrent experimental and computational evidence of these effects are provided by demonstrating polarization switching in ZnO/(Zn,Mg)O/ZnO heterostructures featuring built-in interfacial strain gradients. These results open up an avenue to develop scalable ferroelectrics by controlling strain fluctuations atomistically.

36 MATERIALS SCIENCE↗

Three Birds with One Stone: Improving Performance, Convergence, and System Throughput with NEST

Variational quantum algorithms (VQAs) have the potential to demonstrate quantum utility on near-term quantum computers. However, these algorithms often get executed on the highest-fidelity qubits and computers to achieve the best performance, causing low system throughput. Recent efforts have shown that VQAs can be run on low-fidelity qubits initially and high-fidelity qubits later on to still achieve good performance. We take this effort forward and show that carefully varying the qubit fidelity map of the VQA over its execution using our technique, Nest, does not just (1) improve performance (i.e., help achieve close to optimal results), but also (2) lead to faster convergence. We also use Nest to co-locate multiple VQAs concurrently on the same computer, thus (3) increasing the system throughput, and therefore, balancing and optimizing three conflicting metrics simultaneously.

qaoa↗

A Computational Framework for Control Co-Design of Resilient Cyber–Physical Systems With Applications to Microgrids

Critical infrastructure networks, such as power and transportation networks, can be modelled as cyber-physical systems. As the complexity of such systems grow, there is need for developing metrics and design tools that will co-optimize the physical components of the system and the control policies to guarantee resilience against cyber and natural threats. To that end, we develop a simulation-based control co-design computational framework that will concurrently determine the system and control parameters of a cyber-physical system to meet pre-specified resilience, operational and economic objectives. Here, the capabilities of the developed co-design engine is demonstrated by designing the physical components and control parameters of a microgrid system that will meet its resiliency objectives when subjected to various cyber and physical threats.

97 MATHEMATICS AND COMPUTING↗

Scalable Control Co-design for Resilient-by-Design Cyber Physical Systems

Critical infrastructure networks, such as power and transportation networks, are often modelled as cyber-physical systems. With ever increasing complexity of these systems, there is a need for newer and more relevant metrics and design tools that will co-optimize the physical system components and control policies to guarantee resilience against cyber and natural threats. To this end, a simulation-based control co-design computational framework that will concurrently determine the system and control parameters of a cyber-physical system to meet pre-specified resilience, operational and economic objectives has been developed. The capabilities of the developed co-design engine are demonstrated by designing the physical components and control parameters of a microgrid system that will meet its resiliency objectives when subjected to various cyber and physical threats.

42 ENGINEERING↗

Performance Analysis of Traditional and Data-Parallel Primitive Implementations of Visualization and Analysis Kernels

Measurements of absolute runtime are useful as a summary of performance when studying parallel visualization and analysis methods on computational platforms of increasing concurrency and complexity. We can obtain even more insights by measuring and examining more detailed measures from hardware performance counters, such as the number of instructions executed by an algorithm implemented in a particular way, the amount of data moved to/from memory, memory hierarchy utilization levels via cache hit/miss ratios, and so forth. This work focuses on performance analysis on modern multi-core platforms of three different visualization and analysis kernels that are implemented in different ways: one is "traditional", using combinations of C++ and VTK, and the other uses a data-parallel approach using VTK-m. Our performance study consists of measurement and reporting of several different hardware performance counters on two different multi-core CPU platforms. The results reveal interesting performance differences between these two different approaches for implementing these kernels, results that would not be apparent using runtime as the only metric.

97 MATHEMATICS AND COMPUTING↗

A Task Based Approach for Co-Scheduling Ensemble Workloads on Heterogeneous Nodes

Scientific workflows consist of multiple, connected applications, with data and results flowing from one to another in a pipeline. Traditionally, such workflows are executed in sequential order, storing intermediate data in storage disks. Co-scheduling application workflows concurrently on the same compute nodes would greatly reduce the cost of moving data to/from storage and allow real-time analysis of intermediate results. Nevertheless, most parallel programming runtimes do not allow seamless integration of various applications in a scientific workflow, in part due to the complexity of managing data and resources. The situation is even more complicated for heterogeneous systems. In this work we extend the Minos Computing Library (MCL) runtime to accelerate pipe-lined and parallel workloads where multiple applications are running in the same system. MCL’s asynchronous task library and runtime dynamically manages resources to allow co-scheduling of multiple processes sharing heterogeneous resources. In addition, we design a custom ex- tension of the Open Compute Language (OpenCL) to enable multiple processes to share device memory. We enable MCL to coordinate these shared buffers to allow for easy, fast data sharing between applications. Using malleable micro-benchmarks and two application workflows that combine scientific simulation and AI-based analysis, we show that our method outperforms traditional approaches.

Index Terms—Parallel systems, Scheduling and Task ↗

Concurrent Precipitation of Nb(C,N) and Metastable M 23 C 6 in Alloy 347H at 700°C and 750°C: Computer Simulations and Comparison to Experiment

Here, we present our results for the concurrent precipitation of metastable M 23 C 6 , Nb(C,N) secondary precipitates, and the Nb(C,N) primary crystals in 347H austenitic stainless steel. For precipitation modeling, we have accounted for the elastic contribution to interfacial energy, and for the Fe-spin-polarization for NbC/Fe and M 23 C 6 /Fe interfacial energy values: for NbC/Fe ~ 0.63 J/m 2 . For M 23 C 6 precipitates, an error function was used to describe the interfacial energy growth with particle size. In precipitation simulations, the average size of the primary Nb(C,N) particles remained ~ 1 μm at 700°C and ~ 0.3 μm at 750°C. The M 23 C 6 precipitates at 750°C dissolved after 120 h (our simulations) compared to 300 h (experiments). The Nb(C,N)/Fe interfacial energy was not affected by the nitrogen additions. With these modifications, reasonable agreement with the available experimental data was obtained, which allows using them in the development of the 2nd-phase particle-informed creep theory.

36 MATERIALS SCIENCE↗

Multidisciplinary concurrent optimization framework for multi-phase building design process

Modern day building design projects require multidisciplinary expertise from architects and engineers across various phases of the design (conceptual, preliminary, and detailed) and construction processes. The Architecture Engineering and Construction (AEC) community has recently shifted gears toward leveraging design optimization techniques to make well-informed decisions in the design of buildings. However, most of the building design optimization efforts are either multidisciplinary optimization confined to just a specific design phase (conceptual/preliminary/detailed) or single disciplinary optimization (structural/thermal/daylighting/energy) spanning across multiple phases. Complexity in changing the optimization setup as the design progresses through subsequent phases, interoperability issues between modeling and physics-based analysis tools used at later stages, and the lack of an appropriate level of design detail to get meaningful results from these sophisticated analysis tools are few challenges that limit multi-phase multidisciplinary design optimization (MDO) in the AEC field. Here this paper proposes a computational building design platform leveraging concurrent engineering techniques such as interactive problem structuring, simulation-based optimization using meta models for energy and daylighting (machine learning based) and tradespace visualization. The proposed multi-phase concurrent MDO framework is demonstrated by using it to design and optimize a sample office building for energy and daylighting objectives across multiple phases. Furthermore, limitations of the proposed framework and future avenues of research are listed.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Evaluation of Best Practices in Mitigating Startup Costs on Leadership-Class Supercomputers

Supercomputers at Department of Energy (DOE) National Laboratories face a widening range of workloads, from traditional modeling and simulation to Artificial Intelligence model training or complex multi-stage workflows, and beyond. At DOE Leadership Computing Facilities like the Oak Ridge Leadership Computing Facility (OLCF), these workloads demand concurrent access to large portions of the supercomputer’s resources. Launching a job across massive supercomputers is challenging from the start; the file system struggles with a large backlog of metadata requests as tens of thousands of processes read thousands of the same files, and the compute job cannot start until this is completed. There are multiple existing approaches to calm this metadata storm, ranging from vendor-developed tools like sbcast to National Laboratory-developed tools like Spindle and Copper. In this paper, we benchmark and discuss three common approaches to improving compute job launch latencies on Frontier: Slurm’s sbcast tool, Spindle, and Copper. We evaluate these tools by measuring the launch latencies of four workloads: OSU Microbenchmark’s osu_init, Pynamic, Python import mpi4py, and Python import torch. We provide discussion of the results, highlighting data that meet expectations and that do not meet expectations.

Hagerty, Nick [ORNL] (ORCID:0000000330014414)↗