Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “asynchronous”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23

¬¬Integration of Quantification of Margins and Uncertainties Methodology into Parallel Discrete Event Simulator Framework

Parallel Discrete Events Simulation (PDES) is becoming increasingly important to lab efforts in security and intelligence. It is used to model complex asynchronous systems such as computer networks, satellite systems, vehicular traffic, and human performance. Uncertainty Quantification (UQ) techniques have become a mainstay of Verification and Validation (V&V) efforts on physics simulations. However, PDES models are very different from traditional physics simulations, and research into UQ techniques for PDES is in its infancy. It is not clear which traditional UQ techniques can be applied to PDES, or what new techniques will need to be developed. The goal of this project was to identify existing UQ techniques that can be applied to PDES, and to develop new techniques as necessary. The project implemented or developed techniques to handle issues that do not appear in traditional physics simulations, but are common among PDES, including sampling techniques for high-dimensional homogenous inputs, response surfaces for high-variance heteroskedastic output, and characterization of skewed output distributions. These are foundational UQ techniques that must be used in any complete UQ analysis (Tong 2018).

97 MATHEMATICS AND COMPUTING↗

UPC++ v1.0 Specification, Revision 2020.10.0

UPC++ is a C++11 library providing classes and functions that support Partitioned Global Address Space (PGAS) programming. The key communication facilities in UPC++ are one-sided Remote Memory Access (RMA) and Remote Procedure Call (RPC). All communication operations are syntactically explicit and default to non-blocking; asynchrony is managed through the use of futures, promises and continuation callbacks, enabling the programmer to construct a graph of operations to execute asynchronously as high-latency dependencies are satisfied. A global pointer abstraction provides system-wide addressability of shared memory, including host and accelerator memories. The parallelism model is primarily process-based, but the interface is thread-safe and designed to allow efficient and expressive use in multi-threaded applications. The interface is designed for extreme scalability throughout, and deliberately avoids design features that could inhibit scalability.

97 MATHEMATICS AND COMPUTING↗

UPC++ v1.0 Programmer’s Guide, Revision 2020.10.0

UPC++ is a C++11 library that provides Partitioned Global Address Space (PGAS) programming. It is designed for writing parallel programs that run efficiently and scale well on distributed-memory parallel computers. The PGAS model is single program, multiple-data (SPMD), with each separate constituent process having access to local memory as it would in C++. However, PGAS also provides access to a global address space, which is allocated in shared segments that are distributed over the processes. UPC++ provides numerous methods for accessing and using global memory. In UPC++, all operations that access remote memory are explicit, which encourages programmers to be aware of the cost of communication and data movement. Moreover, all remote-memory access operations are by default asynchronous, to enable programmers to write code that scales well even on hundreds of thousands of cores.

97 MATHEMATICS AND COMPUTING↗

Hardening DOE R&D Software Tools for Web-based Visualization SBIR Phase I Final Report

Ubiquitous web-based visualization is essential to delivering large-scale data visualization to various stakeholders, from the scientist to the board member. These stakeholders will not tolerate a stalled application or a pop-up window asking them to wait for the processing to complete. They require a responsive and interactive visualization environment with high-quality imagery suitable for detailed analysis and boardroom presentations. At Kitware, Inc., we have accomplished web visualization to this point, leveraging state-of-the-art tools like HTML5, CSS3, SVG, Canvas, and WebGL. Solutions that leverage a combination of these technologies are necessary to handle workloads that vary significantly in data size efficiently. However, it is not always practical to move large data to the web client for visualization. Kitware's ParaView as a Service combines client-side visualization using both distributed processing and remote rendering on big data impractical to move. Existing distributed processing and remote rendering solution's interactivity is below the expectations of web-based applications. Our project examined proposed solutions to the areas outlined above in ParaView as a Service. We have investigated concurrent pipelines, streaming images, progressive rendering, and optimization of algorithms and data movement to address these concerns. For the Phase I project, we completed the proposed work plan. As a result, the project produced three prototypes of essential importance for web visualization and the ParaView as a Service community. We created a simple desktop application for an interactive streamline placement prototype, a web-based interactive streamline placement prototype, and a web-based progressive rendering utilizing raytracing prototype. These prototypes relied on the hardening of emerging software toolkits funded by the Department of Energy (DOE) Advanced Scientific Computing Research (ASCR) program (such as ParaView, VTK-m, and Mochi). We blended these components into web-based visualization prototypes that meet the industry's expectations for interactivity and responsiveness. The Phase I project had four essential focus areas: 1. Develop prototype ParaView as a Service backend server using asynchronous, non-blocking design principles. 2. Develop a prototype web application that uses the ParaView as a Service backend server for remote data visualization. 3. Implement image streaming with encoding/compression and progressive rendering capabilities in the proposed platform. 4. Evaluate the prototype developed and summarize observations, including the challenges and pitfalls of our approach. After our successful completion of Phase I, we are strongly positioned to propose a successful Phase II project.

Geveci, Berk↗

CodeFlow: A Code Generation System for Flash-X Orchestration Runtime

We propose the CodeFlow toolchain for Flash-X that realizes the “recipe-to-source” code transformation for Flash-X simulations and that is necessary to achive performance portability. We design a high-level language to express operations of simulations in so-called recipes, which are given as input to the toolchain. The tools of the CodeFlow pipeline include code transformation with tree-based source code representation techniques and code orchestration and generation based on control flow graphs. The generated source code utilizes a new runtime, developed for Flash-X, that orchestrates dynamic and asynchronous data movement and task execution. The functionality of CodeFlow is demonstrated using a hydrodynamic problem with a strong shock.

97 MATHEMATICS AND COMPUTING↗

Domaine-Specific Runtime to Orchestrate Computation on Heterogeneous Platforms

Task-based runtime systems in the past have sought to exploit inherent asynchronicity in the application execution to reduce overall runtime. In the last decade focus shifted to supporting the heterogeneity that is increasingly prevalent in high-performance computing systems. However, existing task-based runtime systems, being general, come with a challenging set of issues such as the complexity of abstractions and overheads. And they still leave much of the burden of exposing the parallelism on the application developers who also have to fit their applications to the runtime systems’ interfaces. In this paper we take a different approach for applications to target heterogeneous systems through domain-specific run-times. Our design aspires to leverage the domain-specific knowledge of a focused class of scientific simulations to pragmatically orchestrate computations in the simulations.

97 MATHEMATICS AND COMPUTING↗

Optimizing and Extending the Functionality of EXARL for Scalable Reinforcement Learning [Slides]

The main goal of the Co-Design Summer School 2021 is to provide algorithmic improvements to EXARL framework by improving performance and by adding functionalities. This presentation includes an introduction to reinforcement learning and to EXARL. The researchers expanded the capability of EXARL by including additional agents like (Asynchronized) Advantage Actor Critic (A2C/A3C) and Twin Delayed Deep Deterministic Policy Gradient (TD3). They also explored algorithmic improvements such as v-trace and Prioritized Experience Replay. They found that A2C/A3C performed best with v-trace and outperformed Deep Q-Network (DQN) on both the CartPole game and the ExaBooster scientific environment. Additionally, they found that TD3 performed as good as the existing Deep Deterministic Policy Gradient (DDPG) agent and that adding Prioritized Experience Replay to DDPG accelerated convergence.

97 MATHEMATICS AND COMPUTING↗

Design and Performance of Kokkos Staging Space toward Scalable Resilient Application Couplings

With the growing number of applications designed for heterogeneous HPC devices, application programmers and users are finding it challenging to compose scalable workflows as ensembles of these applications, that are portable, performant and resilient. The Kokkos C++ library has been designed to simplify this cumbersome procedure by providing an intra-application uniform programming model and portable performance. However, assembling multiple Kokkos-enabled applications into a complex workflow is still a challenge. Although Kokkos enables a uniform programming model, the inter-application data exchange still remains a challenge from both performance and software development cost perspectives. In order to address this issue, we propose Kokkos data staging memory space, an extension of Kokkos' data abstraction (memory space) for heterogeneous computing systems. This new abstraction allows to express data on a virtual shared-space for multiple Kokkos applications, thus extending Kokkos to support inter-application data exchange to build an efficient application workflow. Additionally, we study the effectiveness of asynchronous data layout conversions for applications requiring different memory access patterns for the shared data. Our preliminary evaluation with a synthetic benchmark indicate the effectiveness of this conversion adapted to three different scenarios representing access frequency and use patterns of the shared data.

97 MATHEMATICS AND COMPUTING↗

Performance Improvements for the Griffin Transport Solvers

Griffin is a Multiphysics Object-Oriented Simulation Environment based reactor multiphysics analysis application jointly developed by Idaho National Laboratory and Argonne National Laboratory. Griffin includes a variety of deterministic radiation transport solvers for fixed source, k-eigenvalue, adjoint, and subcritical multiplication, as well as transient solvers for point-kinetics, improved quasi-static, and spatial dynamics. A code assessment performed in FY-20 identified two significant issues with the transport solvers in Griffin: first, the primary heterogeneous SN (discrete ordinates) transport solver based on continuous finite element methods required significant mesh refinement and higher memory usage compared to solvers based on the method of characteristic for equivalent accuracy. Second, the homogeneous PN (spherical harmonics expansion) transport solver did not adequately support polynomial refinement, which is a feature usually required for problems with spatial homogenization and pronounced streaming, typical in fast or gas-cooled reactor systems. To address the first issue, the development effort focused on the more promising discontinuous finite element method (DFEM)-based SN transport solver in Griffin. The addition of an asynchronous parallel transport sweeper and coarse mesh finite difference (CMFD) acceleration have rendered a superior heterogeneous SN transport capability for multiphysics problems that requires far less computing resources in terms of both CPU time and memory usage. This is demonstrated with typical thermal- and fast-spectrum reactor benchmark problems, including 2D Transient Reactor Test, 3D Advanced Burner Test Reactor (ABTR), and 2D and 3D Empire microreactor. For the second issue, the development effort focused on a new transport solver based on the hybrid finite element PN method (HFEM-PN), equivalent to the variational nodal method, as well as a new diffusion solver based on HFEM-Diffusion. This solver is intended for homogenized domains with multiphysics coupling (i.e., supports mesh displacement, seamless temperature feedback, etc.). Initial calculations with the HFEM-Diffusion implementation show very good parallel efficiency for the residual evaluations with the 2D ABTR benchmark. A future development effort will be centered on further improvements to the CMFD, HFEM-PN, and DFEM diffusion solvers to ensure Griffin meets performance and software quality assurance requirements for advanced reactor design and analysis.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

UPC++ v1.0 Programmer’s Guide, Revision 2021.9.0

UPC++ is a C++ library that provides Partitioned Global Address Space (PGAS) programming. It is designed for writing parallel programs that run efficiently and scale well on distributed-memory parallel computers. The PGAS model is single program, multiple-data (SPMD), with each separate constituent process having access to local memory as it would in C++. PGAS additionally provides one-sided Remote Memory Access (RMA) to a global address space, which is allocated in shared segments that are distributed over the processes. UPC++ also features Remote Procedure Call (RPC) communication, making it easy to move computation to operate on data that resides on remote processes. In UPC++, all communication operations are explicit, which encourages programmers to be aware of the cost of communication and data movement. Moreover, all communication operations are asynchronous by default, to enable programmers to write code that scales well even on hundreds of thousands of cores.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

UPC++ v1.0 Specification, Revision 2021.9.0

UPC++ is a C++ library providing classes and functions that support Partitioned Global Address Space (PGAS) programming. The key communication facilities in UPC++ are one-sided Remote Memory Access (RMA) and Remote Procedure Call (RPC). All communication operations are syntactically explicit and default to non-blocking; asynchrony is managed through the use of futures, promises and continuation callbacks, enabling the programmer to construct a graph of operations to execute asynchronously as high-latency dependencies are satisfied. A global pointer abstraction provides system-wide addressability of shared memory, including host and accelerator memories. The parallelism model is primarily process-based, but the interface is thread-safe and designed to allow efficient and expressive use in multi-threaded applications. The interface is designed for extreme scalability throughout, and deliberately avoids design features that could inhibit scalability.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Computer Science Research Needs for Parallel Discrete Event Simulation (PDES)

Historically, scientific computing efforts have demonstrated the clear need for, and effective use of, supercomputing with traditional time-stepped simulations. Nevertheless, there are several areas in the mission spaces of the U.S. Department of Energy and other agencies waiting to tap advanced computing research using a different, discrete event style of modeling, simulation, and analysis. These span a wide spectrum of applications including energy grid resilience, urban planning and policy, transportation science, building technologies, emergency response and planning, environmental impact analysis, computational epidemiology, Internet communications, cyber security, and cyber-physical systems, to name only a few. Even within traditional scientific applications, the role of discrete event modes of execution is increasing in the form of new event-based mathematical solvers such as quantized state integration methods and discrete-continuous hybrid system solvers. Co-design of advanced supercomputing hardware systems is another area that exploits discrete event simulation at its core for effective analyses. Complex systems, entity behaviors and interconnections play a significant role in all these applications, which are mapped to large-scale models with discrete event formulations. To make advancements in all the aforementioned scientific areas, many technical aspects need to be more thoroughly studied and deeply understood in parallel discrete event simulation (PDES). The unique dynamics inherent in a discrete event modeling approach, by their very nature, intersect and influence the entire stack of the computing system, including (a) the unique nature of the instruction sets exercised in PDES workloads without a predominance of high-precision floating point operations, (b) virtual time-constrained multi-threaded execution of many logical processes per processor, (c) extremely variable and difficult to predict network traffic characteristics, (d) interfaces and inter-dependencies with machine learning and artificial intelligence codes at higher software layers, and (e) highly challenging load balancing needs, especially in effectively accounting for accelerated/extremely heterogeneous computing in current and future high-performance computing systems. Efficient and accurate parallel execution of PDES workloads is also dominated by challenges in dealing with their asynchronous concurrency fundamentally present at the model level. Conservative synchronization, optimistic/speculative synchronization, and their hybrid schemes open new questions in fundamental computer science with respect to reversibility of computation and prediction (lookahead) of behaviors inherent within model codes. On the implementation front, there are relatively few scalable, general-purpose parallel discrete event simulators in the world, and even fewer have been studied on emerging hardware platforms. To enable scientific advances using PDES, the research needs in computer science must also be pursued and met in the intersection of the algorithmic and hardware-aware aspects of scalable PDES engines. This report is aimed at capturing a computer science-oriented view of this important area of research in PDES, presenting a sample of important applications with their inherent discrete event technology elements. Needs are outlined in core areas of parallel discrete event research as well as cross-cutting directions in computer science research that positively impact scientific advancements across several important application areas. A selection of priority research opportunities in advanced computing for PDES is identified to serve as reference for key research topics and their order of importance for scientific advancements.

97 MATHEMATICS AND COMPUTING↗

A Hardware and Software Co-design Framework for Energy Efficient Neuromorphic Systems

Neuromorphic systems can be realized by a variety of algorithms and architectures. A common understanding is that spiking neuromorphic designs, which encode information into spatio-temporal spiking events, are both a biologically-accurate and efficient way of processing information. However, representing the information through timing relationships induces sophisticated circuit designs in traditional CMOS-based implementations. In recent years, high-capacity resistive memory (RRAM, aka, memristor) has demonstrated great potential in mimicking synaptic behaviors. Several RRAM-based spiking neuromorphic designs exist, most of which focus on rate coding schemes. These designs simplify circuit implementations of neuron models and explore challenges such as unsatisfactory speed, resolution, and performance. As an alternative, we will explore temporal coding spiking neuromorphic systems that encode information as the relative timing of neuron activations (spikes), which have been proven to be more adaptive and energy-efficient. Developing a neuromorphic system for spiking neural network (SNN) inference and online training, however, faces some major technical challenges: (1) It lacks circuit implementation support for temporal-coding SNN to achieve satisfying power efficiency and accuracy; (2) Although existing research works have investigated memristive synapse and neuron designs for spike-timing-dependent plasticity, the non-ideal conditions in implementation, such as device variations and signal degradation, degrade online learning accuracy of large scale systems; and (3) Non-optimized, inter-layer data traffic in SNNs, leads to unnecessary data communication costs. In this project, we plan to address these challenges by a hardware and software co-design framework that incorporates solutions at the circuit, architecture, and algorithm levels. At the circuit-level, we will elaborate on the in-situ SNN processing element designs for supporting both inference and online training modes. Variation-aware schemes will be studied to improve reliability. At the architecture level, we propose a pipelined, asynchronous architecture to retain the timing resolution of spikes. At the algorithm level, we will investigate an innovative SNN training algorithm for enabling activation sparsification and reducing unnecessary data communication costs. This neuromorphic system will provide an effective solution to real-life energy-constrained applications and significantly contribute to the exploration of next-generation high-performance computing systems under the DOE context.

97 MATHEMATICS AND COMPUTING↗

Smart Contract Architectures and Templates for Blockchain-based Energy Markets (V.1.0)

Within the field of Transactive Energy Systems (TES), there is an active need for tools that can support and accelerate the development of these new grid solutions. Among the many tools available, blockchain stands out as a viable instrument that can help researchers develop decentralized, autonomous, and tamper-resistant grid applications. In this work, we explore the use of smart contracts (SCs), a subset of blockchain technology, and analyze their applicability to facilitating the implementation of TES solutions. In particular, we focus on presenting areas of opportunity and potential drawbacks, along with use cases that can benefit from this technology building upon previous research developed by Pacific Northwest National Laboratory and other research organizations. This work builds upon the fundamentals of TES and smart contract technology to develop a series of software templates that can be used by industry to build TES-oriented grid solutions. These templates are intended to be platform agnostic and take into consideration the unique properties of SCs and distributed ledger storage mechanisms to ensure actual code implementations remain aware of the limitations of the technology. The proposed templates have the potential to enable software architects to mix and match components to satisfy their application requirements, thereby reducing the number of resources required to implement blockchain-based solutions. These templates are divided into two main components—data and behavioral models. The data models are intended to help software engineers represent the underlying grid objects along with their properties in a ledger-based storage system. The behavioral models are used to describe the processes and actions that actors within a system must perform to achieve a given outcome such as registering an asset, placing a bid, and performing bid clearances. These two components are documented in a Unified Modeling Language (UML) format and are intended for use in SC-based implementations, with special behavioral considerations to account for the asynchronous properties of the underlying ledger and the typical execution model of smart contracts. Finally, future research ideas and potential extensions to this work are discussed. In particular, known limitations and potential improvements of the developed product are identified and expected to be addressed in future revisions of the template model.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Software Contribution to the L-CAPE project

The L-CAPE project utilizes asynchronous data from thousands of Linear Accelerator devices and applies data science techniques to detect anomaly of accelerator failure before the incident, as well as automatic labels of accelerator outages. The author describes her contribution progress for the L-CAPE project this summer, as well as suggestions for future interns working on the project.

Tang, Jasmine↗

UPC++ v1.0 Specification (Rev. 2023.9.0)

UPC++ is a C++ library providing classes and functions that support Partitioned Global Address Space (PGAS) programming. The key communication facilities in UPC++ are one-sided Remote Memory Access (RMA) and Remote Procedure Call (RPC). All communication operations are syntactically explicit and default to non-blocking; asynchrony is managed through the use of futures, promises and continuation callbacks, enabling the programmer to construct a graph of operations to execute asynchronously as high-latency dependencies are satisfied. A global pointer abstraction provides system-wide addressability of shared memory, including host and accelerator memories. The parallelism model is primarily process-based, but the interface is thread-safe and designed to allow efficient and expressive use in multi-threaded applications. The interface is designed for extreme scalability throughout, and deliberately avoids design features that could inhibit scalability.

97 MATHEMATICS AND COMPUTING↗

UPC++ v1.0 Programmer’s Guide (Rev. 2023.9.0)

UPC++ is a C++ library that supports Partitioned Global Address Space (PGAS) programming. It is designed for writing efficient, scalable parallel programs on distributed-memory parallel computers. The key communication facilities in UPC++ are one-sided Remote Memory Access (RMA) and Remote Procedure Call (RPC). The UPC++ control model is single program, multiple-data (SPMD), with each separate constituent process having access to local memory as it would in C++. The PGAS memory model additionally provides one-sided RMA communication to a global address space, which is allocated in shared segments that are distributed over the processes. UPC++ also features Remote Procedure Call (RPC) communication, making it easy to move computation to operate on data that resides on remote processes. UPC++ was designed to support exascale high-performance computing, and the library interfaces and implementation are focused on maximizing scalability. In UPC++, all communication operations are syntactically explicit, which encourages programmers to consider the costs associated with communication and data movement. Moreover, all communication operations are asynchronous by default, encouraging programmers to seek opportunities for overlapping communication latencies with other useful work. UPC++ provides expressive and composable abstractions designed for efficiently managing aggressive use of asynchrony in programs. Together, these design principles are intended to enable programmers to write applications using UPC++ that perform well even on hundreds of thousands of cores.

97 MATHEMATICS AND COMPUTING↗

Integration Development and Testing of Rear Transition Monitor for Beam Current Monitoring System

Addressing baseline effects in accelerator environments is crucial for accurate data acquisition and analysis, since baseline effects can obscure signal clarity and impact the reliability of beam current monitoring systems. There are many potential contributors to baseline noise, such as variations in beam dynamics, electromagnetic interference from nearby equipment, or RF interference. Previous applications of noise reduction systems don t sufficiently filter sources of asynchronous noise, so a new algorithm was implemented. A simulation dataset was created to replicate beam conditions and a Red Pitaya FPGA was used to collect data through the streaming application. A Python script was developed to implement noise reduction algorithms and efforts were made to integrate real-time data streaming with the Redis platform and Acnet Front End infrastructure.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗