Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel communication”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

UPC++ v1.0 Specification (Revision 2021.3.0)

UPC++ is a C++11 library providing classes and functions that support Partitioned Global Address Space (PGAS) programming. The key communication facilities in UPC++ are one-sided Remote Memory Access (RMA) and Remote Procedure Call (RPC). All communication operations are syntactically explicit and default to non-blocking; asynchrony is managed through the use of futures, promises and continuation callbacks, enabling the programmer to construct a graph of operations to execute asynchronously as high-latency dependencies are satisfied. A global pointer abstraction provides system-wide addressability of shared memory, including host and accelerator memories. The parallelism model is primarily process-based, but the interface is thread-safe and designed to allow efficient and expressive use in multi-threaded applications. The interface is designed for extreme scalability throughout, and deliberately avoids design features that could inhibit scalability.

97 MATHEMATICS AND COMPUTING↗

UPC++ v1.0 Specification, Revision 2022.9.0

UPC++ is a C++ library providing classes and functions that support Partitioned Global Address Space (PGAS) programming. The key communication facilities in UPC++ are one-sided Remote Memory Access (RMA) and Remote Procedure Call (RPC). All communication operations are syntactically explicit and default to non-blocking; asynchrony is managed through the use of futures, promises and continuation callbacks, enabling the programmer to construct a graph of operations to execute asynchronously as high-latency dependencies are satisfied. A global pointer abstraction provides system-wide addressability of shared memory, including host and accelerator memories. The parallelism model is primarily process-based, but the interface is thread-safe and designed to allow efficient and expressive use in multi-threaded applications. The interface is designed for extreme scalability throughout, and deliberately avoids design features that could inhibit scalability.

97 MATHEMATICS AND COMPUTING↗

UPC++ v1.0 Specification, Revision 2020.3.0

UPC++ is a C++11 library providing classes and functions that support Partitioned Global Address Space (PGAS) programming. The key communication facilities in UPC++ are one-sided Remote Memory Access (RMA) and Remote Procedure Call (RPC). All communication operations are syntactically explicit and default to non-blocking; asynchrony is managed through the use of futures, promises and continuation callbacks, enabling the programmer to construct a graph of operations to execute asynchronously as high-latency dependencies are satisfied. A global pointer abstraction provides system-wide addressability of shared memory, including host and accelerator memories. The parallelism model is primarily process-based, but the interface is thread-safe and designed to allow efficient and expressive use in multi-threaded applications. The interface is designed for extreme scalability throughout, and deliberately avoids design features that could inhibit scalability.

97 MATHEMATICS AND COMPUTING↗

Coupled Lattice Boltzmann Modeling Framework for Pore-Scale Fluid Flow and Reactive Transport

In this paper, we propose a modeling framework for pore-scale fluid flow and reactive transport based on a coupled lattice Boltzmann model (LBM). We develop a modeling interface to integrate the LBM modeling code parallel lattice Boltzmann solver and the PHREEQC reaction solver using multiple flow and reaction cell mapping schemes. The major advantage of the proposed workflow is the high modeling flexibility obtained by coupling the geochemical model with the LBM fluid flow model. Consequently, the model is capable of executing one or more complex reactions within desired cells while preserving the high data communication efficiency between the two codes. Meanwhile, the developed mapping mechanism enables the flow, diffusion, and reactions in complex pore-scale geometries. We validate the coupled code in a series of benchmark numerical experiments, including 2D single-phase Poiseuille flow and diffusion, 2D reactive transport with calcite dissolution, as well as surface complexation reactions. The simulation results show good agreement with analytical solutions, experimental data, and multiple other simulation codes. In addition, we design an AI-based optimization workflow and implement it on the surface complexation model to enable increased capacity of the coupled modeling framework. Compared to the manual tuning results proposed in the literature, our workflow demonstrates fast and reliable model optimization results without incorporating pre-existing domain knowledge.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

UPC++ v1.0 Programmer’s Guide, Revision 2021.9.0

UPC++ is a C++ library that provides Partitioned Global Address Space (PGAS) programming. It is designed for writing parallel programs that run efficiently and scale well on distributed-memory parallel computers. The PGAS model is single program, multiple-data (SPMD), with each separate constituent process having access to local memory as it would in C++. PGAS additionally provides one-sided Remote Memory Access (RMA) to a global address space, which is allocated in shared segments that are distributed over the processes. UPC++ also features Remote Procedure Call (RPC) communication, making it easy to move computation to operate on data that resides on remote processes. In UPC++, all communication operations are explicit, which encourages programmers to be aware of the cost of communication and data movement. Moreover, all communication operations are asynchronous by default, to enable programmers to write code that scales well even on hundreds of thousands of cores.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Decentralized Schemes with Overlap for Solving Graph-Structured Optimization Problems

We present a new algorithmic paradigm for the decentralized solution of graph-structured optimization problems that arise in the estimation and control of network systems. A key and novel design concept of the proposed approach is that it uses overlapping subdomains to promote and accelerate convergence. We show that the algorithm converges if the size of the overlap is sufficiently large and that the convergence rate improves exponentially with the size of the overlap. The proposed approach provides a bridge between fully decentralized and centralized architectures and is flexible in that it enables the implementation of asynchronous schemes, handling of constraints, and balancing of computing, communication, and data privacy needs. The proposed scheme is tested in an estimation problem for a 9241-node power network and we show that it outperforms the alternating direction method of multipliers.

asynchronous↗

A Low Voltage DC Power Electronic Hub to Support Buildings

This paper presents the communication, control, and architecture, for a low voltage (nominal 480V) hybrid AC/DC microgrid for supporting small commercial buildings with critical data center loads. The proposed system is based on a dc-power electronic hub (PEH) that seamlessly integrates renewable energy resources, energy storage, and back-up generation to support commercial building economical energy management and reliability. This PEH concept provides a parallelization of critical power to the commercial and industrial buildings leading to increased efficiency and reliability compared to traditional uninterruptable power supplies. The demonstration of the proposed PEH architecture and controls is conducted through a controller hardware in the loop validation.

Starke, Michael↗

Computer Science Research Needs for Parallel Discrete Event Simulation (PDES)

Historically, scientific computing efforts have demonstrated the clear need for, and effective use of, supercomputing with traditional time-stepped simulations. Nevertheless, there are several areas in the mission spaces of the U.S. Department of Energy and other agencies waiting to tap advanced computing research using a different, discrete event style of modeling, simulation, and analysis. These span a wide spectrum of applications including energy grid resilience, urban planning and policy, transportation science, building technologies, emergency response and planning, environmental impact analysis, computational epidemiology, Internet communications, cyber security, and cyber-physical systems, to name only a few. Even within traditional scientific applications, the role of discrete event modes of execution is increasing in the form of new event-based mathematical solvers such as quantized state integration methods and discrete-continuous hybrid system solvers. Co-design of advanced supercomputing hardware systems is another area that exploits discrete event simulation at its core for effective analyses. Complex systems, entity behaviors and interconnections play a significant role in all these applications, which are mapped to large-scale models with discrete event formulations.

97 MATHEMATICS AND COMPUTING↗

Laser Spectroscopy of Exotic Atoms and Molecules Containing Octupole-Deformed Nuclei

This project investigated the nuclear, atomic, and molecular structure of exotic atoms and molecules containing actinide isotopes. These short-lived radioactive systems are challenging to produce and study in the laboratory, yet they offer unique opportunities for fundamental science. These nuclei are predicted to exhibit pear-shaped (octupole) deformation, a rare collective nuclear behavior that dramatically enhances their sensitivity to fundamental physics phenomena such as time-reversal- and parity-violating effects. Such enhancements make them ideal probes for exploring open questions in our understanding of the universe such as the origin of the matter–antimatter asymmetry of the universe. To realize these measurements, the project led the development of a new laser spectroscopy experiment at MIT, and later commissioned at the Facility for Rare Isotope Beams (FRIB) at Michigan State University: the Resonance Ionization Spectroscopy Experiment (RISE). RISE combines the spectroscopic precision of collinear laser spectroscopy with the sensitivity of particle-detection techniques, enabling measurements of rare isotopes produced at rates as low as a few ions per second. The beamline was designed, built, and installed, and was successfully commissioned at FRIB during the grant period. RISE is now a permanent capability of the FRIB facility, and has produced several results on the study of rare atoms and molecules for nuclear structure and fundamental symmetries. In parallel with the FRIB program, the project contributed to the first precision laser-spectroscopy measurements of short-lived radioactive molecules. Working with international collaborators at CERN's ISOLDE facility, the team conducted pioneering experiments on radium monofluoride (RaF) and actinium monofluoride (AcF). The results from this work have been published in major journals of science, including Nature, Science, Nature Physics, Nature Communications, and Physical Review Letters. These findings have guided future experiments on the laser cooling of radioactive molecules, opening a new platform for precision tests of fundamental symmetries.

38 RADIATION CHEMISTRY, RADIOCHEMISTRY, AND NUCLEA↗

ParFlow Sand Tank: A tool for groundwater exploration

The ParFlow Sand Tank model is an open source application designed to allow users to interactively simulate and visualize groundwater movement through the subsurface. The app is designed for both research and education; teaching hydrogeology concepts and making it easy explore and run sophisticated groundwater simulations. Our goal is to support increased accessibility and usability of research grade hydrology tools for research and teaching. The Sand Tank application simulates groundwater and surface water fluxes as well as contaminant transport in real time using the integrated physical hydrology model ParFlow (Kollet & Maxwell, 2006; Maxwell & Miller, 2005; Osei-Kuffuor et al., 2014) and the particle tracking code EcoSlim (Maxwell et al., 2019). ParFlow is a numerical hydrology model that simulates spatially distributed groundwater and surface water flow. It is a well established research tool with more than 90 publications documenting its development use to advance our understanding of groundwater dynamics and groundwater surface water interactions from the hillslope to the continental scale e.g. (Condon et al., 2020; Condon & Maxwell, 2019; Maxwell & Condon, 2016). It is designed for efficient parallel computation and has been run on many platforms spanning from laptops to supercomputers. However, one of the challenges of ParFlow is that it requires significant training and hydrologic expertise to develop simulations. The Sand Tank application makes this model accessible to anyone for education and exploration. Our application uses ParFlow for its simulation backend and ParaView for the data loading and processing. The communication infrastructure relies on the ParaViewWeb framework. We use model templates deployed in Docker images to setup the Sand Tank framework. Users can build the application locally or interact with it through our web deployment. When interacting with a template users can interactively change model parameters like subsurface processes or pump/inject water into the subsurface and watch the system respond to their changes in real time as the simulation runs. Additionally, our template setup will allow more advanced users to build custom templates of increasing complexity for both research and educational purposes.

54 ENVIRONMENTAL SCIENCES↗

Robust Decentralized Secondary Control Scheme for Inverter-based Power Networks

Inverter-dominated microgrids are quickly becoming a key building block of future power systems. They rely on centralized controllers that can provide reliability and resiliency in extreme events. Nonetheless, communication failures due to cyber-physical attacks or natural disasters can make autonomous operation of islanded microgrids challenging. This paper examines a unified decentralized secondary control scheme that is robust to inverter clock synchronization errors and can be seamlessly applied to grid-following or grid-forming control architectures. The proposed scheme overcomes the well-known stability problem that arises from parallel operation of local integral controllers. Theoretical guarantees for stability are provided along with criteria to appropriately tune the secondary control gains to achieve good frequency regulation performance while ensuring fair power sharing. The efficacy of our approach is demonstrated through simulations on a 5-bus microgrid with four grid-forming inverters.

24 POWER TRANSMISSION AND DISTRIBUTION↗

A minimally invasive, efficient method for propagation of full-field uncertainty in solid dynamics

In this work, we present a minimally invasive method for forward propagation of material property uncertainty to full-field quantities of interest in solid dynamics. Full-field uncertainty quantification enables the design of complex systems where quantities of interest, such as failure points, are not known a priori. The method, motivated by the well-known probability density function (PDF) propagation method of turbulence modeling, uses an ensemble of solutions to provide the joint PDF of desired quantities at every point in the domain. A small subset of the ensemble is computed exactly, and the remainder of the samples are computed with approximation of the evolution equations based on those exact solutions. Although the proposed method has commonalities with traditional interpolatory stochastic collocation methods applied directly to quantities of interest, it is distinct and exploits the parameter dependence and smoothness of the driving term of the evolution equations. The implementation is model independent, storage and communication efficient, and straightforward. We demonstrate its efficiency, accuracy, scaling with dimension of the parameter space, and convergence in distribution with two problems: a quasi-one-dimensional bar impact, and a two material notched plate impact. For the bar impact problem, we provide an analytical solution to PDF of the solution fields for method validation. With the notched plate problem, we also demonstrate good parallel efficiency and scaling of the method.

42 ENGINEERING↗

Clang UPC2C Translator (Clang UPC2C) v9.0.1-1

Clang Unified Parallel C 2 C (Clang UPC2C) translator compiles programs written in the UPC (Unified Parallel C) language to ISO C99, with calls to the Berkeley UPC runtime system. The Clang UPC2C compiler extends the capabilities of the Clang LLVM C frontend to comply with the UPC Language Specification version 1.3. The compiler generates programs that run on a wide variety of systems ranging from workstations to leadership-class supercomputers, in conjunction with the Berkeley UPC runtime and GASNet communication system. In addition to the standard UPC libraries, Clang UPC2C also provides access to Berkeley UPC library extensions.

Hargrove, Paul↗

PETSc/TAO developments for GPU-based early exascale systems

The Portable Extensible Toolkit for Scientific Computation (PETSc) library provides scalable solvers for nonlinear time-dependent differential and algebraic equations and for numerical optimization via the Toolkit for Advanced Optimization (TAO). PETSc is used in dozens of scientific fields and is an important building block for many simulation codes. During the U.S. Department of Energy’s Exascale Computing Project, the PETSc team has made substantial efforts to enable efficient utilization of the massive fine-grain parallelism present within exascale compute nodes and to enable performance portability across exascale architectures. We recap some of the challenges that designers of numerical libraries face in such an endeavor, and then discuss the many developments we have made, which include the addition of new GPU backends, features supporting efficient on-device matrix assembly, better support for asynchronicity and GPU kernel concurrency, and new communication infrastructure. In conclusion, we evaluate the performance of these developments on some pre-exascale systems as well as the early exascale systems Frontier and Aurora, using compute kernel, communication layer, solver, and mini-application benchmark studies, and then close with a few observations drawn from our experiences on the tension between portable performance and other goals of numerical libraries.

Exascale Computing Project (ECP)↗

Clang UPC Compiler (Clang UPC) v3.9.1-1

Clang Unified Parallel C (Clang UPC) provides a compilation and execution environment for programs written in the UPC (Unified Parallel C) language. The Clang UPC compiler extends the capabilities of the Clang LLVM C compiler to comply with the UPC Language Specification version 1.3. It includes support for UPC collectives and a configurable pointer-to-shared representation. The compiler generates programs that run on a wide variety of systems ranging from workstations to leadership-class supercomputers, in conjunction with the Berkeley UPC runtime and GASNet communication system. This compiler generates assembly code / object code directly for Intel processors and IBM PowerPC.

Hargrove, Paul↗

Wave Tank Testing Report for Controls Validation of a Heaving Point Absorber

The core objectives of this project is to improve the power capture of three different wave energy conversion (WEC) devices by more than 50% using an advanced control system and validate the attained improvements using wave tank and full scale testing. In parallel, we will bring along the development of a wave prediction system that is required to enable effective control and test it at full scale. The purposes of this report are to: 1. Plan and document the 1/25th scale device testing at the wave-tank facility; 2. Document the test article, setup and methodology, sensor and instrumentation, mooring, electronics, wiring, and data flow and quality assurance; 3. Communicate the testing results between the associated members; 4. Facilitate reviews that will help to ensure all aspects (risk, safety, testing procedures, etc.); 5. Provide a systematic guide to setting up, executing and decommissioning the experiment.

16 TIDAL AND WAVE POWER↗

Direct numerical simulations for hybrid rocket boundary layers: Performance modeling and scaling

This paper presents a comprehensive performance and scaling analysis of direct numerical simulations for reacting boundary layers, focusing on slab burner configurations. Using a PETSc-based finite volume CFD framework, the study evaluates the scalability and computational cost of flow, chemistry, and radiation evaluations across 2D and 3D simulations. Polymethyl methacrylate (PMMA) is the fuel with pure O 2 as the oxidizer, modeled using a detailed chemical kinetics mechanism with 113 species and 660 reactions. A ray-tracing-based radiation solver, designed for distributed memory applications, is implemented to model radiation heat transfer. Parallel scalability is analyzed for the coupled flow, chemistry, and radiation heat transfer processes. Weak and strong scaling studies are conducted on up to 15,000 computational ranks, revealing robust performance when flow cells exceed 200 per rank. Chemistry evaluations dominate the computational cost in large 3D simulations, accounting for approximately 40% of the total runtime, while flow processes contribute around 35%, and radiation solver contributions remain below 10% due to reduced evaluation frequencies. GPU accelerated chemistry evaluation, implemented with Zero-RK, demonstrates significant promise, achieving up to a 4x speedup for workloads exceeding 30,000 cells per GPU. However, diminishing returns are observed for smaller workloads due to CPU-GPU communication overhead. This study identifies key challenges, including memory bottlenecks and the effects of domain partitioning on flow scalability, while highlighting the potential of GPU-accelerated chemistry to reduce computational costs. In conclusion, these findings provide realizable run configurations for 2D, 3D, and GPU-accelerated cases, offering insights for optimizing reactive flow solvers.

CFD Scalability↗

Steerable terahertz beams using surface waves on an active metasurface

The development of dynamic components for controlling wave fronts in the sub-terahertz region of the electromagnetic spectrum has emerged as a frontier research topic for many applications in sensing and communications. One approach which has attracted much attention involves the use of active metasurfaces, tiled arrays of sub-wavelength elements with properties that can be reconfigured via external actuation. In nearly all cases, these metasurfaces are employed as either transmissive or reflective elements, taking advantage of their strong and tunable interaction with free-space electromagnetic waves. These interactions can be significantly enhanced through the use of surface waves propagating parallel to the metasurface array, although very few studies have exploited this option. Here, we integrate a metasurface into the interior of a parallel-plate waveguide in a configuration explicitly designed to exploit this surface-wave geometry. We show that varying the electrical properties of the active metasurface changes the wave vector of the guided mode, and thereby alters the emission angle of radiation out-coupled through a leaky-wave slot aperture. These results, which are consistent with numerical simulations, represent a new approach to broadband beam steering suitable for the sub-terahertz spectral range.

77 NANOSCIENCE AND NANOTECHNOLOGY↗