Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel communication”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Computer Science Research Needs for Parallel Discrete Event Simulation (PDES)

Historically, scientific computing efforts have demonstrated the clear need for, and effective use of, supercomputing with traditional time-stepped simulations. Nevertheless, there are several areas in the mission spaces of the U.S. Department of Energy and other agencies waiting to tap advanced computing research using a different, discrete event style of modeling, simulation, and analysis. These span a wide spectrum of applications including energy grid resilience, urban planning and policy, transportation science, building technologies, emergency response and planning, environmental impact analysis, computational epidemiology, Internet communications, cyber security, and cyber-physical systems, to name only a few. Even within traditional scientific applications, the role of discrete event modes of execution is increasing in the form of new event-based mathematical solvers such as quantized state integration methods and discrete-continuous hybrid system solvers. Co-design of advanced supercomputing hardware systems is another area that exploits discrete event simulation at its core for effective analyses. Complex systems, entity behaviors and interconnections play a significant role in all these applications, which are mapped to large-scale models with discrete event formulations. To make advancements in all the aforementioned scientific areas, many technical aspects need to be more thoroughly studied and deeply understood in parallel discrete event simulation (PDES). The unique dynamics inherent in a discrete event modeling approach, by their very nature, intersect and influence the entire stack of the computing system, including (a) the unique nature of the instruction sets exercised in PDES workloads without a predominance of high-precision floating point operations, (b) virtual time-constrained multi-threaded execution of many logical processes per processor, (c) extremely variable and difficult to predict network traffic characteristics, (d) interfaces and inter-dependencies with machine learning and artificial intelligence codes at higher software layers, and (e) highly challenging load balancing needs, especially in effectively accounting for accelerated/extremely heterogeneous computing in current and future high-performance computing systems. Efficient and accurate parallel execution of PDES workloads is also dominated by challenges in dealing with their asynchronous concurrency fundamentally present at the model level. Conservative synchronization, optimistic/speculative synchronization, and their hybrid schemes open new questions in fundamental computer science with respect to reversibility of computation and prediction (lookahead) of behaviors inherent within model codes. On the implementation front, there are relatively few scalable, general-purpose parallel discrete event simulators in the world, and even fewer have been studied on emerging hardware platforms. To enable scientific advances using PDES, the research needs in computer science must also be pursued and met in the intersection of the algorithmic and hardware-aware aspects of scalable PDES engines. This report is aimed at capturing a computer science-oriented view of this important area of research in PDES, presenting a sample of important applications with their inherent discrete event technology elements. Needs are outlined in core areas of parallel discrete event research as well as cross-cutting directions in computer science research that positively impact scientific advancements across several important application areas. A selection of priority research opportunities in advanced computing for PDES is identified to serve as reference for key research topics and their order of importance for scientific advancements.

97 MATHEMATICS AND COMPUTING↗

Distributed Real-time Plume Monitoring for Deep Sea Mineral Extraction​

In the emerging industry of deep-sea mining for minerals and deposits (e.g. polymetallic nodules for nickel, cobalt, copper, and manganese), more data is required to understand the effects of sediment plume generation and predict the distribution of disturbed sediment. There are two main sources of plume generation, the first being at the active mining site where the “collector” directly removes the top layer of the sea floor. The other is the “midwater plume” consisting of unwanted sediment that was collected during extraction that is pumped back into the aphotic zone. The vast majority of plume generation is caused by the collector, causing detrimental and long-lasting impacts on seafloor ecosystems due to the lack of wave activity or strong currents at the sea floor. Therefore, it is crucial to invest in the infrastructure to support the study and constant monitoring over a large area of the sea floor where plume generation is present. Due to the limited number of usable channels and power requirements, current subsea wireless communications technologies are not well suited to instrumenting the large areas of the sea floor needed to monitor plume migration. The scope of this effort is to transition experimental demonstrations of high-bandwidth, full-duplex scalable underwater laser communications to the seafloor in an open ocean environment. Specifically tackling challenges associated with the dynamic nature of the subsea world, including but not limited to, deployment logistics, sustainability, and range. The goal is to enable the internet of underwater things for deep sea industries by broadening the capabilities of subsea communications. By using high-precision laser transmitters, many of the challenges current subsea optical systems face can be circumvented, such as power consumption, interference, and bandwidth limitations. This approach lends itself to wireless interlinking multi-node networks, in series or parallel, facilitating the implementation of a wide array of sensor types. This interlinking allows all the data gathered from the network to be processed through a single hardline uplink to the surface, lowering the complexity required for near real-time data processing. Additionally, the laser control systems produce metadata that can be used to help characterize the water column between the nodes. Combining data from various sensors such as turbidity, temperature, current velocity with metadata such as beam attenuation and deflection can produce a high-resolution model of sea floor conditions around an active mining zone. The resulting near real-time model can be used to optimize location and flow rate of the mining operation to minimize and quantify the environmental impact.

Mons, Ishan↗

Scalable Asynchronous Domain Decomposition Solvers

We discuss how parallel implementations of linear iterative solvers generally alternate between phases of data exchange and phases of local computation. Increasingly large problem sizes and more heterogeneous compute architectures make load balancing and the design of low latency network interconnects that are able to satisfy the communication requirements of linear solvers very challenging tasks. In particular, global communication patterns such as inner products become increasingly limiting at scale. We explore the use of asynchronous communication based on one-sided Message Passing Interface primitives in the context of domain decomposition solvers. In particular, a scalable asynchronous two-level Schwarz method is presented. We discuss practical issues encountered in the development of a scalable solver and show experimental results obtained on a state-of-the-art supercomputer system that illustrate the benefits of asynchronous solvers in load balanced as well as load imbalanced scenarios. Using the novel method, we can observe speedups of up to four times over its classical synchronous equivalent.

97 MATHEMATICS AND COMPUTING↗

GLUE Code: A framework handling communication and interfaces between scales

Many scientific applications are inherently multiscale in nature. Such complex physical phenomena often require simultaneous execution and coordination of simulations spanning multiple time and length scales. This is possible by combining expensive small-scale simulations (such as molecular dynamics simulations) with larger scale simulations (such continuum limit/hydro solvers) to allow for considerably larger systems using task and data parallelism. However, the granularity of the tasks can be very large and often leads to load imbalance. Traditionally, we use approximations to streamline the computation of the more costly interactions and this introduces trade-offs between simulation cost and accuracy. In recent years, the available computational power and the advances in machine learning have made computing these scale-bridging interactions and multiscale simulations more feasible. One driving application has been plasma modeling in inertial confinement fusion (ICF), which is fundamentally multiscale in nature. This requires deep understanding of how to extrapolate microscopic information into macroscopically relevant scales. For example, in ICF one needs an accurate understanding of the connection between experimental observables and the underlying microphysics. The properties of the larger scales are often affected by the microscale behavior incorporated usually into the equations of state and ionic and electronic transport coefficients (Liboff, 1959; Rinderknecht et al., 2014; Rosenberg et al., 2015; Ross et al., 2017). Instead of incorporating this information using reliable molecular dynamics (MD) simulations, one often needs to use theoretical models, due to the inability of MD to reach engineering scales (Glosli et al., 2007; Marinak et al., 1998). One approach to resolve this issue is by coupling two MD simulations of different scales via force interpolation, e.g., the AdResS method (Krekeler et al., 2018; Nagarajan et al., 2013). Another approach, which we will pursue in the scope of this work, is by enabling scale bridging between MD simulations and meso/macro-scale models through the development and support of application programming interfaces that these different applications can interact through.

54 ENVIRONMENTAL SCIENCES↗

Lessons Learned on the Interface Between Quantum and Conventional Networking

The future Quantum Internet is expected to be based on a hybrid architecture with core quantum transport capabilities complemented by conventional networking. Practical and foundational considerations indicate the need for conventional control and data planes that (i) utilize extensive existing telecommunications fiber infrastructure, and (ii) provide parallel conventional data channels needed for quantum networking protocols. We propose a quantum-conventional network (QCN) harness to implement a new architecture to meet these requirements. The QCN control plane carries the control and management traffic, whereas its data plane handles the conventional and quantum data communications. We established a local area QCN connecting three quantum laboratories over dedicated fiber and conventional network connections. We describe considerations and tradeoffs for layering QCN functionalities, informed by our recent quantum entanglement distribution experiments conducted over this network.

Alshowkan, Muneer↗

Electrical control of magnetism by electric field and current-induced torques

The remanent magnetization of ferromagnets has long been studied and used to store binary information. While early magnetic memory designs relied on magnetization switching by locally generated magnetic fields, key insights in condensed matter physics later suggested the possibility of doing it by electrical means instead. In the 1990s, Slonczewski and Berger formulated the concept of current-induced spin torques in magnetic multilayers through which a spin-polarized current generated by a first ferromagnet may be used to switch the magnetization of a second one. This discovery drove the development of spin-transfer-torque magnetic random-access memories (MRAMs). More recent fundamental research revealed other types of current-induced torques named spin-orbit torques (SOTs) and will lead to a new generation of devices including SOT MRAMs and skyrmion-based devices. Parallel to these advances, multiferroics and their magnetoelectric coupling, first investigated experimentally in the 1960s, experienced a renaissance. Dozens of multiferroic compounds with new magnetoelectric coupling mechanisms were discovered and high-quality multiferroic films were synthesized (notably of BiFeO 3 ), also leading to novel device concepts for information and communication technology such as the magnetoelectric spin-orbit (MESO) transistor. The story of the electrical switching of magnetization, which is discussed in this review, is that of a dance between fundamental research (in spintronics, condensed matter physics, and materials science) and technology (MRAMs, MESO transistors, microwave emitters, spin diodes, skyrmion-based devices, components for neuromorphics, etc.). This pas de deux has led to major scientific and technological breakthroughs in recent decades (such as the conceptualization of pure spin currents, the observation of magnetic skyrmions, and the discovery of spin-charge interconversion effects). As a result, this field has not only propelled MRAMs into consumer electronics products but also fueled discoveries in adjacent research areas such as ferroelectrics or magnonics. Here, in this review, recent advances in the control of magnetism by electric fields and by current-induced torques are covered. Fundamental concepts in these two directions are reviewed first, their combination is then discussed, and finally current various families of devices harnessing the electrical control of magnetic properties for various application fields are addressed. The review concludes by giving perspectives in terms of both emerging fundamental physics concepts and new directions in materials science.

36 MATERIALS SCIENCE↗

Heterogeneous graphics processing unit for scheduling thread groups for execution on variable width SIMD units

A compute unit configured to execute multiple threads in parallel is presented. The compute unit includes one or more single instruction multiple data (SIMD) units and a fetch and decode logic. The SIMD units have differing numbers of arithmetic logic units (ALUs), such that each SIMD unit can execute a different number of threads. The fetch and decode logic is in communication with each of the SIMD units, and is configured to assign the threads to the SIMD units for execution based on such differing numbers of ALUs.

97 MATHEMATICS AND COMPUTING↗

Optimal Control Strategy With Efficiency and Reliability Improvement for Offshore DC Microgrids

Offshore microgrids, due to their remote location and lack of external energy support, face significant challenges in wide-range load operation and maintenance. Consequently, efficiency and reliability are critical concerns for converters in offshore dc microgrids. This article presents an optimal control strategy aimed at enhancing both efficiency and reliability. A normalized nonlinear relationship between power loss and thermal stress of a paralleled converter is first established. Based on this, a dual-objective optimization function with an active weight function as well as a system overall performance index is established. The active weight function dynamically adjusts the control priority based on converter efficiency and switching device thermal stress. Then, the optimal power-sharing strategy is derived by the Lagrange multiplier method with the proposed optimal function. Additionally, to accommodate a wide load range, an optimal selection strategy for operating converter combinations is proposed, requiring only low-bandwidth communication. Experiment verification is given to validate the effectiveness of the proposed control strategy. The experiment results demonstrate that the proposed control strategy can improve the overall performance of offshore microgrids by optimizing efficiency and reliability.

24 POWER TRANSMISSION AND DISTRIBUTION↗

SAIGE-GPU: accelerating genome- and phenome-wide association studies using GPUs

Genome-wide association studies (GWAS) at biobank scale are computationally intensive, especially for admixed populations requiring robust statistical models. SAIGE is a widely used method for generalized linear mixed-model GWAS but is limited by its CPU-based implementation, making phenome-wide association studies impractical for many research groups. We developed SAIGE-GPU, a GPU-accelerated version of SAIGE that replaces CPU-intensive matrix operations with GPU-optimized kernels. The core innovation is distributing genetic relationship matrix calculations across GPUs and communication layers. Applied to 2068 phenotypes from 635 969 participants in the Million Veteran Program, including diverse and admixed populations, SAIGE-GPU achieved a 5-fold speedup in mixed model fitting on supercomputing infrastructure and cloud platforms. We further optimized the variant association testing step through multi-core and multi-trait parallelization. Deployed on Google Cloud Platform and Azure, the method provided substantial cost and time savings. Source code and binaries are available for download at https://github.com/saigegit/SAIGE/tree/SAIGE-GPU-1.3.3. A code snapshot is archived at Zenodo for reproducibility (DOI: [10.5281/zenodo.17642591]). SAIGE-GPU is available in a containerized format for use across HPC and cloud environments and is implemented in R/C++ and runs on Linux systems.

Rodriguez, Alex [Argonne National Laboratory (ANL)↗

Vulnerability of a VOC-Based Inverter Due to Noise Injection and Its Mitigation

Even though virtual oscillator control (VOC)-based inverters do not communicate with each other, they need to make local measurements for control. The impact of tampering with these measured or sensed signals on the performance of a VOC-based inverter and synchronization of multiple such inverters is an important but open-ended issue. As such, this letter explores the impact of intentional side-channel noise intrusion (SNI) on the synchronization of VOC-based communication-free self-synchronizing inverters (CFSIs). Two different scenarios are investigated via experimental and analytical studies using a half-bridge neutral point clamped (NPC) single-phase CFSI. Furthermore, they address the impact of SNI on the ability of a CFSI to ensure a stable 60-Hz limit cycle and on the parallel operation of two such CFSIs to ensure synchronism to a common 60-Hz load frequency.

42 ENGINEERING↗

Probing Particle Impingement in Boilers Using High-Performance Computing with Parallel CPUs and GPUs

The major goals of the project are to calculate and analyze particle impingement within boilers, quantify effects of particulates in boilers, and predict damage rates of boilers under different cycling modes. Collectively, these initiatives develop insight into existing coal plant challenges using advanced modeling tools, particularly those leveraging high-performance computing resources. High-performance CFD computing forms a central theme in this project that will employ a high degree of coordination and communication between these initiatives to realize a final, rigorously sound, and validated computational capability upon completion. These results will create a holistic, comprehensive, systems-level assessment of damage rates under different cycling modes. Together, these objectives will develop critical insight into damage mechanisms in existing coal plant challenges for accurately and efficiently assessing operating performance in fossil energy power plants.

20 FOSSIL-FUELED POWER PLANTS↗

5G Securely Energized and Resilient: (5G-SER) (Final Report)

NREL's work on 5G integration with physical power systems (versus simulated systems demonstrated in task 3) under 5G-Securely Energized and Resilient (5G-SER) achieved a major milestone in with the completion of Task 4 activities. After many months (6+) of planning and development along with scaling challenges along the way, a full 5G end-to-end network with physical hardware components (physical inverter, physical power panel, physical battery, etc.) was deployed in a containerized environment. In parallel, a distributed controls architecture for a microgrid powering 5G resources was also successfully modified from its previous instantiation for a simulated environment to work with these physical components. These accomplishments set the stage to use newly setup 5G technology with physical microgrid components to test the feasibility of 5G wireless with physical systems to enable resilient communications between controls and distributed solar and storage resources while exploring ways to configure 5G components to survive power disturbances. The results detailed in the report below show that even utilizing physical components, 5G wireless systems were able to provide resilient results that were similar to the results from the simulated environment Task 3).

24 POWER TRANSMISSION AND DISTRIBUTION↗

Efficient Parallelization of Irregular Applications on GPU Architectures

With the enlarging computation capacity of general Graphics Processing Units (GPUs), leveraging GPUs to accelerate parallel applications has become a critical topic in academia and industry. However, a wide range of irregular applications with the computation-/memory-intensive nature cannot easily achieve high GPU utilization. The challenges mainly involve the following aspects: first, data dependence leads to coarse-grained kernel and inefficient parallelism; second, heavy GPU memory usage may cause frequent memory evictions and extra overhead of I/O; third, specific computation patterns produce memory redundancies; last, workload balance and data reusability conjunctly benefit the overall performance, but there may exist a dynamic trade-off between them. Targeting these challenges, this dissertation proposes multiple optimizations to accelerate two real-world applications: many-body correlation functions to simulate nuclear physics in a large-scale scientific system; the other is the eALS-based matrix factorization recommendation system. To accelerate the calculations of many-body correlation functions, this dissertation presents three frameworks in GPU memory management and multi-GPU scheduling. Firstly, an optimized systematic GPU memory management framework, MemHC, utilizes a series of new memory reduction designs in GPU memory allocation, CPU/GPU communications, and GPU memory oversubscription. Secondly, an enhanced multi-GPU scheduling framework, MICCO, particularly by taking both data dimension (e.g., data reuse and data eviction) and computation dimension into account. MICCO designs a heuristic scheduling algorithm and a machine learning-based regression model to generate the optimal settings of a proposed new concept to manage the trade-off. Thirdly, a locality-aware multi-GPU scheduling framework. This scheduler leverages pipeline batch generation with a looking-ahead strategy by building local dependency graphs for memory transfer reduction and better data reuse, achieving up to 79.92% memory cost reduction and 1.67x speedup. To parallelize the eALS-based recommendation system, this dissertation proposes an efficient CPU/GPU heterogeneous recommendation system, HEALS. HEALS employs newly designed architecture-adaptive data formats to achieve load balance and good data locality on CPU and GPU. To mitigate the data dependence, HEALS presents a CPU/GPU collaboration model for both task parallelism and data parallelism with multiple kernel computation optimizations. In summary, this dissertation efficiently accelerates two typical irregular applications on GPUs by building four frameworks, including CPU/GPU collaboration, GPU memory management, and multi-GPU scheduling.

Wang, Qihan↗

UPC++ v1.0 Specification, Revision 2020.10.0

UPC++ is a C++11 library providing classes and functions that support Partitioned Global Address Space (PGAS) programming. The key communication facilities in UPC++ are one-sided Remote Memory Access (RMA) and Remote Procedure Call (RPC). All communication operations are syntactically explicit and default to non-blocking; asynchrony is managed through the use of futures, promises and continuation callbacks, enabling the programmer to construct a graph of operations to execute asynchronously as high-latency dependencies are satisfied. A global pointer abstraction provides system-wide addressability of shared memory, including host and accelerator memories. The parallelism model is primarily process-based, but the interface is thread-safe and designed to allow efficient and expressive use in multi-threaded applications. The interface is designed for extreme scalability throughout, and deliberately avoids design features that could inhibit scalability.

97 MATHEMATICS AND COMPUTING↗

UPC++ v1.0 Specification, Revision 2021.9.0

UPC++ is a C++ library providing classes and functions that support Partitioned Global Address Space (PGAS) programming. The key communication facilities in UPC++ are one-sided Remote Memory Access (RMA) and Remote Procedure Call (RPC). All communication operations are syntactically explicit and default to non-blocking; asynchrony is managed through the use of futures, promises and continuation callbacks, enabling the programmer to construct a graph of operations to execute asynchronously as high-latency dependencies are satisfied. A global pointer abstraction provides system-wide addressability of shared memory, including host and accelerator memories. The parallelism model is primarily process-based, but the interface is thread-safe and designed to allow efficient and expressive use in multi-threaded applications. The interface is designed for extreme scalability throughout, and deliberately avoids design features that could inhibit scalability.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

UPC++ v1.0 Specification (Rev. 2023.9.0)

UPC++ is a C++ library providing classes and functions that support Partitioned Global Address Space (PGAS) programming. The key communication facilities in UPC++ are one-sided Remote Memory Access (RMA) and Remote Procedure Call (RPC). All communication operations are syntactically explicit and default to non-blocking; asynchrony is managed through the use of futures, promises and continuation callbacks, enabling the programmer to construct a graph of operations to execute asynchronously as high-latency dependencies are satisfied. A global pointer abstraction provides system-wide addressability of shared memory, including host and accelerator memories. The parallelism model is primarily process-based, but the interface is thread-safe and designed to allow efficient and expressive use in multi-threaded applications. The interface is designed for extreme scalability throughout, and deliberately avoids design features that could inhibit scalability.

97 MATHEMATICS AND COMPUTING↗

UPC++ v1.0 Specification (Revision 2022.3.0)

UPC++ is a C++ library providing classes and functions that support Partitioned Global Address Space (PGAS) programming. The key communication facilities in UPC++ are one-sided Remote Memory Access (RMA) and Remote Procedure Call (RPC). All communication operations are syntactically explicit and default to non-blocking; asynchrony is managed through the use of futures, promises and continuation callbacks, enabling the programmer to construct a graph of operations to execute asynchronously as high-latency dependencies are satisfied. A global pointer abstraction provides system-wide addressability of shared memory, including host and accelerator memories. The parallelism model is primarily process-based, but the interface is thread-safe and designed to allow efficient and expressive use in multi-threaded applications. The interface is designed for extreme scalability throughout, and deliberately avoids design features that could inhibit scalability.

97 MATHEMATICS AND COMPUTING↗

UPC++ v1.0 Specification, Revision 2023.3.0

UPC++ is a C++ library providing classes and functions that support Partitioned Global Address Space (PGAS) programming. The key communication facilities in UPC++ are one-sided Remote Memory Access (RMA) and Remote Procedure Call (RPC). All communication operations are syntactically explicit and default to non-blocking; asynchrony is managed through the use of futures, promises and continuation callbacks, enabling the programmer to construct a graph of operations to execute asynchronously as high-latency dependencies are satisfied. A global pointer abstraction provides system-wide addressability of shared memory, including host and accelerator memories. The parallelism model is primarily process-based, but the interface is thread-safe and designed to allow efficient and expressive use in multi-threaded applications. The interface is designed for extreme scalability throughout, and deliberately avoids design features that could inhibit scalability.

97 MATHEMATICS AND COMPUTING↗