Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel systems”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Magnetized laser–plasma interactions in high-energy-density systems: Parallel propagation

In this study, we investigate parametric processes in magnetized plasmas, driven by a large-amplitude pump light wave. Our focus is on laser–plasma interactions relevant to high-energy-density (HED) systems, such as the National Ignition Facility and the Sandia MagLIF concept. We present a self-contained derivation of a “parametric” dispersion relation for magnetized three-wave interactions, meaning the pump wave is included in the equilibrium, similar to the unmagnetized work of Drake et al., Phys. Fluids 17, 778 (1974). For this, we use a multi-species plasma fluid model and Maxwell's equations. The application of an external B field causes right- and left-polarized light waves to propagate with differing phase velocities. This leads to Faraday rotation of the polarization, which can be significant in HED conditions. Phase-matching and linear wave dispersion relations show that Raman and Brillouin scattering have modified spectra due to the background B field, though this effect is usually small in systems of current practical interest. We study a scattering process we call stimulated whistler scattering, where a light wave decays to an electromagnetic whistler wave (ω≲ω ce ) and a Langmuir wave. This only occurs in the presence of an external B field, which is required for the whistler wave to exist.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Formal Definitions and Performance Comparison of Consistency Models for Parallel File Systems

The semantics of HPC storage systems are defined by the consistency models to which they abide. Storage consistency models have been less studied than their counterparts in memory systems, with the exception of the POSIX standard and its strict consistency model. The use of POSIX consistency imposes a performance penalty that becomes more significant as the scale of parallel file systems increases and the access time to storage devices, such as node-local solid storage devices, decreases. While some efforts have been made to adopt relaxed storage consistency models, these models are often defined informally and ambiguously as by-products of a particular implementation. Here in this work, we establish a connection between memory consistency models and storage consistency models and revisit the key design choices of storage consistency models from a high-level perspective. Further, we propose a formal and unified framework for defining storage consistency models and a layered implementation that can be used to easily evaluate their relative performance for different I/O workloads. Finally, we conduct a comprehensive performance comparison of two relaxed consistency models on a range of commonly seen parallel I/O workloads, such as checkpoint/restart of scientific applications and random reads of deep learning applications. We demonstrate that for certain I/O scenarios, a weaker consistency model can significantly improve the I/O performance. For instance, in small random reads that are typically found in deep learning applications, session consistency achieved a 5x improvement in I/O bandwidth compared to commit consistency, even at small scales.

97 MATHEMATICS AND COMPUTING↗

Hardware-in-the-Loop Investigation of Emissions Challenges in Hybrid Medium- and Heavy-Duty Powertrains Using a Pre-Production Diesel-Electric Parallel Hybrid System With and Without Stop-Start Operation

Hybrid electric powertrains are a growing market in medium- and heavy-duty applications. There is a lack of available information to understand the challenges in the integration of engine platforms into electrified powertrains, such as cold-start, restart, and load-reduction effects on emissions and emission control devices. Results from the Heavy Heavy-Duty Diesel Truck (HHDDT) cycle using a conventional medium-duty diesel engine were compared with those of a parallel hybrid architecture. Oak Ridge National Laboratory in collaboration with the US Department of Energy and Odyne Systems, LLC developed a powertrain in a hardware-in-the-loop environment, integrating the Odyne Systems, LLC medium-duty parallel hybrid system, which was used for the hybrid portion of this study. Experiments under the HHDDT cycle showed increasing improvements in fuel consumption and engine-out emissions with the integration of stop/start, hybrid, and hybrid with stop/start. However, the effects of load reduction and exhaust temperature on the thermal management strategy have shown an increase in fueling in the second part of the HHDDT cycle. Four configurations of medium-duty electrification were studied and contributed to building a unique data set containing combustion, emissions, and system integration data. Each electrification level was compared with the conventional baseline. The calibration of the conventional engine was not altered for this study. Opportunities to tailor the combustion process were identified with the stop/start strategy.

Lerin, Chloe↗

Collection of Disk Failure Events from Alpine, the Parallel File System for Summit Supercomputer

This dataset contains disk (HDD) failure events collected from the Alpine storage system of the Summit supercomputer, hosted at OLCF, spanning from January 4, 2019, to December 21, 2023 (a total of 4 years, 11 months, and 18 days), covering 89% of its operational lifetime. It includes 3,766 disk failure events, each recorded with its detection timestamp (in ISO 8601 format) and detailed by its location within the storage system - rack, enclosure, and drive slot number.

97 MATHEMATICS AND COMPUTING↗

TriC: Distributed-memory Triangle Counting by Exploiting the Graph Structure

Graph analytics has emerged as an important tool in the analysis of large scale data from diverse application domains such as social networks, cyber security and bioinformatics. Counting the number of triangles in a graph is a fundamental kernel with several applications such as detecting the community structure of a graph or in identifying important vertices in a graph. The ubiquity of massive datasets is driving the need to scale graph analytics on parallel systems. However, numerous challenges exist in efficiently parallelizing graph algorithms, especially on distributed-memory systems. Irregular memory accesses and communication patterns, low computation to communication ratios, and the need for frequent synchronization are some of the leading challenges. In this paper, we present TriC, our distributed-memory implementation of triangle counting in graphs using the Message Passing Interface (MPI), as a submission to the 2020 GraphChallenge competition. Using a set of synthetic and real-world inputs from the challenge, we demonstrate a speedup of up to 90x relative to previous work on 32 processor-cores of a NERSC Cori node. We also provide details from distributed runs with up to8192 processes along with strong scaling results. The observations presented in this work provide an understanding of the system-level bottlenecks at scale that specifically impact sparse-irregular workloads and will therefore benefit other efforts to parallelize graph algorithms.

Halappanavar, Mahantesh↗

Paralleling of LLC Resonant Converters

The LLC resonant converter is a popular, variable switching frequency DC-DC converter that may be controlled using two methods: charge and frequency control. In this paper, the application of LLC resonant converters to input-parallel, output-parallel system is studied. In this respect, the models of output-port I-V characteristics and small-signal output impedance of the charge controlled LLC converter are proposed. In addition, a mathematical framework is developed for droop-based paralleled DC-DC systems. Here, it distinctly identifies the output DC voltage and circulating current modes of stability, even in systems comprising of non-identical converters.

42 ENGINEERING↗

Using a Grid-Forming Inverter to Stabilize a Low-Inertia Power System - Maui Hawaiian Island

As power systems around the world integrate greater amounts of wind and solar photovoltaic power, periods of very high instantaneous power shares of inverters, the primary interfacing technology for these generation sources, are complicating system stability and control. The contemporary, primary mode of inverter operation, grid-following, which explicitly assumes the presence of a local, stable voltage waveform, yields operational inadequacy at high instantaneous power shares potentially leading to instability due to the low-inertia conditions, as well as the correlated reduced voltage forming capacity on the respective system. Parallel connected grid-forming inverters, which directly regulate the local voltage, are a solution that is expected to bolster system stability and mitigate the shortcomings of the grid-following technology. In this paper, grid-following and two types of grid-forming inverter control, the traditional linear droop and the recently introduced nonlinear exponential droop (Droop-e), are simulated on a low-inertia, H = 0.48s, high inverter-based resource scenario, 97%, with a validated electromagentic transient domain model of the Hawaiian island of Maui power system. The benefit of a single grid-forming device over its grid-following counterpart is significant, both in terms frequency deviation and voltage stability. Further, the superiority of the Droop-e and the associated secondary power sharing control over both grid-following and linear droop grid-forming technologies is displayed, with improved nadir and rate of change of frequency over the linear droop control.

droop control↗

Design Considerations for GPU-based Mixed Integer Programming on Parallel Computing Platforms

Mixed Integer Programming (MIP) is a powerful abstraction in combinatorial optimization that finds real-life application across many significant sectors. The recent proliferation of graphical processing unit (GPU)-based accelerated computing architectures in large-scale parallel computing or supercomputing presents new opportunities as well as challenges in the advancement of MIP solver technology to effectively use the new accelerated computing platforms and scale to large parallel systems. Here, we recount the conventional processor-based strategies and focus on configurations where the most promising intersection lies between parallel MIP solver approaches and the specific strengths of accelerated parallel platforms. We note that the best potential lies in solving problems whose individual matrix sizes (of the linear program relaxation) fit entirely within one accelerator's memory and whose branch-and-bound (or branch-and-cut) trees cannot be fully contained within a small number of computational nodes. Additionally, we identify ideal features of computational linear algebra support on GPU accelerators that would help advance this direction of scalable parallel solution of MIP problems on GPU-based accelerated computing architectures.

Perumalla, Kalyan↗

Invited Paper: Benchmarking and Optimizing Data Movement on Emerging Heterogeneous Architectures

As supercomputers evolve, nodes are continually increasing in complexity. As a result, each generation of parallel systems brings new performance challenges. For instance, on recent systems inter-node communication has outperformed inter-socket, resulting in poor performance of many node-aware communication optimizations. Communication optimizations are critical for the performance and scalability of parallel applications, but are dependent on the parallel architecture, which varies significantly among recent generations of supercomputers. Furthermore, this paper investigates the performance of various paths of data movement on recent generations of systems, and analyzes the increased complexity of communication, particularly on recent heterogeneous systems. The paper also introduces MPI Advance, a communication library that enables optimizations to be created based on benchmark analysis of each emerging system.

benchmarking↗

Design-time performance modeling of compositional parallel programs

Performance models are powerful instruments for understanding the performance of parallel systems and uncovering their bottlenecks. Already during system design, performance models can help ponder alternative development options. However, creating a performance model – whether theoretically or empirically – for an entire application that does not exist yet is challenging. In this paper, we propose to generate performance models of full programs from performance models of their components using formal composition operators derived from parallel design patterns. As long as the design of the overall system follows such a pattern, its performance model can be predicted with reasonable accuracy without an actual implementation. In conclusion, we demonstrate our approach with design patterns of varying complexity, including pipeline, task pool, and eventually MapReduce, which is representative of a broad class of data-analytics applications.

97 MATHEMATICS AND COMPUTING↗

Towards Lightweight Data Integration Using Multi-Workflow Provenance and Data Observability

Modern large-scale scientific discovery requires multidisciplinary collaboration across diverse computing facilities, including High Performance Computing (HPC) machines and the Edge-to-Cloud continuum. Integrated data analysis plays a crucial role in scientific discovery, especially in the current AI era, by enabling Responsible AI development, FAIR, Reproducibility, and User Steering. However, the heterogeneous nature of science poses challenges such as dealing with multiple supporting tools, cross-facility environments, and efficient HPC execution. Building on data observability, adapter system design, and provenance, we propose MIDA: an approach for lightweight runtime Multi-workflow Integrated Data Analysis. MIDA defines data observability strategies and adaptability methods for various parallel systems and machine learning tools. With observability, it intercepts the dataflows in the background without requiring instrumentation while integrating domain, provenance, and telemetry data at runtime into a unified database ready for user steering queries. We conduct experiments showing end-to-end multi-workflow analysis integrating data from Dask and MLFlow in a real distributed deep learning use case for materials science that runs on multiple environments with up to 276 GPUs in parallel. We show near-zero overhead running up to 100,000 tasks on 1,680 CPU cores on the Summit supercomputer.

Santos Souza, Renan↗