Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Distributed Asynchronous”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

87 records · Page 5

Dispatch Manager for NEML2 Constitutive Model Calculations Embedded in MOOSE

This report describes the extended capabilities of the NEML2 constitutive modeling library, including a flexible and efficient work dispatching system designed to leverage both CPU and GPU resources. This enhancement addresses one of the primary computational challenges in large-scale simulations: the ability to distribute and execute batches of material model evaluations across heterogeneous computing devices. The new dispatch system introduces a modular set of dispatcher and scheduler classes that coordinate the flow of data and execution between devices. The dispatcher is responsible for efficiently packaging work, managing device-specific memory operations, and synchronizing results. This modularity allows for extensibility, making it straightforward to integrate additional computing backends in the future. From an implementation standpoint, the dispatcher system interfaces seamlessly with NEML2's existing models. They handle device-aware tensor operations, optimize memory transfers, and support asynchronous execution when applicable. This design ensures that batches of material points can be evaluated concurrently, substantially improving throughput compared to previous single-device or serial implementations. These improvements not only enhance the raw performance of NEML2 but also improve its usability in multiscale and high-fidelity simulations, where the simultaneous evaluation of large material point batches is critical. Benchmarks included in the report demonstrate the system’s scalability, highlighting its effectiveness when leveraging modern GPU architectures.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

The 2025 “Hacking Limnology” Workshop Series and DSOS Virtual Summit: A Half Decade of Data‐Intensive Aquatic Science

The 5th Aquatic Ecosystem MOdeling Network—Junior (AEMON-J) “Hacking Limnology” Workshop and 6th Virtual Summit: Incorporating Data Science and Open Science in the Aquatic Sciences (DSOS) convened 21–25 July 2025. As in previous years (Fig. 1; Meyer and Zwart 2020; Meyer et al. 2021b, 2021c, 2022, 2024), the virtual workshops and summit were free of charge, the content was formatted to allow for broad engagement from a globally distributed audience, and workshop materials and recordings were made available on the AEMON-J/DSOS archive (Meyer et al. 2021a). In contrast to previous years, which primarily focused on inland aquatic ecosystems, this year's workshops and summit showcased a notable plurality of ecosystem types, with workshops spanning marine, riverine, and lacustrine environments. The weeklong event brought together researchers and practitioners interested in the nexus of data science, open science, and the aquatic sciences, hosting between 47 and 65 attendees at a single time and a higher number of registrants (n = 389), who might opt to access the material asynchronously.

Meyer, Michael F. [US Geological Survey, Portland,↗

DenKv: Addressing Design Trade-offs of Key-value Stores for Scientific Applications

High-performance computing (HPC) facilities have employed flash-based storage tier near to compute nodes to absorb high I/O demand by HPC applications during periodic system-level checkpoints. To accelerate these checkpoints, proxy-based distributed key-value stores (PD-KVS) gained particular attention for their flexibility to support multiple backends and different network configurations. PD-KVS rely internally on monolithic KVS, such as LevelDB or RocksDB, to exploit the KV interface and query support. However, PD-KVS are unaware of the high redundancy factor in checkpoint data, which can be up to GBs to TBs, and therefore, tend to generate high write and space amplification on these storage layers. In this paper, we propose DenKv which is deduplication-extended node-local LSM-tree-based KVS. DenKv employs asynchronous partially inline dedup (APID) and aims to maintain the performance characteristics of LSM-tree-based KVS while reducing the write and space amplification problems. We implemented DenKv atop BlobDB and showed that our proposed solution maintains performance while reducing write amplification up to 2× and space amplification by 4× on average.

Khan, Awais↗

Asynchronous x-ray multiprobe data acquisition for x-ray transient absorption spectroscopy

Laser pump X-ray Transient Absorption (XTA) spectroscopy offers unique insights into photochemical and photophysical phenomena. X-ray Multiprobe data acquisition (XMP DAQ) is a technique that acquires XTA spectra at thousands of pump-probe time delays in a single measurement, producing highly self-consistent XTA spectral dynamics. In this work, we report two new XTA data acquisition techniques that leverage the high performance of XMP DAQ in combination with High Repetition Rate (HRR) laser excitation: HRR-XMP and Asynchronous X-ray Multiprobe (AXMP). HRR-XMP uses a laser repetition rate up to 200 times higher than previous implementations of XMP DAQ and proportionally increases the data collection efficiency at each time delay. This allows HRR-XMP to acquire more high-quality XTA data in less time. AXMP uses a frequency mismatch between the laser and x-ray pulses to acquire XTA data at a flexibly defined set of pump-probe time delays with a spacing down to a few picoseconds. AXMP introduces a novel pump-probe synchronization concept that acquires data in clusters of time delays. Further, the temporally inhomogeneous distribution of acquired data improves the attainable signal statistics at early times, making the AXMP synchronization concept useful for measuring sub-nanosecond dynamics with photon-starved techniques like XTA. In this paper, we demonstrate HRR-XMP and AXMP by measuring the laser-induced spectral dynamics of dilute aqueous solutions of Fe(CN) 6 4₋ and [Fe II (bpy) 3 ] 2+ (bpy: 2,2'-bipyridine), respectively.

47 OTHER INSTRUMENTATION↗

Chapter 9: Impact of Variable Renewable Energy Sources on Bulk Power System Planning and Operations

Wind and solar photovoltaics (PV) have experienced remarkable growth in recent years, with many consequent benefits within and outside of power systems. At the same time, wind and solar PV have unique characteristics relative to the historically dominant dispatchable technologies like coal, gas, and nuclear power plants that have required and will continue to require changes in power system planning and operations. This chapter discusses planning and operational challenges of integrating wind and solar PV into bulk power systems. We first present the key characteristics of wind and solar PV that differentiate it from conventional technologies, such as variable and uncertain electricity generation, asynchronous interconnection to the power system, and near-zero marginal costs. We then link these characteristics to power system planning and operational challenges at low through high wind and solar penetrations. Finally, we discuss near- and long-term solutions to those challenges, such as diversifying the generation mix and wind and solar fleets, improving system flexibility, diversifying ancillary service products, and integrating generation and transmission planning.

bulk power system↗

Traveler: Navigating Task Parallel Traces for Performance Analysis

Understanding the behavior of software in execution is a key step in identifying and fixing performance issues. This is especially important in high performance computing contexts where even minor performance tweaks can translate into large savings in terms of computational resource use. To aid performance analysis, developers may collect an execution trace —a chronological log of program activity during execution. As traces represent the full history, developers can discover a wide array of possibly previously unknown performance issues, making them an important artifact for exploratory performance analysis. However, interactive trace visualization is difficult due to issues of data size and complexity of meaning. Traces represent nanosecond-level events across many parallel processes, meaning the collected data is often large and difficult to explore. The rise of asynchronous task parallel programming paradigms complicates the relation between events and their probable cause. Here, to address these challenges, we conduct a continuing design study in collaboration with high performance computing researchers. We develop diverse and hierarchical ways to navigate and represent execution trace data in support of their trace analysis tasks. Through an iterative design process, we developed Traveler , an integrated visualization platform for task parallel traces. Traveler provides multiple linked interfaces to help navigate trace data from multiple contexts. We evaluate the utility of Traveler through feedback from users and a case study, finding that integrating multiple modes of navigation in our design supported performance analysis tasks and led to the discovery of previously unknown behavior in a distributed array library.

97 MATHEMATICS AND COMPUTING↗

Real Time Phasor Analytics (RTPA) and RTPA-SCR System Strength Online Tool

This presentation showcases the Real-Time Phasor Analytics (RTPA) framework for monitoring inertia and assessing system strength in power grids. RTPA is an open-source tool designed to standardize access to data from Power Management Units (PMUs) and Phasor Data Concentrators (PDCs). It facilitates real-time connectivity to multiple PDCs in accordance with the IEEE C37.118-2 standard and supports asynchronous data stream integration. Additionally, RTPA can simulate a PDC server streaming C37.118-2 data and provides Python bindings for seamless interaction with the framework, eliminating the need for direct Rust programming.

24 POWER TRANSMISSION AND DISTRIBUTION↗

BULKI-Store v0.3.2

BULKI-Store is a distributed object storage system optimized for high-performance computing environments. Built with a Rust core and Python bindings, it efficiently manages scientific and machine learning datasets across HPC clusters. The system employs a client-server architecture with MPI integration, enabling seamless scaling on supercomputers like Perlmutter. BULKI-Store's object-oriented approach provides intuitive data organization with rich metadata support, contrasting with traditional file-based solutions. Key optimizations include selective checkpoint loading, unified checkpoint files, and object chunking for large data transfers. For machine learning workloads, BULKI-Store offers advantages through fine-grained access patterns, dynamic data sharing between training instances, and reduced memory pressure. Memory management features include strategic Python GC calls, minimized data copies, and batch processing capabilities. The system leverages Rayon's thread pool for asynchronous data prefetching and supports multiple CPU architectures (ARM64, x86, AMD, RISC-V). By combining performance optimizations with developer-friendly APIs, BULKI-Store addresses the complex data management challenges of modern HPC applications while maintaining compatibility across heterogeneous computing environments.

Zhang, Wei [Lawrence Berkeley National Laboratory ↗

SVM-Based Synchronized Fault Detection for 100% Renewable Microgrids

Traditional protection schemes face significant challenges when applied to microgrids with high penetrations of renewables with inverter-based resources (IBRs). The proliferation of advanced sensing and communication technologies has generated copious data, offering an opportunity to overcome these limitations using data-driven machine learning approaches. This work proposes a novel approach based on a support vector machine (SVM) for detecting faults within a 100% renewable microgrid. The approach encompasses a systematic offline training stage for the development of a linear SVM-based fault detection algorithm. This process covers offline data collection from the microgrid under study, the extraction of features such as positive- and negative-sequence components and the total harmonic distortion of the voltage and current measurements of the relays, and the design of the linear SVM-based classifier. During the online implementation, however, different classifiers can exhibit asynchronicity in detecting the fault inception at different subcycle-to-cycle period-level delays. To circumvent this asynchronicity issue, a separate algorithm is developed for each relay to estimate the fault inception time as close to the real fault time. The performance of the proposed SVM-based synchronized fault detection method is evaluated using online time-domain simulation studies on a microgrid test system. The results corroborate the reliability of the fault detection scheme when tested under various fault cases (fault types, locations, and impedances) and non-fault cases during both grid-tied and islanded operation modes.

100% microgrid↗

SVM-Based Synchronized Fault Detection for 100% Renewable Microgrids: Preprint

Traditional protection schemes face significant challenges when applied to microgrids with high penetrations of renewables with inverter-based resources (IBRs). The proliferation of advanced sensing and communication technologies has generated copious data, offering an opportunity to overcome these limitations using data-driven machine learning approaches. This work proposes a novel approach based on a support vector machine (SVM) for detecting faults within a 100% renewable microgrid. The approach encompasses a systematic offline training stage for the development of a linear SVM-based fault detection algorithm. This process covers offline data collection from the microgrid under study, the extraction of features such as positive- and negative-sequence components and the total harmonic distortion of the voltage and current measurements of the relays, and the design of the linear SVM-based classifier. During the online implementation, however, different classifiers can exhibit asynchronicity in detecting the fault inception at different subcycle-to-cycle period-level delays. To circumvent this asynchronicity issue, a separate algorithm is developed for each relay to estimate the fault inception time as close to the real fault time. The performance of the proposed SVM-based synchronized fault detection method is evaluated using online time-domain simulation studies on a microgrid test system. The results corroborate the reliability of the fault detection scheme when tested under various fault cases (fault types, locations, and impedances) and non-fault cases during both grid-tied and islanded operation modes.

100% microgrid↗

Hardening DOE R&D Software Tools for Web-based Visualization SBIR Phase I Final Report

Ubiquitous web-based visualization is essential to delivering large-scale data visualization to various stakeholders, from the scientist to the board member. These stakeholders will not tolerate a stalled application or a pop-up window asking them to wait for the processing to complete. They require a responsive and interactive visualization environment with high-quality imagery suitable for detailed analysis and boardroom presentations. At Kitware, Inc., we have accomplished web visualization to this point, leveraging state-of-the-art tools like HTML5, CSS3, SVG, Canvas, and WebGL. Solutions that leverage a combination of these technologies are necessary to handle workloads that vary significantly in data size efficiently. However, it is not always practical to move large data to the web client for visualization. Kitware's ParaView as a Service combines client-side visualization using both distributed processing and remote rendering on big data impractical to move. Existing distributed processing and remote rendering solution's interactivity is below the expectations of web-based applications. Our project examined proposed solutions to the areas outlined above in ParaView as a Service. We have investigated concurrent pipelines, streaming images, progressive rendering, and optimization of algorithms and data movement to address these concerns. For the Phase I project, we completed the proposed work plan. As a result, the project produced three prototypes of essential importance for web visualization and the ParaView as a Service community. We created a simple desktop application for an interactive streamline placement prototype, a web-based interactive streamline placement prototype, and a web-based progressive rendering utilizing raytracing prototype. These prototypes relied on the hardening of emerging software toolkits funded by the Department of Energy (DOE) Advanced Scientific Computing Research (ASCR) program (such as ParaView, VTK-m, and Mochi). We blended these components into web-based visualization prototypes that meet the industry's expectations for interactivity and responsiveness. The Phase I project had four essential focus areas: 1. Develop prototype ParaView as a Service backend server using asynchronous, non-blocking design principles. 2. Develop a prototype web application that uses the ParaView as a Service backend server for remote data visualization. 3. Implement image streaming with encoding/compression and progressive rendering capabilities in the proposed platform. 4. Evaluate the prototype developed and summarize observations, including the challenges and pitfalls of our approach. After our successful completion of Phase I, we are strongly positioned to propose a successful Phase II project.

Geveci, Berk↗

TBAA20: Task-Based Algorithms and Applications

The new challenges posed by Exascale system architectures have resulted in difficulty achieving a desired scalability using traditional distributed ­memory runtimes. Task­-based programming models show promise in addressing these challenges, providing application developers with a productive and performant approach to programming on next generation systems. Empirical studies show that task-based models can overcome load ­balancing issues that are inherent to traditional distributed ­memory runtimes, and that task-­based runtimes perform comparably to those systems when balanced. This panel is designed to explore the advantages of task-­based programming models on modern and future HPC systems from an industry, university, and national lab perspective. It aims at gathering application experts and proponents of these models to present concrete and practical examples of using task­-based runtimes to overcome the challenges posed by Exascale system architectures. This report describes the objectives, activities, and outcomes of the panel TBAA: Task­-Based Algorithms and Applications which was held at the International Conference for High Performance Computing, Networking, Storage, and Analysis (SC 20) on November 18, 2020.

97 MATHEMATICS AND COMPUTING↗

Cybersecurity Enhancement for Multi-Infeed High-Voltage DC Systems

Composed of multiple two-terminal high-voltage DC (HVDC) transmission systems, a multi-infeed HVDC (MIDC) system exchanges massive power among multiple asynchronous AC systems. However, as an intrinsically cyber-physical system, an MIDC system could suffer from cyber-attacks, leading to massive power mismatches in multiple AC systems, and resulting in catastrophic consequences. Since the sequential responses of an MIDC system and interconnected AC systems are in different timescales, this paper first establishes a two-timescale model to evaluate the sequential impacts caused by cyber-attacks. Then, an event-triggered cyber-defense strategy is proposed to enhance the cybersecurity of an MIDC system by mitigating multiple non-simultaneous cyber-attacks. Whenever new cyber-attack events occur, the proposed cyber-defense strategy, which is mathematically modeled as a mixed-integer quadratic programming problem, is executed on-line and updated in an event-triggered manner. Here, simulation results on an MIDC system demonstrate that the low-cost and almost blind cyber-attacks can cause severe frequency deviations, and the proposed strategy can mitigate multi-cyber-attacks effectively.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Designing and prototyping extensions to the Message Passing Interface in MPICH

As HPC system architectures and the applications running on them continue to evolve, the MPI standard itself must evolve. The trend in current and future HPC systems toward powerful nodes with multiple CPU cores and multiple GPU accelerators makes efficient support for hybrid programming critical for applications to achieve high performance. However, the support for hybrid programming in the MPI standard has not kept up with recent trends. The MPICH implementation of MPI provides a platform for implementing and experimenting with new proposals and extensions to fill this gap and to gain valuable experience and feedback before the MPI Forum can consider them for standardization. Here, in this work, we detail six extensions implemented in MPICH to increase MPI interoperability with other runtimes, with a specific focus on heterogeneous architectures. First, the extension to MPI generalized requests lets applications integrate asynchronous tasks into MPI’s progress engine. Second, the iovec extension to datatypes lets applications use MPI datatypes as a general-purpose data layout API beyond just MPI communications. Third, a new MPI object, MPIX_Stream, can be used by applications to identify execution contexts beyond MPI processes, including threads and GPU streams. MPIX stream communicators can be created to make existing MPI functions thread-aware and GPU-aware, thus providing applications with explicit ways to achieve higher performance. Fourth, MPIX Streams are extended to support the enqueue semantics for offloading MPI communications onto a GPU stream context. Fifth, thread communicators allow MPI communicators to be constructed with individual threads, thus providing a new level of interoperability between MPI and on-node runtimes such as OpenMP. Lastly, we present an extension to invoke MPI progress, which lets users spawn progress threads with fine-grained control to adapt the communication performance to their application designs. We describe the design and implementation of these extensions, provide usage examples, and highlight their expected benefits with performance results.

97 MATHEMATICS AND COMPUTING↗

Optimized allocation method of the VSC-MTDC system for frequency regulation reserves considering the cost

In recent years, the interconnection of asynchronous power grids through the VSC-MTDC system has been proposed and extensively studied in light of the potential benefits of economical bulk power exchanges and frequency regulation reserves sharing. This paper proposed an optimized allocation method for sharing frequency regulation reserves among the interconnected power systems and the corresponding frequency regulation control of the VSC-MTDC system under emergency frequency deviation events. Firstly, the frequency regulation reserve classification is proposed. In the classification, the available frequency response capacity reserves of each interconnection are divided into commercial reserves and regular reserves. While the commercial reserves are procured through long-term contracts, the regular reserves are purchased based on market prices of frequency regulation services. Secondly, based on the proposed frequency regulation reserve classification, a novel frequency regulation control is then introduced for the VSC-MTDC system. This control method could minimize the costs of the disturbed power grid for the needed frequency response supports from the other power grids. Simulation verifications are performed on a modified IEEE 39 bus system and a highly reduced power system model representing the North American grids. The simulation verification indicates that the developed frequency regulation control significantly reduced ancillary service costs of the disturbed power grid.

24 POWER TRANSMISSION AND DISTRIBUTION↗