Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “asynchronous”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Optimizing neural networks

A system and method design and optimize neural networks. The system and method include a data store that stores a plurality of gene vectors that represent diverse and distinct neural networks and an evaluation queue stored with the plurality of gene vectors. Secondary nodes construct, train, and evaluate the neural network and automatically render a plurality of fitness values asynchronously. A primary node executes a gene amplification on a select plurality of gene vectors, a crossing-over of the amplified gene vectors, and a mutation of the crossing-over gene vectors automatically and asynchronously, which are then transmitted to the evaluation queue. The process continuously repeats itself by processing the gene vectors inserted into the evaluation queue until a fitness level is reached, a network's accuracy level plateaus, a processing time period expires, or when some stopping condition or performance metric is met or exceeded.

97 MATHEMATICS AND COMPUTING↗

Data exchange and processing synchronization in distributed systems

Systems, methods, techniques and apparatuses of asynchronous communication is distributed systems are disclosed. One exemplary embodiment is a method determining, with a plurality of agent nodes structured to communicate asynchronously in a distributed system, a first set of iterations including an iteration determined by each of the plurality of agent nodes; determining, with a first agent node of the plurality of agent nodes, a local vector clock; receiving, with the first agent node, a first iteration of the first set of iterations and a remote vector clock determined based on the first iteration; updating, with the first agent node, the local vector clock based on the received remote vector clock; and determining a first iteration of a second set of iterations based on the first set of iterations after determining all iterations of the first set of iterations have been received based on the local vector clock.

Cintuglu, Mehmet H.↗

SVM-Based Synchronized Fault Detection for 100% Renewable Microgrids: Preprint

Traditional protection schemes face significant challenges when applied to microgrids with high penetrations of renewables with inverter-based resources (IBRs). The proliferation of advanced sensing and communication technologies has generated copious data, offering an opportunity to overcome these limitations using data-driven machine learning approaches. This work proposes a novel approach based on a support vector machine (SVM) for detecting faults within a 100% renewable microgrid. The approach encompasses a systematic offline training stage for the development of a linear SVM-based fault detection algorithm. This process covers offline data collection from the microgrid under study, the extraction of features such as positive- and negative-sequence components and the total harmonic distortion of the voltage and current measurements of the relays, and the design of the linear SVM-based classifier. During the online implementation, however, different classifiers can exhibit asynchronicity in detecting the fault inception at different subcycle-to-cycle period-level delays. To circumvent this asynchronicity issue, a separate algorithm is developed for each relay to estimate the fault inception time as close to the real fault time. The performance of the proposed SVM-based synchronized fault detection method is evaluated using online time-domain simulation studies on a microgrid test system. The results corroborate the reliability of the fault detection scheme when tested under various fault cases (fault types, locations, and impedances) and non-fault cases during both grid-tied and islanded operation modes.

100% microgrid↗

Stage-local partitioned two-step runge-kutta methods for large systems of ordinary differential equations

We introduce stage-local partitioned two-step Runge-Kutta methods are an extension of standard two-step Runge-Kutta methods, which are an alternative to the standard additive two-step Runge-Kutta methods currently existing in the literature. Furthermore, these new schemes are designed with an eye towards truly N-partitioned systems and leverage local stage approximations to make several computationally interesting approximations viable. Specifically, the focus on local stage approximations makes possible the construction of truly asynchronous schemes, in the parallel sense, possible. In addition, we show that an implicit-explicit approach to these schemes can lead to methods that require the inversion of only local nonlinear systems.

Applied Dynamical Systems↗

Efficient distributed continual learning for steering experiments in real-time

Deep learning has emerged as a powerful method for extracting valuable information from large volumes of data. However, when new training data arrives continuously (i.e., is not fully available from the beginning), incremental training suffers from catastrophic forgetting (i.e., new patterns are reinforced at the expense of previously acquired knowledge). Training from scratch each time new training data becomes available would result in extremely long training times and massive data accumulation. Rehearsal-based continual learning has shown promise for addressing the catastrophic forgetting challenge, but research to date has not addressed performance and scalability. To fill this gap, we propose an approach based on a distributed rehearsal buffer that efficiently complements data-parallel training on multiple GPUs to achieve high accuracy, short runtime, and scalability. It leverages a set of buffers (local to each GPU) and uses several asynchronous techniques for updating these local buffers in an embarrassingly parallel fashion, all while handling the communication overheads necessary to augment input minibatches using unbiased, global sampling. We further propose a generalization of rehearsal buffers to support both classification and generative learning tasks, as well as more advanced rehearsal strategies (notably Dark Experience Replay, leveraging knowledge distillation). We illustrate this approach with a real-life HPC streaming application from the domain of ptychographic image reconstruction. Furthermore, we run extensive experiments on up to 128 GPUs of the ThetaGPU supercomputer to compare our approach with baselines representative of training-from-scratch (the upper bound in terms of accuracy) and incremental training (the lower bound). Results show that rehearsal-based continual learning achieves a top-5 validation accuracy close to the upper bound, while simultaneously exhibiting a runtime close to the lower bound.

Asynchronous data management↗

Structured Adaptive Mesh Refinement Adaptations to Retain Performance Portability With Increasing Heterogeneity

Adaptive mesh refinement (AMR) is an important method that enables many mesh-based applications to run at effectively higher resolution within limited computing resources by allowing high resolution only where really needed. This advantage comes at a cost, however: greater complexity in the mesh management machinery and challenges with load distribution. With the current trend of increasing heterogeneity in hardware architecture, AMR presents an orthogonal axis of complexity. Additionally, the usual techniques, such as asynchronous communication and hierarchy management for parallelism and memory that are necessary to obtain reasonable performance are very challenging to reason about with AMR. Different groups working with AMR are bringing different approaches to this challenge. Here, we examine the design choices of several AMR codes and also the degree to which demands placed on them by their users influence these choices.

42 ENGINEERING↗

Decentralized Schemes with Overlap for Solving Graph-Structured Optimization Problems

We present a new algorithmic paradigm for the decentralized solution of graph-structured optimization problems that arise in the estimation and control of network systems. A key and novel design concept of the proposed approach is that it uses overlapping subdomains to promote and accelerate convergence. We show that the algorithm converges if the size of the overlap is sufficiently large and that the convergence rate improves exponentially with the size of the overlap. The proposed approach provides a bridge between fully decentralized and centralized architectures and is flexible in that it enables the implementation of asynchronous schemes, handling of constraints, and balancing of computing, communication, and data privacy needs. The proposed scheme is tested in an estimation problem for a 9241-node power network and we show that it outperforms the alternating direction method of multipliers.

asynchronous↗

Approach for energy efficient building design during early phase of design process

Energy consumption in the building sector is about 40% of total energy consumed globally and is trending upwards, along with its contribution to greenhouse gas (GHG) emissions. Given the adverse impacts of GHG emissions, it is crucial to integrate energy efficiency into building designs. The most significant opportunities for enhancing energy performance are present during the initial phases of building design, when there is less impact of other design constraints. Various tools exist for simulating different design options and providing feedback in terms of energy consumption and comfort parameters. These simulation outputs must then be analyzed to derive design solutions. This paper presents an innovative approach that utilizes user input parameters, processes them through cloud computing, and outputs easily understandable strategies for energy-efficient building design. The methodology employs Asynchronous Distributed Task Queues (DTQ) - a more scalable and reliable alternative to conventional speedup techniques-for conducting parametric energy simulations in the cloud. The goal of this approach is to assist design teams in identifying, visualizing, and prioritizing energy-saving design strategies from a range of possible solutions for each project. Furthermore, a tool ‘eDOT’ has been developed utilizing the discussed methodology. Unlike existing tools, eDOT leverages artificial intelligence to dynamically generate and provide design strategies during the early phases of design process. By simplifying the simulation process, eDOT enables design teams to make informed, data-driven decisions without needing to interpret complex simulation outputs. A case study simulated for two locations is provided in this paper to demonstrate the effectiveness of eDOT, further underscoring its practical impact on energy-efficient building design.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Minimizing Inter‐Lattice Strain to Stabilize Li‐Rich Cathode by Order–Disorder Control

Li-rich Mn-based layered (LMR) cathodes with anionic redox chemistry show great potential for next-generation sustainable Li-ion battery (LIB) applications due to the low cost and high energy density. However, the asynchronous structural evolutions with cycling in the heterogeneous composite structure of LMR lead to serious lattice strain and thus fast electrochemical decay, which hinders the commercialization of LMR cathodes. Here, in this study, an order–disorder coherent LMR cathode is demonstrated that exhibits a higher average voltage (by 0.25 V), negligible voltage decay (97.6% voltage retention after 100 cycles at 100 mA g −1 ), and enhanced cycling stability (98% capacity retention after 200 cycles at 100 mA g −1 ) compared to its layered oxide counterparts. It is proposed that this order–disorder coherent structure design can promote a more synchronous and homogeneous structure evolution during charge and discharge, thus minimizing lattice strain, which significantly prevents layer collapse and collective degradation at high voltage, improving the electrochemical stability. The study displays the feasibility of optimizing the performance of Li-rich cathode materials through a dedicated order–disorder structure control for sustainable energy storage.

Xu, Shenyang [Peking University, Shenzhen (China)]↗

Enabling Modular Autonomous Feedback‐Loops in Materials Science through Hierarchical Experimental Laboratory Automation and Orchestration

Abstract Materials acceleration platforms (MAPs) operate on the paradigm of integrating combinatorial synthesis, high‐throughput characterization, automatic analysis, and machine learning. Within a MAP, one or multiple autonomous feedback loops may aim to optimize materials for certain functional properties or to generate new insights. The scope of a given experiment campaign is defined by the range of experiment and analysis actions that are integrated into the experiment framework. Herein, the authors present a method for integrating many actions within a hierarchical experimental laboratory automation and orchestration (HELAO) framework. They demonstrate the capability of orchestrating distributed research instruments that can incorporate data from experiments, simulations, and databases. HELAO interfaces laboratory hardware and software distributed across several computers and operating systems for executing experiments, data analysis, provenance tracking, and autonomous planning. Parallelization is an effective approach for accelerating knowledge generation provided that multiple instruments can be effectively coordinated, which the authors demonstrate with parallel electrochemistry experiments orchestrated by HELAO. Efficient implementation of autonomous research strategies requires device sharing, asynchronous multithreading, and full integration of data management in experimental orchestration, which to the best of the authors’ knowledge, is demonstrated for the first time herein.

36 MATERIALS SCIENCE↗

Unique Features of Polarization in Ferroelectric Ionic Conductors

Abstract Ferroelectrics that are also ionic conductors offer possibilities for novel applications with high tunability, especially if the same atomic species causes both phenomena. In particular, at temperatures just below the Curie temperature, polarized states may be sustainable as the mobile species is driven in a controlled way over the energy barrier that governs ionic conduction, resulting in unique control of the polarization. This possibility is recently demonstrated in CuInP 2 S 6 , a layered ferroelectric ionic conductor in which Cu ions cause both ferroelectricity and ionic conduction. Here, it is shown that the commonly used approach to calculate the polarization of evolving atomic configurations in ferroelectrics using the modern theory of polarization, namely concerted (synchronous) migration of the displacing ions, is not well suited to describe the polarization evolution as the Cu ions cross the van der Waals gaps. An asynchronous Cu‐migration scheme is introduced, which reflects the physical process by which Cu ions migrate, resolves the difficulties, and describes the polarization evolution both for normal ferroelectric switching and for transitions across the van der Waals gaps, providing a single framework to discuss ferroelectric ionic conductors.

O'Hara, Andrew↗

PCET‐Driven Reactivity of Neptunyl(VI) Yields Oxo‐Bridged Np(V) and Np(IV) Species

Two unconventional polynuclear complexes of neptunium (Np) featuring mono-mathematical equation -oxo motifs have been accessed by proton-coupled electron transfer (PCET) reactivity involving the dissolution of neptunyl(VI) diacetate dihydrate (NpO 2 (OAc) 2 (H 2 O) 2 ∙ HOAc) in methanol followed by addition of a pentadentate Schiff-base ligand. One complex is a mixed-valent [Np V ,Np IV , Np V ] trimer with two bridging mathematical equation μ 2 -oxos and the other is a [Np V , Np V ] dimer featuring a single mathematical equation μ 2 -oxo. In both complexes the outer Np centers are capped with terminal oxo ligands. Spectroscopic and spectrokinetic studies aimed at elucidating mechanistic details of complex formation in this system show that intermediate multinuclear [Np V O 2 (OAc)] n species form prior to metal chelation by the ligand; electrolysis experiments demonstrate that production of Np(V) gives rise to asynchronous proton transfer that does not occur otherwise (in the Np(VI) state) as well as condensation with loss of H 2 O and formation of the polynuclear complexes. We attribute the oxo-deficient nature of these products, with respect to conventional actinyl ([AnO 2 ] m+ ) species, to the reduction/condensation reaction sequence of PCET.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Interfacing HDF5 with a scalable object‐centric storage system on hierarchical storage

Summary Object storage technologies that take advantage of multitier storage on HPC systems are emerging. However, to use these technologies at present, applications have to be modified significantly from current I/O libraries. HDF5, a widely used I/O middleware on HPC systems, provides a virtual object layer (VOL) that allows applications to connect to different storage mechanisms transparently without requiring significant code modifications. We recently designed the proactive data containers (PDC) object‐centric storage system that provides the capabilities of transparent, asynchronous, and autonomous data movement taking advantage of multiple storage tiers—a decision that has so far been left upon the user on most current systems. To enable PDC's features through HDF5 without modifying application codes, we have developed an HDF5 VOL connector that interfaces with PDC. We present in this article the connector interface and evaluate its performance on Cori, a Cray XC40 supercomputer located at the National Energy Research Scientific Computing Center (NERSC). Our evaluation demonstrates up to an 8× improvement compared with HDF5 that has the most recent optimizations.

Mu, Jingqing↗

h5bench: A unified benchmark suite for evaluating HDF5 I/O performance on pre‐exascale platforms

Summary Parallel I/O is a critical technique for moving data between compute and storage subsystems of supercomputers. With massive amounts of data produced or consumed by compute nodes, high‐performant parallel I/O is essential. I/O benchmarks play an important role in this process; however, there is a scarcity of I/O benchmarks representative of current workloads on HPC systems. Toward creating representative I/O kernels from real‐world applications, we have created h5bench , a set of I/O kernels that exercise hierarchical data format version 5 (HDF5) I/O on parallel file systems in numerous dimensions. Our focus on HDF5 is due to the parallel I/O library's heavy usage in various scientific applications running on supercomputing systems. The various tests benchmarked in the h5bench suite include I/O operations (read and write), data locality (arrays of basic data types and arrays of structures), array dimensionality (one‐dimensional arrays, two‐dimensional meshes, three‐dimensional cubes), I/O modes (synchronous and asynchronous). In this paper, we present the observed performance of h5bench executed along several of these dimensions on existing supercomputers (Cori and Summit) and pre‐exascale platforms (Perlmutter, Theta, and Polaris). h5bench measurements can be used to identify performance bottlenecks and their root causes and evaluate I/O optimizations. As the I/O patterns of h5bench are diverse and capture the I/O behaviors of various HPC applications, this study will be helpful to the broader supercomputing and I/O community.

97 MATHEMATICS AND COMPUTING↗

The 2025 “Hacking Limnology” Workshop Series and DSOS Virtual Summit: A Half Decade of Data‐Intensive Aquatic Science

The 5th Aquatic Ecosystem MOdeling Network—Junior (AEMON-J) “Hacking Limnology” Workshop and 6th Virtual Summit: Incorporating Data Science and Open Science in the Aquatic Sciences (DSOS) convened 21–25 July 2025. As in previous years (Fig. 1; Meyer and Zwart 2020; Meyer et al. 2021b, 2021c, 2022, 2024), the virtual workshops and summit were free of charge, the content was formatted to allow for broad engagement from a globally distributed audience, and workshop materials and recordings were made available on the AEMON-J/DSOS archive (Meyer et al. 2021a). In contrast to previous years, which primarily focused on inland aquatic ecosystems, this year's workshops and summit showcased a notable plurality of ecosystem types, with workshops spanning marine, riverine, and lacustrine environments. The weeklong event brought together researchers and practitioners interested in the nexus of data science, open science, and the aquatic sciences, hosting between 47 and 65 attendees at a single time and a higher number of registrants (n = 389), who might opt to access the material asynchronously.

Meyer, Michael F. [US Geological Survey, Portland,↗

LC-MEMENTO: A Memory Model for Accelerated Architectures

With the advent of heterogeneous architectures, in particular, with the ubiquity of multi-GPU systems, it is becoming increasingly important to manage device memory efficiently in order to reap the benefits of the additional core count. To date, such responsibility mainly falls on the programmer where device-to-host data communication (and vice versa), if not done properly, may incur costly memory transfer operations and synchronization. The problem may be compounded by additional requirement to maintain system-wide memory consistency that may involve expensive synchronization overhead. In this paper, we present Location Consistency Memory Model for Enhanced Transfer Operations (LC-MEMENTO). This framework considers incorporating runtime techniques for multi-GPU memory management to support relaxed synchronization semantics and memory transfer operations automatically. Specifically, we implement a relaxed form of a memory consistency model based on the Location Consistency (LC) in an Asynchronous Many-Task Runtime (ARTS) and demonstrate that, this memory model enables additional optimization opportunities for the three representative applications encompassing different computational patterns (scientific computation, graphs, data streaming, etc.).

Memory Models, Accelerators, Adaptive Optimization↗

Making Uintah Performance Portable for Department of Energy Exascale Testbeds

To help ease ports to forthcoming Department of Energy (DOE) exascale systems, testbeds have been made available to select users. These testbeds are helpful for preparing codes to run on the same hardware and similar software as in their respective exascale systems. This paper describes how the Uintah Computational Framework, an open-source asynchronous many-task (AMT) runtime system, has been modified to be performance portable across the DOE Crusher, DOE Polaris, and DOE Sunspot testbeds in preparation for portable simulations across the exascale DOE Frontier and DOE Aurora systems. The Crusher, Polaris, and Sunspot testbeds feature the AMD MI250X, NVIDIA A100, and Intel PVC GPUs, respectively. This performance portability has been made possible by extending Uintah’s intermediate portability layer [18] to additionally support the Kokkos::HIP, Kokkos::OpenMPTarget, and Kokkos::SYCL back-ends. This paper also describes notable updates to Uintah’s support for Kokkos, which were required to make this extension possible. Results are shown for a challenging radiative heat transfer calculation, central to the University of Utah’s predictive boiler simulations. These results demonstrate single-source portability across AMD-, NVIDIA-, and Intel-based GPUs using various Kokkos back-ends.

Holmen, John↗

Rethinking Programming Paradigms in the QC-HPC Context

Programming for today’s quantum computers is making significant strides toward modern workflows compatible with high performance computing (HPC), but fundamental challenges still remain in the integration of these vastly different technologies. Quantum computing (QC) programming languages share some common ground, as well as their emerging runtimes and algorithmic modalities. In this short paper, we explore avenues of refinement for the quantum processing unit (QPU) in the context of many-tasks management, asynchronous or otherwise, in order to understand the value it can play in linking QC with HPC. Through examples, we illustrate how its potential for scientific discovery might be realized.

Wong, Elaine↗