Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “distributed computation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

RLC4CLR (Reinforcement Learning Controller for Critical Load Restoration Problems)

RLC4CLR demonstrates using a reinforcement learning controller (RLC) to solve a critical load restoration (CLR) problem, which improves the grid resilience after a substation outage event. RLC4CLR consists of two parts. (1) RL environment: This environment encapsulates the CLR problem to be solved and provides interfacing functions to follow the standard OpenAI Gym format. A power system simulator, i.e., OpenDSS, is included to provide the power flow solution. Controller inputs and outputs (RL state and action) as well as the reward are defined in this environment as well. In summary, the RL environment is the problem formulation from which the RL agent can learn. (2) RL training script: The training script enables the RL agent to learn its control policy by interacting with the RL environment. For RL training, an open-sourced RL library, i.e., RLlib, is leveraged which is based on a distributed computing framework (Ray). The training script is designed to be able to be run on both local machine or the NREL HPC system. Other components of RLC4CLR include input data, e.g., grid model (standard IEEE test feeders), and other files used for results analysis.

Zhang, Xiangyu↗

mkite

mkite is a distributed computing platform for materials simulation. mkite is built with the server-client pattern, decoupling production databases from client runners. When used in combination with message brokers, mkite enables any available client to perform calculations without prior hardware specification on the server side. Furthermore, the software enables the creation of complex workflows with multiple inputs and branches, facilitating the exploration of combinatorial chemical spaces. The mkite suite provides recipes and tools to interact with package such as VASP, but is extensible to any other simulation package. Finally, mkite helps keeping the provenance of calculations in a SQL database. A complete description of the software is available at https://arxiv.org/abs/2301.08841.

Schwalbe Koda, Daniel↗

msdlive-cli-distro

MSD-LIVE, the MultiSector Dynamics – Living, Intuitive, Value-adding, Environment, is a flexible and scalable data and code management system combined with a distributed computational platform that will enable MSD researchers to document and archive their data, run their models and analysis tools, and share their data, software, and multi-model workflows within a robust Community of Practice. MSD-LIVE will facilitate a new open, collaborative, resource-rich, technology-facilitated, community-driven way of doing MSD research.

Lansing, Carina↗

torc (Torc Workflow Management System) [SWR-24-127]

This software package orchestrates execution of a workflow of jobs on distributed computing resources. It is optimized for use on HPCs with Slurm, but also can be used in the cloud and on local computers. Please refer to the documentation at https://nrel.github.io/torc

Thom, Daniel [National Renewable Energy Laboratory↗

Cadmus v1.0

This software package is for distributed computation of persistent (co)homology. It is primarily intended for Topological Data Analysis (TDA) audience and scientists who use TDA methods and apply them to datasets that are too big to handle on a single machine. There are very few codes available for this; Cadmus outperforms DIPHA, if the user is interested in cohomology. The accompanying paper explaining the algorithm was accepted to ALENEX 26.

Nigmetov, Arnur [Lawrence Berkeley National Labora↗

Detector Interface for Streaming, Control, and Open-source integration (DISCO) v1.0.0

This suite consists of a multi-package ecosystem featuring detector emulators, EPICS areaDetector drivers, and remote server frameworks designed for the Advanced Light Source (ALS). Engineered for high-bandwidth devices—including VFCCD, Timepix3, Timepix4, and related pixel detectors—the software simulates hardware, wraps vendor SDKs into remote-callable servers, and integrates with open-source control systems. Key Capabilities: Distributed SDK Architecture: Server packages wrap hardware-specific SDKs, allowing areaDetector drivers to execute remote framework calls. This isolates proprietary libraries from the EPICS IOC, enhancing stability and enabling distributed computing across beamline networks. Device Support: Custom drivers for VFCCD, the Timepix family, and similar sensors optimize the data path from hardware control to high-speed transport. Full-Stack Emulation: Sophisticated emulator packages allow end-to-end pipeline testing and software development without requiring physical hardware or beam time. Integrated Workflows: Supports high-bandwidth streaming for real-time analysis and robust, metadata-rich file-based workflows (e.g., HDF5/NeXus). By standardizing interfaces across heterogeneous hardware, this suite reduces technical debt. It provides the ALS with a scalable, open-source solution to manage massive data rates within a unified control environment.

Mahl, Johannes [Lawrence Berkeley National Laborat↗

Designing and prototyping extensions to the Message Passing Interface in MPICH

As HPC system architectures and the applications running on them continue to evolve, the MPI standard itself must evolve. The trend in current and future HPC systems toward powerful nodes with multiple CPU cores and multiple GPU accelerators makes efficient support for hybrid programming critical for applications to achieve high performance. However, the support for hybrid programming in the MPI standard has not kept up with recent trends. The MPICH implementation of MPI provides a platform for implementing and experimenting with new proposals and extensions to fill this gap and to gain valuable experience and feedback before the MPI Forum can consider them for standardization. Here, in this work, we detail six extensions implemented in MPICH to increase MPI interoperability with other runtimes, with a specific focus on heterogeneous architectures. First, the extension to MPI generalized requests lets applications integrate asynchronous tasks into MPI’s progress engine. Second, the iovec extension to datatypes lets applications use MPI datatypes as a general-purpose data layout API beyond just MPI communications. Third, a new MPI object, MPIX_Stream, can be used by applications to identify execution contexts beyond MPI processes, including threads and GPU streams. MPIX stream communicators can be created to make existing MPI functions thread-aware and GPU-aware, thus providing applications with explicit ways to achieve higher performance. Fourth, MPIX Streams are extended to support the enqueue semantics for offloading MPI communications onto a GPU stream context. Fifth, thread communicators allow MPI communicators to be constructed with individual threads, thus providing a new level of interoperability between MPI and on-node runtimes such as OpenMP. Lastly, we present an extension to invoke MPI progress, which lets users spawn progress threads with fine-grained control to adapt the communication performance to their application designs. We describe the design and implementation of these extensions, provide usage examples, and highlight their expected benefits with performance results.

97 MATHEMATICS AND COMPUTING↗

LLM Benchmarking with LLaMA2: Evaluating Code Development Performance Across Multiple Programming Languages

The rapid evolution of large language models (LLMs) has opened new possibilities for automating various tasks in software development. This paper evaluates the capabilities of the LLaMA 2-70B model in automating these tasks for scientific applications written in commonly used programming languages. Using representative test problems, we assess the model's capacity to generate code, documentation, and unit tests, as well as its ability to translate existing code between commonly used programming languages. Our comprehensive analysis evaluates the compilation, runtime behavior, and correctness of the generated and translated code. Additionally, we assess the quality of automatically generated code, documentation, and unit tests. Here, our results indicate that while LLaMA 2-70B frequently generates syntactically correct and functional code for simpler numerical tasks, it encounters substantial difficulties with more complex, parallelized, or distributed computations, requiring considerable manual corrections. We identify key limitations and suggest areas for future improvements to better leverage AI-driven automation in scientific computing workflows.

97 MATHEMATICS AND COMPUTING↗

Exploring OpenSNAPI Use Cases and Evolving Requirements [Slides]

Emerging system architectures are rapidly transforming in order to meet shifting requirements. Motivated by expanding data volumes, energy efficiency concerns, and the omnipresent need to improve performance, architectures are increasingly adopting a data-centric approach. At the core of this concept is the goal of minimizing data motion and instead processing data in-situ to the greatest degree possible. Therefore, data-centric designs, in contrast to conventional CPU-centric models, typically distribute compute capabilities throughout the architecture. As part of this paradigm shift, a novel class of devices known as data processing units (DPUs), alongside CPUs and GPUs, are quickly forming a third pillar of data-centric systems. These devices, which include smart network adapters and switches, seek to offload computation on data at the network edge as well as in-flight within the network fabric. The Open Smart Network API (OpenSNAPI) project seeks to develop a unified API for DPU devices. In our previous talks, we introduced the OpenSNAPI project and detailed our investigations regarding the viability of offloading compute intensive kernels to BlueField DPUs. In contrast, in this talk we detail our efforts to offload application-level file I/O to the DPU. We also discuss plans and early efforts to explore in-network compute capabilities. Finally, we describe our observations with respect to the evolving design of OpenSNAPI.

97 MATHEMATICS AND COMPUTING↗

User Manual - HydraGNN: Distributed PyTorch Implementation of Multi-Headed Graph Convolutional Neural Networks

This document serves as user manual for HydraGNN, a scalable graph neural network (GNN) architecture that allows for a simultaneous prediction of multiple target properties using multi-task learning (MTL). The HydraGNN architecture is constructed by successive superposition of three different sets of layers. The first set is made of message-passing layers to exchange information across nodes in the graph and use this to update the nodal features. The second set is made of global pooling layers that aggregate information from all the nodes in the graph and map it into a scalar, and is needed only for global target properties that are related to the entire graph. The third set of layers is dedicated to the implementation of MTL, which is enabled by forking of the architecture into separate heads, each one of them dedicated to the predictive task of one specific target property. Through an object-oriented programming paradigm, HydraGNN is templated over different message-passing policies, which allows for a user-friendly hyperparameter study to assess the sensitivity of the predictive performance of the HydraGNN architecture on a specific dataset with respect to the choice of the message-passing policy. The object-oriented paradigm used by HydraGNN also allows for a user-friendly inclusion of newly developed message passing policies within the existing framework. HydraGNN supports distributed computing capabilities for scalable data reading and scalable training on leadership-class supercomputers.

97 MATHEMATICS AND COMPUTING↗

Creating Unit Tests for GlideinWMS using AI tools

GlideinWMS is a workload management system that uses distributed computing to complete tasks, also known as jobs. It is particularly useful for high-throughput computing that’s used in research projects. It relies on Glideins, which are pilot jobs that pull jobs from a queue and provide resources for their completion, based on the jobs requirements. These decisions are made based on resource availability and job requirements. We used new AI tools to add unit tests to GlideinWMS.

Baburashvili, Ilya↗

Enabling Innovative Analysis on Heterogeneous Clusters through HTCdaskgateway

High energy particle (HEP) physics research is going through fundamental changes as we move to collect larger amounts of data from the Large Hadron Collider (LHC). Analysis facilities and distributed computing, through HTCs, have come together to create the next pythonic generation of analysis by utilizing HTCdaskgateway, a Dask gateway extension, allowing users to spawn workers compatible with both their analysis and heterogeneous clusters in line with authentication requirements. This is enabling physicists to engage with scientific python in ways they had not before because of domain specific C++ tools. An example of HTCdaskgateway’s use is Fermilab’s Elastic Analysis Facility.

Chavez, Elise [U. Wisconsin, Madison (main)]↗

Drugsniffer: An Open Source Workflow for Virtually Screening Billions of Molecules for Binding Affinity to Protein Targets

The SARS-CoV2 pandemic has highlighted the importance of efficient and effective methods for identification of therapeutic drugs, and in particular has laid bare the need for methods that allow exploration of the full diversity of synthesizable small molecules. While classical high-throughput screening methods may consider up to millions of molecules, virtual screening methods hold the promise of enabling appraisal of billions of candidate molecules, thus expanding the search space while concurrently reducing costs and speeding discovery. Here, we describe a new screening pipeline, called drugsniffer, that is capable of rapidly exploring drug candidates from a library of billions of molecules, and is designed to support distributed computation on cluster and cloud resources. As an example of performance, our pipeline required ~40,000 total compute hours to screen for potential drugs targeting three SARS-CoV2 proteins among a library of ~3.7 billion candidate molecules.

59 BASIC BIOLOGICAL SCIENCES↗

Containerization in ATLAS Software Development and Data Production

The ATLAS experiment's software production and distribution on the grid benefits from a semi-automated infrastructure that provides up-to-date information about software usability and availability through the CVMFS dis-tribution service for all relevant systems. The software development process uses a Continuous Integration pipeline involving testing, validation, packag-ing and installation steps. For opportunistic sites that can not access CVMFS, containerized releases are needed. These standalone containers are currently created manually to support Monte-Carlo data production at such sites. In this paper we will describe an automated procedure for the containerization of AT-LAS software releases in the existing software development infrastructure, its motivation, integration and testing in the distributed computing system.

97 MATHEMATICS AND COMPUTING↗

Integration of Rucio in Belle II

The Belle II experiment, which started taking physics data in April 2019, will multiply the volume of data currently stored on its nearly 30 storage elements worldwide by one order of magnitude to reach about 340 PB of data (raw and Monte Carlo simulation data) by the end of operations. To tackle this massive increase and to manage the data even after the end of the data taking, it was decided to move the Distributed Data Management software from a homegrown piece of software to a widely used Data Management solution in HEP and beyond : Rucio. This contribution describes the work done to integrate Rucio with Belle II distributed computing infrastructure as well as the migration strategy that was successfully performed to ensure a smooth transition.

97 MATHEMATICS AND COMPUTING↗

A Modularized Urban Scale Building Energy Modeling Framework Designed with An Open Mind

In recent years, physics-based building energy modeling (BEM) has started being used to evaluate the performance of buildings in the context of connected communities and on an urban scale to study their aggregated energy use, interactions, and impacts on the energy supply infrastructure and environment. The development of urban-scale BEM solutions needs extensive effort. Existing attempts tend to focus on different aspects of BEM on an urban scale, such as collecting as-built building data from different information sources, integrating geometry modeling with geographic information systems (GISs), representing operational and occupancy profiles, automating workflow, processing and visualizing the results, and conducting large-scale simulations. Urban-scale BEM development would benefit from multi-disciplinary research areas and from an open platform to adopt advancements on data sources and tools. For these purposes, this research proposes a modularized bottom-up model creation and simulation framework that is built on the state-of-the-art BEM tools and can accommodate different building stock data. This framework uses a standardized schema to describe building design and operational characteristics, and it can be instantiated from different building survey datasets with heterogeneous structures. The paper demonstrates how thousands of surveyed buildings from the 2012 U.S. Energy Information Administration’s Commercial Buildings Energy Consumption Survey (CBECS) were one-to-one converted to EnergyPlus models through the schema and the model generation process, then simulated with distributed computing, and their results are summarized.

Lei, Xuechen↗

Editorial: Edge computation and digital distribution networks

The development of digital technologies is penetrating all areas of energy revolution. Based on the in-depth integration of advanced digital technologies, distribution networks are gradually transforming into digital distribution networks (DDNs) with tremendous changes from the structure to the operation mode. DDNs is the digitalized appearance of the physical distribution network, in which ubiquitous connections and massive data are the basic characteristics (Huo et al., 2022). It is an important task to utilize the massive data and propose novel operation modes to construct more efficient and intelligent distribution networks (Jian et al., 2022). Among the advanced digital technologies in DDNs, edge computing has received wide attention (Zhao et al., 2022). It has superior performance in local sensing and intelligent computation, which can effectively relieve huge communication pressure. However, the limited computing resources and the complex computing tasks at the edge side significantly challenge the collaboration of distribution network regulation and advanced digital technologies (Hu et al., 2022). It is necessary to find out proper methods to utilize advanced digital technology to construct DDNs. This Research Topic is organized to introduce the recent progress in the construction, operation and advanced computational methods for DDNs. Finally, seven papers have been accepted, which can be sorted into the following three categories: 1) Evolution and technical features of DDNs, 2) Intelligent operation control of DDNs, 3) Advanced simulation for large-scale DDNs. The three sections below respectively introduce the major research and contributions of the papers covered in each category.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Assessment of Local Observation of Atomic Ordering in Alloys via the Radial Distribution Function: A Computational and Experimental Approach

As a powerful analytical technique, atom probe tomography (APT) has the capacity to acquire the spatial distribution of millions of atoms from a complex sample. However, extracting information at the Ångstrom-scale on atomic ordering remains a challenge due to the limits of the APT experiment and data analysis algorithms. The development of new computational tools enable visualization of the data and aid understanding of the physical phenomena such as disorder of complex crystalline structures. Here, we report progress towards this goal using two steps. We describe a computational approach to evaluate atomic ordering in the crystal structure by generating radial distribution functions (RDF). Atomic ordering is rendered as the Fractional Cumulative Radial Distribution Function (FCRDF) which allows for greater visibility of local compositions at short range in the structure. Further, we accommodate in the analysis additional parameters such as uncertainty in the atomic coordinates and the atomic abundance to ascertain short-range ordering in APT data sets. We applied the FCRDF analysis to synthetic and experimental APT data sets for Ni 3 Al. The ability to observe a signal of atomic ordering consistent with the known L1 2 crystal structure is heavily dependent on spatial uncertainty, irrespective of abundance. Detection of atomic ordering is subject to an upper limit of spatial uncertainty of atoms described with Gaussian distributions with a standard deviation of 1.3 Å. The FCRDF analysis was also applied to the APT data set for a six-component alloy, Al 1.3 CoCrCuFeNi. In this case, we are currently able to visualize elemental segregation at the nanoscale, though unambiguous identification of atomic ordering at the Ångstrom (nearest-neighbor) scale remains a goal.

36 MATERIALS SCIENCE↗