Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Distributed System and Computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Obfuscation for high-performance computing systems

An example technique includes initializing, by an obfuscation computing system, communications with nodes in a distributed computing platform. The nodes include compute nodes that provide resources in the distributed computing platform and a controller node that performs resource management of the resources. The obfuscation computing system serves as an intermediary between the controller node and the compute nodes. The technique further includes outputting an interactive user interface (UI) providing a selection between a first privilege level and a second privilege level, and performing one of: based on the selection being for the first privilege level, a first obfuscation mechanism for the distributed computing platform to obfuscate digital traffic between a user computing system and the nodes, or based on the selection being for the second privilege level, a second obfuscation mechanism for the distributed computing platform to obfuscate digital traffic between the user computing system and the nodes.

Aloisio, Scott↗

System, method, and computer-accessible medium for remote sensing of the electrical distribution grid with hypertemporal imaging

An exemplary system, method, and computer-accessible medium for determining a property(ies) regarding an electrical grid(s) can be provided, which can include, for example, receiving a video(s) of the electrical grid(s), determining a flicker(s) in the electrical grid(s) based on the video(s), and determining the property(ies) based on the flicker(s). The flicker(s) can be a 120 Hertz flicker. The flicker(s) can be a flicker in a light(s) recorded in the video(s). A frequency and a phase of the flicker(s) can be determined.

Bianco, Federica B.↗

Cloud Services Enable Efficient AI-Guided Simulation Workflows across Heterogeneous Resources

Applications which fuse machine learning and simulation are rarely best served by a single computing resource. Highly parallel simulation codes are best deployed on super- computers, while AI tasks used to decide which simulations to perform may be best suited to specialized accelerators. Here we present a Function-as-a-Service (FaaS) system for executing complex, distributed computational campaigns that achieves performance parity with conventional workflow systems without the complexities of secure network connections between compute providers. One innovation enabling high performance is a subsystem that directly moves task data between sites, separate from the cloud-hosted FaaS system used to distribute task instructions. We also introduce a flexible scheduling system that allows us access factor of 2 trade offs between the amount of resources required to solve a problem at each compute site. We anticipate that this system will upgrade multi-site applications from demonstration projects to routine practice in computational science.

Ward, Logan↗

Quantum-Inspired Power System Reliability Assessment

To enable an in-depth study of power system operation and planning, the assessment of standard reliability indices is inevitable. The Monte Carlo Simulation (MCS) approach is a broadly used method in replacing the analytical methods in reliability indices assessment. The accuracy of MCS, however, highly depends on the sampling size, and hence, a complicated system with large number of components requires a large sampling size and daunting computational effort. To address this shortcoming, we, in this paper attempt to take advantage of potentials of the quantum computing (QC) for power system reliability assessment by realizing the following contributions: 1) an innovative quantum model designed for reliability assessment; 2) a quantum circuit that achieves the quadratic speed up compared to the classical MCS method; 3) an efficient quantum amplitude estimation (QAE) algorithm to accurately evaluate the reliability indices. The accuracy and efficacy of the quantum reliability method are extensively verified and demonstrated on both radial and mesh distribution systems.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Design and implementation of dynamic I/O control scheme for large scale distributed file systems

In this paper, we have analyzed the input/output (I/O) activities of Cori, which is a high-performance computing system at the National Energy Research Scientific Computing Center at Lawrence Berkeley National Laboratory. Our analysis results indicate that most users do not adjust storage configurations but rather use the default settings. In addition, owing to the interference from many applications running simultaneously, the performance varies based on the system status. To configure file systems autonomously in complex environments, we developed DCA-IO, a dynamic distributed file system configuration adjustment algorithm that utilizes the system log information to adjust storage configurations automatically. Our scheme aims to improve the application performance and avoid interference from other applications without user intervention. Moreover, DCA-IO uses the existing system logs and does not require code modifications, an additional library, or user intervention. To demonstrate the effectiveness of DCA-IO, we performed experiments using I/O kernels of real applications in both an isolated small-sized Lustre environment and Cori. Our experimental results shows that our scheme can improve the performance of HPC applications by up to 263% with the default Lustre configuration.

97 MATHEMATICS AND COMPUTING↗

Scalable Circuit Cutting and Scheduling in a Resource-constrained and Distributed Quantum System

Despite quantum computing's rapid development, current systems remain limited in practical applications due to their limited qubit count and quality. Various technologies, such as superconducting, trapped ions, and neutral atom quantum computing technologies are progressing towards a fault tolerant era, however they all face a diverse set of challenges in scalability and control. Recent efforts have focused on multi-node quantum systems that connect multiple smaller quantum devices to execute larger circuits. Future demonstrations hope to use quantum channels to couple systems, however current demonstrations can leverage classical communication with circuit cutting techniques. This involves cutting large circuits into smaller subcircuits and reconstructing them post-execution. However, existing cutting methods are hindered by lengthy search times as the number of qubits and gates increases. Additionally, they often fail to effectively utilize the resources of various worker configurations in a multi-node system. To address these challenges, we introduce FitCut, a novel approach that transforms quantum circuits into weighted graphs and utilizes a community-based, bottom-up approach to cut circuits according to resource constraints, e.g., qubit counts, on each worker. FitCut also includes a scheduling algorithm that optimizes resource utilization across workers. Implemented with Qiskit and evaluated extensively, FitCut significantly outperforms the Qiskit Circuit Knitting Toolbox, reducing time costs by factors ranging from 3 to 2000 and improving resource utilization rates by up to 3.88 times on the worker side, achieving a system-wide improvement of 2.86 times.

Kan, Shuwen [Fordham University]↗

Design and implementation of I/O performance prediction scheme on HPC systems through large-scale log analysis

Abstract Large-scale high performance computing (HPC) systems typically consist of many thousands of CPUs and storage units used by hundreds to thousands of users simultaneously. Applications from large numbers of users have diverse characteristics, such as varying computation, communication, memory, and I/O intensity. A good understanding of the performance characteristics of each user application is important for job scheduling and resource provisioning. Among these performance characteristics, I/O performance is becoming increasingly important as data sizes rapidly increase and large-scale applications, such as simulation and model training, are widely adopted. However, predicting I/O performance is difficult because I/O systems are shared among all users and involve many layers of software and hardware stack, including the application, network interconnect, operating system, file system, and storage devices. Furthermore, updates to these layers and changes in system management policy can significantly alter the I/O behavior of applications and the entire system. To improve the prediction of the I/O performance on HPC systems, we propose integrating information from several different system logs and developing a regression-based approach to predict the I/O performance. Our proposed scheme can dynamically select the most relevant features from the log entries using various feature selection algorithms and scoring functions, and can automatically select the regression algorithm with the best accuracy for the prediction task. The evaluation results show that our proposed scheme can predict the write performance with up to 90% prediction accuracy and the read performance with up to 99% prediction accuracy using the real logs from the Cori supercomputer system at NERSC.

97 MATHEMATICS AND COMPUTING↗

Neural Networks-Based Inverter Control: Modeling and Adaptive Optimization for Smart Distribution Networks

The optimal voltage control of inverter-based resources, especially under the high penetration of solar photovoltaics, is critical to the stability of the distribution power system. However, the computational complexity as well as the coordinated operation performance of the voltage control optimization in the distribution power system limits the real-time applications. To mitigate this issue, a model-free based adaptive optimal control scheme for the smart inverter is proposed to maximize the active power generation, minimize the power loss, and maintain the bus voltages in smart distribution networks. An inverter-based optimization model for coordinated operation is first established, considering the uncertainties of renewable power generation. Subsequently, by collecting the data and control strategies, the neural networks (NNs) based algorithm is proposed to efficiently predict the best possible control strategy. The main objective of this scheme is to accurately predict candidate optimal solutions with near-negligible feasibility and optimization gaps, with the advantage of avoiding complicated iteration-based numerical algorithms. Thereafter, the co-simulation among OpenDSS, MATLAB, and Python is set up to fully take advantage of the three individual software. Experiments are conducted based on different control parameter characteristics and structures of NNs. Finally, the results reveal that an average mean squared error of 0.013 and 1 ms response time are achieved, which is lower than some state-of-the-art methods.

42 ENGINEERING↗

From Reproducible Edge–Cloud Experimentation to Real-World Practice: The E2Clab Experience

Reproducibility is already difficult in distributed systems; on the computing continuum, it becomes substantially harder. Applications that span sensing devices, edge and fog resources, and cloud platforms must be evaluated across heterogeneous hardware, variable network conditions, cross-layer orchestration decisions, and long-running workflow lifecycles. We use E2Clab as a case study to examine these challenges and their implications for experimental methodology. We explain why reproducible experimentation is harder on the continuum, then revisit E2Clab as an initial response based on explicit modeling of infrastructure, workflow lifecycle, and artifacts. Lastly, we discuss how its evolution toward more realistic application settings can be understood through the lens of Translational Computer Science. We argue that reproducible continuum experimentation requires methods that are rigorous enough for research while remaining adaptable to real-world practice.

42 ENGINEERING↗

Creating Unit Tests for GlideinWMS using AI tools

GlideinWMS is a workload management system that uses distributed computing to complete tasks, also known as jobs. It is particularly useful for high-throughput computing that’s used in research projects. It relies on Glideins, which are pilot jobs that pull jobs from a queue and provide resources for their completion, based on the jobs requirements. These decisions are made based on resource availability and job requirements. We used new AI tools to add unit tests to GlideinWMS.

Baburashvili, Ilya↗

Custom Accessors: Enabling Scalable Data Ingestion, (Re-)Organization, and Analysis on Distributed Systems

The emerging class of high velocity and high volume data analytic workflows comprise interwoven data ingestion, organization, and processing stages, with ingestion and organization steps often contributing comparable or even higher computational costs than actual processing steps. Since complex workflows consist of a variety of phases that view and use data differently, being able to construct efficient, scalable, distributed data structures (arrays, vectors, sets, maps, and multi-maps) is essential and requires custom methods to extend and shrink containers, analyze and position data, and, maintain globallyconsistent meta-data. In this paper, we propose a novel datastructure access paradigm based on the concept of Accessors. At a high level, accessors are customizable callable objects that can modify the behavior of insert, read, update, and delete operations for distributed containers while preserving atomicity guarantees. Accessors provide a very clean and natural way to implement a variety of programming patterns, e.g., conditional insertion/deletion and cascading computations, which would be otherwise hard (or even impossible) to express in parallel and distributed settings without using locks. We demonstrate the practicality and usefulness of our approach with two representative use cases and study the performance of these applications on a distributed High-Performance Computing system. Our analysis highlights that our proposed abstraction allows for an effective overlapping and concurrent execution of different workflow steps (e.g., data ingestion and analysis), which in a conventional analytics pipeline would execute sequentially, contributing cumulatively to the overall latency.

Castellana, Vito G. [BATTELLE (PACIFIC NW LAB)] (O↗

Mass Transport in Membrane Systems: Flow Regime Identification by Fourier Analysis

The numerical calculation of local mass distributions in membrane systems by computational fluid dynamics (CFD) offers indispensable benefits. However, the concept to calculate such distributions in response to separate variations of operation conditions (OCs) makes it difficult to address overall, flow-physics-related questions, which require the consideration of the collective interaction of OCs. It is shown that such understanding-related relationships can be obtained by the analytical solution of the advection–diffusion equation considered. A Fourier series model (FSM) is presented, which provides exact solutions of an advection–diffusion equation for a wide range of OCs. On this basis, a new zeroth-order model is developed, which is very simple and as accurate as the complete FSM for all conditions of practical relevance. Advection-dominated blocked and diffusion-dominated unblocked flow regimes are identified (depending on a Péclet number which compares the flow geometry with a length scale imposed by the flow), which implies relevant requirements for the use of lab results for pilot- and full-scale applications. Analyses reveal the equivalence of variations of OCs, which offers a variety of options to accomplish desired flow regime changes.

Heinz, Stefan (ORCID:0000000248712416)↗

Establishing metrics to quantify spatial similarity in spherical and red blood cell distributions

As computational power increases and systems with millions of red blood cells can be simulated, it is important to note that varying spatial distributions of cells may affect simulation outcomes. Since a single simulation may not represent the ensemble behavior, many different configurations may need to be sampled to adequately assess the entire collection of potential cell arrangements. In order to determine both the number of distributions needed and which ones to run, we must first establish methods to identify well-generated, randomly placed cell distributions and to quantify distinct cell configurations. We utilize metrics to assess (1) the presence of any underlying structure to the initial cell distribution and (2) similarity between cell configurations. We propose the use of the radial distribution function to identify long-range structure in a cell configuration and apply it to a randomly distributed and structured set of red blood cells. To quantify spatial similarity between two configurations, we make use of the Jaccard index, and characterize sets of red blood cell and sphere initializations. As an extension to our work submitted to the International Conference on Computational Science, we significantly increase our data set size from 72 to 1048 cells, include a similar set of studies using spheres, compare the effects of varying sphere size, and utilize the Jaccard index distribution to probe sets of extremely similar configurations. Our results show that the radial distribution function can be used as a metric to determine long-range structure in both distributions of spheres and RBCs. We determine that the ideal case of spheres within a cube versus bi-concave shaped cells within a cylinder affects the shape of the Jaccard index distributions, as well as the range of Jaccard values, showing that both the shape of particle and the domain may play a role. Furthermore, we also find that the distribution is able to capture very similar configurations through Jaccard index values greater than 95% when appending several nearly identical configurations into the data set.

59 BASIC BIOLOGICAL SCIENCES↗

Recommendations for Distributed Energy Resource Patching

While computer systems, software applications, and operational technology (OT)/Industrial Control System (ICS) devices are regularly updated through automated and manual processes, there are several unique challenges associated with distributed energy resource (DER) patching. Millions of DER devices from dozens of vendors have been deployed in home, corporate, and utility network environments that may or may not be internet-connected. These devices make up a growing portion of the electric power critical infrastructure system and are expected to operate for decades. During that operational period, it is anticipated that critical and noncritical firmware patches will be regularly created to improve DER functional capabilities or repair security deficiencies in the equipment. The SunSpec/Sandia DER Cybersecurity Workgroup created a Patching Subgroup to investigate appropriate recommendations for the DER patching, holding fortnightly meetings for more than nine months. The group focused on DER equipment, but the observations and recommendations contained in this report also apply to DERMS tools and other OT equipment used in the end-to-end DER communication environment. The group found there were many standards and guides that discuss firmware lifecycles, patch and asset management, and code-signing implementations, but did not singularly cover the needs of the DER industry. This report collates best practices from these standards organizations and establishes a set of best practices that may be used as a basis for future national or international patching guides or standards.

97 MATHEMATICS AND COMPUTING↗

S-QGPU: Shared quantum gate processing unit for distributed quantum computing

We propose a distributed quantum computing (DQC) architecture in which individual small-sized quantum computers are connected to a shared quantum gate processing unit (S-QGPU). The S-QGPU comprises a collection of hybrid two-qubit gate modules for remote gate operations. In contrast to conventional DQC systems, where each quantum computer is equipped with dedicated communication qubits, S-QGPU effectively pools the resources (e.g., the communication qubits) together for remote gate operations, and, thus, significantly reduces the cost of not only the local quantum computers but also the overall distributed system. Our preliminary analysis and simulation show that S-QGPU's shared resources for remote gate operations enable efficient resource utilization. When not all computing qubits (also called data qubits) in the system require simultaneous remote gate operations, S-QGPU-based DQC architecture demands fewer communication qubits, further decreasing the overall cost. Alternatively, with the same number of communication qubits, it can support a larger number of simultaneous remote gate operations more efficiently, especially when these operations occur in a burst mode.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Visualizing Fault Induced Traveling Waves in Medium Voltage Systems: Preprint

Traveling waves are induced in power systems during most transient events in the grid. These waves travel close to the speed of light in overhead lines and 50% to 60% the speed of light in underground cables. Even though traveling wave-based protection schemes for transmission systems are available commercially, traveling waves in medium voltage distribution networks are still in research space. Compared to transmission system, medium voltage distribution systems contain more reflections and refractions. Thus, visualization is challenging and is critical in locating faults in distribution network. To address this visualization challenge, this paper presents an open-source tool to visualize the traveling waves using Bewley lattice approach. The developed visualization tool will be useful for the protection and safety engineers to detect and triangulate fault locations in the medium voltage systems and isolate the faults.

Bewley Lattice↗