Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Distributed System and Computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Determining Levels of Detail for Simulators of Parallel and Distributed Computing Systems via Automated Calibration

There are two sources of inaccuracy when simulating parallel and distributed computing systems: (i) a simulator implemented at an insufficient level of detail; and (ii) incorrectly calibrated simulation parameter values. Increasing the simulator’s level of detail can improve accuracy, but at the cost of higher space, time, and/or software complexity. Furthermore, evaluating the intrinsic accuracy of a simulator requires that its parameters be well-calibrated. Making decisions regarding the level of detail is thus challenging. We propose a methodology for instantiating the simulation calibration process and a framework for automating this process, which makes it possible to pick appropriate levels of detail for any simulator. We demonstrate the usefulness of our approach via two case studies for two different domains.

McDonald, Jessie [University of Hawaii at Manoa, H↗

Position Papers for the ASCR Workshop on Cybersecurity and Privacy for Scientific Computing Ecosystems

At the request of the Department of Energy's (DOE) Office of Advanced Scientific Computing Research (ASCR), this program committee has been tasked with organizing a workshop to identify basic research needs in cybersecurity and privacy to better support DOE's science and energy mission. As part of the process, the program committee is soliciting community input in the form of position papers to help identify significant use cases, facility issues, and other barriers to enabling verifiably trustworthy computational science while preserving data confidentiality as appropriate for scientific workflows of interest to DOE. The program committee will review these position papers and based on the fit of their area of expertise and interest, selected contributors will have the opportunity to participate in the workshop currently planned as a virtual event November 3-5th, 2021. The thrust areas that will be explored by this workshop are the following: (1) Algorithms for secure, scalable, privacy-enhancing technologies and frameworks, including: Federated AI/ML, Differential privacy, Randomized algorithms, Adversarial modeling & simulation, Graph algorithms, and Formal methods; (2) Platforms to support the entire scientific-computing ecosystem, including edge computing for large-scale experiments, focusing on heterogeneous systems and distributed systems, including: Heterogeneous computing systems, Distributed computing systems, and Secure data architectures; and (3) Data workflows to allow agile use of data while preserving integrity and privacy, making the important properties verifiable either at runtime or post-computation, including: Integrity and provenance and Data management infrastructure. Topics that are out-of-scope for the workshop include discussing specific proposed solutions or areas that are clearly out of DOE's fundamental and applied-sciences mission scope, e.g., cryptography, enterprise security, and general-operations technology.

97 MATHEMATICS AND COMPUTING↗

DGaaS: GPU as a Service on Distributed Computing System

In the rapidly evolving landscape of scientific computing, Graphics Processing Units (GPUs) have become indispensable for their unparalleled ability to handle parallel tasks in complex calculations, simulations, and data analysis. Their utility is further magnified in machine learning and AI applications, where they significantly accelerate model training and predictive analytics. Within this context, the Triton Inference Server emerges as a pivotal open-source tool, specializing in AI inferencing and optimizing GPU utilization across various platforms and frameworks. This paper presents an in-depth study on distributed High Throughput Computing (HTC), specifically focusing on the HTCondor framework and its resource provisioning tools, GlideinWMS and HEPCloud. These systems enable large-scale scientific experiments like CMS and DUNE to efficiently access and utilize vast computational resources. The paper explores the core architectural components of GlideinWMS, including jobs, user pools, and worker nodes, and discusses their integration with GPUs and the Triton server. The primary aim of this research is to develop a solution that optimizes GPU utilization by leveraging Glideins and containers. This approach allows computational jobs, particularly those involving AI models, to use GPUs only when essential, thereby facilitating efficient sharing of limited GPU resources. To validate this architecture, the study conducted three key tests involving custom scripts, container-based servers, and Triton server deployments. However, the study faces challenges, notably in locating the Triton server and ensuring secure remote access. To address these issues, future work will focus on developing a proxy mechanism and enhancing security protocols. In conclusion, this study offers a comprehensive roadmap for effective and efficient GPU utilization in distributed High Throughput Computing. It aims to contribute significantly to the scientific community by solving pressing problems and implementing robust solutions in collaboration with the GlideinWMS and HEPCloud teams. The research sets the stage for a more efficient, scalable, and cost-effective paradigm in scientific computing.

97 MATHEMATICS AND COMPUTING↗

RuralAI in Tomato Farming: Integrated Sensor System, Distributed Computing, and Hierarchical Federated Learning for Crop Health Monitoring

Precision horticulture is evolving due to scalable sensor deployment and machine learning (ML) integration. These advancements boost the operational efficiency of individual farms, balancing the benefits of analytics with autonomy requirements. However, given concerns that affect wide geographic regions (e.g., climate change), there is a need to apply models that span farms. Federated learning (FL) has emerged as a potential solution. FL enables decentralized ML across different farms without sharing private data. Traditional FL assumes simple two-tier network topologies and, thus, falls short of operating on more complex networks found in real-world agricultural scenarios. Networks vary across crops and farms and encompass various sensor data modes, extending across jurisdictions. New hierarchical FL (HFL) approaches are needed for more efficient and context-sensitive model sharing, accommodating regulations across multiple jurisdictions. Here, we present the RuralAI architecture deployment for tomato crop monitoring, featuring sensor field units for soil, crop, and weather data collection. HFL with personalization is used to offer localized and adaptive insights. Model management, aggregation, and transfers are facilitated via a flexible approach, enabling seamless communication between local devices, edge nodes, and the cloud.

60 APPLIED LIFE SCIENCES↗

BigPanDA monitoring system evolution in the ATLAS Experiment

Monitoring services play a crucial role in the day-to-day operation of distributed computing systems. The ATLAS Experiment at LHC uses the Production and Distributed Analysis workload management system (PanDA WMS), which allows a million computational jobs to run daily at over 170 computing centers of the WLCG and opportunistic resources, utilizing 600k cores simultaneously on average. The BigPanDA monitor is an essential part of the monitoring infrastructure for the ATLAS Experiment that provides a wide range of views, from top-level summaries to a single computational job and its logs. Over the past few years of the PanDA WMS advancement in the ATLAS Experiment, several new components were developed, such as Harvester, iDDS, Data Carousel, and Global Shares. Due to its modular architecture, the BigPanDA monitor naturally grew into a platform where the relevant data from all PanDA WMS components and accompanying services are accumulated and displayed in the form of interactive charts and tables. Moreover the system has been adopted by other experiments beyond HEP. In this paper we describe the evolution of the BigPanDA monitor system, the development of new modules, and the integration process into other experiments.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

New Results on Communication- and Memory-Aware Load Balancing Model and Algorithms

While load balancing in distributed-memory computing has been well-studied, we present an innovative approach to this problem: a unified, reduced-order model that combines three key components to describe “work” in a distributed system: computation, communication, and memory. Our model enables an optimizer to explore complex tradeoffs in task placement, such as augmented parallelism, at the expense of data replication increasing memory usage. We propose a fully distributed, heuristic-based load balancing optimization algorithm, and demonstrate that it quickly finds close-to-optimal solutions. We formalize the complex optimization problem as a mixed-integer linear program, and compare it to our strategy. Finally, we show that when applied to an electromagnetics code, our approach obtains up to 2.3x speedups for the imbalanced execution.

97 MATHEMATICS AND COMPUTING↗

Software and computing for Run 3 of the ATLAS experiment at the LHC

The ATLAS experiment has developed extensive software and distributed computing systems for Run 3 of the LHC. These systems are described in detail, including software infrastructure and workflows, distributed data and workload management, database infrastructure, and validation. The use of these systems to prepare the data for physics analysis and assess its quality are described, along with the software tools used for data analysis itself. An outlook for the development of these projects towards Run 4 is also provided.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

DIRAC current, upcoming and planned capabilities and technologies

DIRAC is the interware for building and operating large scale distributed computing systems. It is adopted by multiple collaborations from various scientific domains for implementing their computing models. DIRAC provides a framework and a rich set of ready-to-use services for Workload, Data and Production Management tasks of small, medium and large scientific communities having different computing requirements. The base functionality can be easily extended by custom components supporting community specific workflows. DIRAC is at the same time an aging project, and a new DiracX project is taking shape for replacing DIRAC in the long term. This contribution will highlight DIRAC’s current, upcoming and planned capabilities and technologies, and how the transition to DiracX will take place. Examples include, but are not limited to, adoption of security tokens and interactions with Identity Provider services, integration of Clouds and High Performance Computers, interface with Rucio, improved monitoring and deployment procedures.

97 MATHEMATICS AND COMPUTING↗

Systems and methods for distributed power system model calibration

A computing device for distributed power system model calibration is provided. The computing device is programmed to receive event data and model response data associated with a model to simulate, wherein the model includes a plurality of parameters, divide the event data into a plurality of sets, wherein each set includes associated parameters, and transmit the plurality of sets of event data to a plurality of client nodes. Each client node of the plurality of client nodes is programmed to analyze a corresponding set of event data to determine updated parameters for the model. The computing device is further programmed to receive a plurality of updated parameters for the model from the plurality of client nodes and analyze the received plurality of updated parameters to determine at least one adjusted parameter.

Wang, Honggang↗

A Reference Implementation for a Quantum Message Passing Interface

Practical applications of quantum computing are currently limited by the number of qubits that can be set with reasonable fidelities for each system. Therefore, a distributed quantum computing system with multiple quantum computers coherently connected is highly demanding. To realize the internode communication of quantum information, the software interface, Quantum Message Passing Interface (QMPI), leveraging the framework built for classical MPI but taking advantage of quantum teleportation to communicate between different quantum nodes was proposed. In this project, we develop the QMPI with point-to-point and collective operations in Qiskit and characterize its performance by demonstrating the application implementations. Moreover, we developed a new technique for optimizing collective communication of the distributed quantum programs with Multi-Controlled Toffoli gates. This technique beats the state-of-the-art in terms of fidelity and the number of remote EPR pairs consumed in both simulations and experiments.

Shi, Yue↗

Obfuscation for high-performance computing systems

An example method includes initializing, by an obfuscation computing system, communications with nodes in a distributed computing platform, the nodes including one or more compute nodes and a controller node, and performing at least one of: (a) code-level obfuscation for the distributed computing platform to obfuscate interactions between an external user computing system and the nodes, wherein performing the code-level obfuscation comprises obfuscating data associated with one or more commands provided by the user computing system and sending one or more obfuscated commands to at least one of the nodes in the distributed computing platform; or (b) system-level obfuscation for the distributed computing platform, wherein performing the system-level obfuscation comprises at least one of obfuscating system management tasks that are performed to manage the nodes or obfuscating network traffic data that is exchanged between the nodes.

Powers, Judson↗

A Sparse Distributed Gigascale Resolution Material Point Method

In this paper, we present a four-layer distributed simulation system and its adaptation to the Material Point Method (MPM). The system is built upon a performance portable C++ programming model targeting major High-Performance-Computing (HPC) platforms. A key ingredient of our system is a hierarchical block-tile-cell sparse grid data structure that is distributable to an arbitrary number of Message Passing Interface (MPI) ranks. We additionally propose strategies for efficient dynamic load balance optimization to maximize the efficiency of MPI tasks. Our simulation pipeline can easily switch among backend programming models, including OpenMP and CUDA, and can be effortlessly dispatched onto supercomputers and the cloud. Finally, we construct benchmark experiments and ablation studies on supercomputers and consumer workstations in a local network to evaluate the scalability and load balancing criteria. We demonstrate massively parallel, highly scalable, and gigascale resolution MPM simulations of up to 1.01 billion particles for less than 323.25 seconds per frame with 8 OpenSSH-connected workstations.

97 MATHEMATICS AND COMPUTING↗

Extremely Scalable Distributed Computation of Contour Trees via Pre-Simplification

Contour trees offer an abstract representation of the level set topology in scalar fields and are widely used in topological data analysis and visualization. However, applying contour trees to large-scale scientific datasets remains challenging due to scalability limitations. Recent developments in distributed hierarchical contour trees have addressed these challenges by enabling scalable computation across distributed systems. Building on these structures, advanced analytical tasks—such as volumetric branch decomposition and contour extraction—have been introduced to facilitate large-scale scientific analysis. Despite these advancements, such analytical tasks substantially increase memory usage, which hampers scalability. In this paper, we propose a pre-simplification strategy to significantly reduce the memory overhead associated with analytical tasks on distributed hierarchical contour trees. We demonstrate enhanced scalability through strong scaling experiments, constructing the largest known contour tree—comprising over half a trillion nodes with complex topology—in under 15 minutes on a dataset containing 550 billion elements.

Li, Mingzhe [University of Utah]↗

System, method, and computer-accessible medium for remote sensing of the electrical distribution grid with hypertemporal imaging

An exemplary system, method, and computer-accessible medium for determining a property(ies) regarding an electrical grid(s) can be provided, which can include, for example, receiving a video(s) of the electrical grid(s), determining a flicker(s) in the electrical grid(s) based on the video(s), and determining the property(ies) based on the flicker(s). The flicker(s) can be a 120 Hertz flicker. The flicker(s) can be a flicker in a light(s) recorded in the video(s). A frequency and a phase of the flicker(s) can be determined.

Bianco, Federica B.↗

Quantum-Inspired Power System Reliability Assessment

To enable an in-depth study of power system operation and planning, the assessment of standard reliability indices is inevitable. The Monte Carlo Simulation (MCS) approach is a broadly used method in replacing the analytical methods in reliability indices assessment. The accuracy of MCS, however, highly depends on the sampling size, and hence, a complicated system with large number of components requires a large sampling size and daunting computational effort. To address this shortcoming, we, in this paper attempt to take advantage of potentials of the quantum computing (QC) for power system reliability assessment by realizing the following contributions: 1) an innovative quantum model designed for reliability assessment; 2) a quantum circuit that achieves the quadratic speed up compared to the classical MCS method; 3) an efficient quantum amplitude estimation (QAE) algorithm to accurately evaluate the reliability indices. The accuracy and efficacy of the quantum reliability method are extensively verified and demonstrated on both radial and mesh distribution systems.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗