Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Distributed Computing Resources”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Resource distribution under spatiotemporal uncertainty of disease spread: Stochastic versus robust approaches

We consider the problem of optimizing locations of distribution centers (DCs) and plans for distributing resources such as test kits and vaccines, under spatiotemporal uncertainties of disease spread and demand for the resources. We aim to balance the operational cost (including costs of deploying facilities, shipping, and storage) and quality of service (reflected by demand coverage), while ensuring equity and fairness of resource distribution across multiple populations. We compare a sample-based stochastic programming (SP) approach with a distributionally robust optimization (DRO) approach using a moment-based ambiguity set. Numerical studies are conducted on instances of distributing COVID-19 vaccines in the United States and test kits, to compare SP and DRO models with a deterministic formulation using estimated demand and with the current resource distribution plans implemented in the US. We demonstrate the results over distinct phases of the pandemic to estimate the cost and speed of resource distribution depending on scale and coverage, and show the “demand-driven” properties of the SP and DRO solutions. Furthermore, our results further indicate that if the worst-case unmet demand is prioritized, then the DRO approach is preferred despite of its higher overall cost. Nevertheless, the SP approach can provide an intermediate plan under budgetary restrictions without significant compromises in demand coverage.

97 MATHEMATICS AND COMPUTING↗

From Reproducible Edge–Cloud Experimentation to Real-World Practice: The E2Clab Experience

Reproducibility is already difficult in distributed systems; on the computing continuum, it becomes substantially harder. Applications that span sensing devices, edge and fog resources, and cloud platforms must be evaluated across heterogeneous hardware, variable network conditions, cross-layer orchestration decisions, and long-running workflow lifecycles. We use E2Clab as a case study to examine these challenges and their implications for experimental methodology. We explain why reproducible experimentation is harder on the continuum, then revisit E2Clab as an initial response based on explicit modeling of infrastructure, workflow lifecycle, and artifacts. Lastly, we discuss how its evolution toward more realistic application settings can be understood through the lens of Translational Computer Science. We argue that reproducible continuum experimentation requires methods that are rigorous enough for research while remaining adaptable to real-world practice.

42 ENGINEERING↗

Operational experience and R&D results using the Google Cloud for High-Energy Physics in the ATLAS experiment

The ATLAS experiment at CERN relies on a Worldwide Distributed Computing Grid infrastructure to support its physics program at the Large Hadron Collider. ATLAS has integrated cloud computing resources to complement its Grid infrastructure and conducted an R&D program on Google Cloud Platform. These initiatives leverage key features of commercial cloud providers: lightweight configuration and operation, elasticity and availability of diverse infrastructures. Here this paper examines the seamless integration of cloud computing services as a conventional Grid site within the ATLAS workflow management and data management systems, while also offering new setups for interactive, parallel analysis. It underscores pivotal results that enhance the on-site computing model and outlines several R&D projects that have benefited from large-scale, elastic resource provisioning models. Furthermore, this study discusses the impact of cloud-enabled R&D projects in three domains: accelerators and AI/ML, ARM CPUs and columnar data analysis techniques.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Dispatch Manager for NEML2 Constitutive Model Calculations Embedded in MOOSE

This report describes the extended capabilities of the NEML2 constitutive modeling library, including a flexible and efficient work dispatching system designed to leverage both CPU and GPU resources. This enhancement addresses one of the primary computational challenges in large-scale simulations: the ability to distribute and execute batches of material model evaluations across heterogeneous computing devices. The new dispatch system introduces a modular set of dispatcher and scheduler classes that coordinate the flow of data and execution between devices. The dispatcher is responsible for efficiently packaging work, managing device-specific memory operations, and synchronizing results. This modularity allows for extensibility, making it straightforward to integrate additional computing backends in the future. From an implementation standpoint, the dispatcher system interfaces seamlessly with NEML2's existing models. They handle device-aware tensor operations, optimize memory transfers, and support asynchronous execution when applicable. This design ensures that batches of material points can be evaluated concurrently, substantially improving throughput compared to previous single-device or serial implementations. These improvements not only enhance the raw performance of NEML2 but also improve its usability in multiscale and high-fidelity simulations, where the simultaneous evaluation of large material point batches is critical. Benchmarks included in the report demonstrate the system’s scalability, highlighting its effectiveness when leveraging modern GPU architectures.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

CSR in the Probabilistic Resource Adequacy Suite (PRAS): Fast Interregional Transmission Models for Renewable Integration Studies

NREL's Probabilistic Resource Adequacy Suite (PRAS) was developed for use in renewable integration studies, to assess the resource adequacy of large interconnected power systems with potentially significant interregional transfers of power. This presentation provides an overview of the tool and the methods it uses to solve multi-regional transmission and storage dispatch problems in a computationally-efficient manner.

composite system reliability↗

Characterizing, Modeling, and Accurately Simulating Power and Energy Consumption of I/O-intensive Scientific Workflows

While distributed computing infrastructures can provide infrastructure-level techniques for managing energy consumption, application-level energy consumption models have also been developed to support energy-efficient scheduling and resource provisioning algorithms. In this work, we analyze the accuracy of a widely-used application-level model that has been developed and used in the context of scientific workflow executions. To this end, we profile two production scientific workflows on a distributed platform instrumented with power meters. We then conduct an analysis of power and energy consumption measurements. This analysis shows that power consumption is not linearly related to CPU utilization and that I/O operations significantly impact power, and thus energy, consumption. We then propose a power consumption model that accounts for I/O operations, including the impact of waiting for these operations to complete, and for concurrent task executions on multi-socket, multi-core compute nodes. We implement our proposed model as part of a simulator that allows us to draw direct comparisons between real-world and modeled power and energy consumption. Here, we find that our model has high accuracy when compared to real-world executions. Furthermore, our model improves accuracy by about two orders of magnitude when compared to the traditional models used in the energy-efficient workflow scheduling literature.

97 MATHEMATICS AND COMPUTING↗

Modeling Data Flows with Network Calculus in Cyber-Physical Systems: Enabling Feature Analysis for Anomaly Detection Applications

The electric grid is becoming increasingly cyber-physical with the addition of smart technologies, new communication interfaces, and automated grid-support functions. Because of this, it is no longer sufficient to only study the physical system dynamics, but the cyber system must also be monitored as well to examine cyber-physical interactions and effects on the overall system. To address this gap for both operational and security needs, cyber-physical situational awareness is needed to monitor the system to detect any faults or malicious activity. Techniques and models to understand the physical system (the power system operation) exist, but methods to study the cyber system are needed, which can assist in understanding how the network traffic and changes to network conditions affect applications such as data analysis, intrusion detection systems (IDS), and anomaly detection. In this paper, we examine and develop models of data flows in communication networks of cyber-physical systems (CPSs) and explore how network calculus can be utilized to develop those models for CPSs, with a focus on anomaly and intrusion detection. This provides a foundation for methods to examine how changes to behavior in the CPS can be modeled and for investigating cyber effects in CPSs in anomaly detection applications.

97 MATHEMATICS AND COMPUTING↗

A PSCAD Library Component Featuring a Reduced-Order IBR Model for EMT-Based Fault Studies

This paper presents a fully implemented inverter reduce-order-model (ROM) in an EMT simulation (PSCAD) library component for direct user utilization in protection studies. The developed inverter ROM has the following features: Equivalent to a full IBR inverter model with positive- and negative-sequence current formulation and representation A python script is developed to fully automate this process, including training data generation, ROM parameter training, updating parameters, and model verification and validation. With this PSCAD ROM library component, protection engineers can utilize a trustworthy, accurate ROM for protection studies in an easy-to-use and streamlined manner.

24 POWER TRANSMISSION AND DISTRIBUTION↗

An Open-Source Parallel EMT Simulation Framework: Preprint

As the integration level of inverter-based resources (IBR) increases, ensuring the reliable operation of the bulk power systems requires the use of electromagnetic transient (EMT) simulation tools to identify and mitigate system-wide stability risks. Conducting EMT studies for large-scale, IBR-rich grids, however, is challenging due to the inherent computational bottleneck caused by the underlying high-fidelity models and required small time steps. This paper introduces ParaEMT: an open-source, generic EMT simulation framework designed to accelerate simulations by leveraging advanced parallel computational technologies, such as high-performance computers. This paper presents a comprehensive exposition of ParaEMT, covering its modeling library, simulation strategy, framework structure, operational procedures, and auxiliary features, alongside its extensible parallel computational architecture. Notably, ParaEMT is a publicly accessible and modularized framework written in Python, thereby facilitating future development and the integration of new models and algorithms. The accuracy and efficiency of ParaEMT are demonstrated by rigorous validations via multiple case studies.

electromagnetic transient simulation↗

Modifications to Sandia's MDT and WNTR tools for ERMA

ERMA is leveraging Sandia’s Microgrid Design Toolkit (MDT) [1] and adding significant new features to it. Development of the MDT was primarily funded by the Department of Energy, Office of Electricity Microgrid Program with some significant support coming from the U.S. Marine Corps. The MDT is a software program that runs on a Microsoft Windows PC. It is an amalgamation of several other software capabilities developed at Sandia and subsequently specialized for the purpose of microgrid design. The software capabilities include the Technology Management Optimization (TMO) application for optimal trade-space exploration, the Microgrid Performance and Reliability Model (PRM) for simulation of microgrid operations, and the Microgrid Sizing Capability (MSC) for preliminary sizing studies of distributed energy resources in a microgrid.

97 MATHEMATICS AND COMPUTING↗

Neural Networks-Based Inverter Control: Modeling and Adaptive Optimization for Smart Distribution Networks

The optimal voltage control of inverter-based resources, especially under the high penetration of solar photovoltaics, is critical to the stability of the distribution power system. However, the computational complexity as well as the coordinated operation performance of the voltage control optimization in the distribution power system limits the real-time applications. To mitigate this issue, a model-free based adaptive optimal control scheme for the smart inverter is proposed to maximize the active power generation, minimize the power loss, and maintain the bus voltages in smart distribution networks. An inverter-based optimization model for coordinated operation is first established, considering the uncertainties of renewable power generation. Subsequently, by collecting the data and control strategies, the neural networks (NNs) based algorithm is proposed to efficiently predict the best possible control strategy. The main objective of this scheme is to accurately predict candidate optimal solutions with near-negligible feasibility and optimization gaps, with the advantage of avoiding complicated iteration-based numerical algorithms. Thereafter, the co-simulation among OpenDSS, MATLAB, and Python is set up to fully take advantage of the three individual software. Experiments are conducted based on different control parameter characteristics and structures of NNs. Finally, the results reveal that an average mean squared error of 0.013 and 1 ms response time are achieved, which is lower than some state-of-the-art methods.

42 ENGINEERING↗

Neumann Series Based Voltage Sensitivity Analysis for Three Phase Distribution System

In this letter, a simplified voltage sensitivity analysis technique that can provide accurate estimates of voltage change across the network for a given change in bus power injections in a three-phase unbalanced distribution network is proposed. This technique is derived from the first-order approximation of the Neumann series, which allows maintaining the accuracy of the solution while the computational effort is reduced. Here, the proposed technique is tested on a 559-bus unbalanced distribution system with multiple distributed generation resources. The results show that the average error in the voltage estimates with the proposed method is not more than 0.3% with the execution time of similar order relative to the state-of-the-art sensitivity analysis methods.

42 ENGINEERING↗

Multi-Agent Control Planes for Quantum Networks: A Scalable Architecture for Autonomous Quantum Internet Management

Quantum networks are expected to enable distributed quantum computing, secure communication, and global entanglement distribution. However, operating such networks presents significant challenges, including stochastic quantum processes, fragile entanglement resources, dynamic topology, and cross-layer control requirements. Current quantum network control architectures largely rely on centralized or hierarchical controllers inspired by classical software-defined networking (SDN). While effective for small testbeds, these approaches face scalability, latency, and reliability limitations as quantum networks grow. This paper proposes a multi-agent control plane architecture for quantum networks. In this design, intelligent software agents operate at quantum nodes, repeaters, and orchestration layers, collectively managing entanglement generation, routing, purification, and scheduling. The distributed intelligence of the agent system allows the network to adapt dynamically to quantum hardware variability and environmental noise. We argue that multi-agent systems provide significant advantages over centralized control approaches, including scalability, resilience, local autonomy, and real-time adaptation. The paper discusses architectural design principles, agent coordination mechanisms, and research challenges in deploying multi-agent control planes for the emerging quantum Internet.

Alnajjar, Anees [ORNL] (ORCID:0000000237101601)↗

Lowering entry barriers to developing custom simulators of distributed applications and platforms with SimGrid

Researchers in parallel and distributed computing (PDC) often resort to simulation because experiments conducted using a simulator can be for arbitrary experimental scenarios, are less resource-, labor-, and time-consuming than their real-world counterparts, and are perfectly repeatable and observable. Many frameworks have been developed to ease the development of PDC simulators, and these frameworks provide different levels of accuracy, scalability, versatility, extensibility, and usability. Further, the SimGrid framework has been used by many PDC researchers to produce a wide range of simulators for over two decades. Its popularity is due to a large emphasis placed on accuracy, scalability, and versatility, and is in spite of shortcomings in terms of extensibility and usability. Although SimGrid provides sensible simulation models for the common case, it was difficult for users to extend these models to meet domain-specific needs. Furthermore, SimGrid only provided relatively low-level simulation abstractions, making the implementation of a simulator of a complex system a labor-intensive undertaking. In this work we describe developments in the last decade that have contributed to vastly improving extensibility and usability, thus lowering or removing entry barriers for users to develop custom SimGrid simulators.

97 MATHEMATICS AND COMPUTING↗

A Generic and Multifunctional Electromagnetic Transient Model for Grid-Following Inverters

This article presents a generic and multi-functional electromagnetic transient (EMT) dynamic model of grid following (GFL) inverter-based resource (IBR) using the PSCAD software platform. The features of the developed model includes the flexibility in selecting various types and combinations of DC sources covering PV modules, battery modules, ideal DC source module, as well as flexibility in selecting either switched or averaged model of inverter. This model also covers exhaustive lists of controller logic covering open-loop/closed-loop PQ dispatch control, DC voltage and AC terminal voltage control along with the conventional current control designed in dq-domain, ate- domain and positive-negative sequence domain. Moreover, this model is equipped with the flexibility in selecting various types of current limiting schemes that includes saturation-based as well as latching-based current limiter, anti-windup protection. Moreover, the EMT model is agnostic to the MVA rating and is suitable for interfacing transmission systems by being complaint with the IEEE Std. 2800. The generality in the power circuits and the multi-functional options in operation and control of the developed EMT model makes it suitable for both academia and industry to study various power system aspects not limited to but such as fault behavior of GFL IBR and impacts on protection system, transient stability of a system interfaced with large number of GFL IBRs etc.

14 SOLAR ENERGY↗

How to Build a Quantum Supercomputer: Scaling from Hundreds to Millions of Qubits

In the span of four decades, quantum computation has evolved from an intellectual curiosity to a potentially realizable technology. Today, small-scale demonstrations have become possible for quantum algorithmic primitives on hundreds of physical qubits and proof-of-principle error-correction on a single logical qubit. Nevertheless, despite significant progress and excitement, the path toward a full-stack scalable technology is largely unknown. There are significant outstanding quantum hardware, fabrication, software architecture, and algorithmic challenges that are either unresolved or overlooked. These issues could seriously undermine the arrival of utility-scale quantum computers for the foreseeable future. Here, we provide a comprehensive review of these scaling challenges. We show how the road to scaling could be paved by adopting existing semiconductor technology to build much higher-quality qubits, employing system engineering approaches, and performing distributed quantum computation within heterogeneous high-performance computing infrastructures. These opportunities for research and development could unlock certain promising applications, in particular, efficient quantum simulation/learning of quantum data generated by natural or engineered quantum systems. To estimate the true cost of such promises, we provide a detailed resource and sensitivity analysis for classically hard quantum chemistry calculations on surface-code error-corrected quantum computers given current, target, and desired hardware specifications based on superconducting qubits, accounting for a realistic distribution of errors. Furthermore, we argue that, to tackle industry-scale classical optimization and machine learning problems in a cost-effective manner, heterogeneous quantum-probabilistic computing with custom-designed accelerators should be considered as a complementary path toward scalability.

Mohseni, Masoud↗