Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Multi-Agent”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

162 records · Page 9

Reinforcement Learning to Enhance Optimal Operation of Resilient Community Energy Systems

This paper presents a novel model-free multi-agent Reinforcement Learning (RL) control method to enhance the resilience of community energy systems in island mode, which coordinates multiple objectives without the necessity of identifying system models that require expert knowledge. Specifically, a community-level coordinator agent is designed to allocate renewable energy resources among different buildings, and multiple building-level agents are developed to optimize load schedules based on limited energy resources and requirements of building loads and occupants’ comfort. In a two-day evaluation, our RL approach demonstrated a similar performance against MPC without requiring system models and formulation of optimization problems as required in MPC.

ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATION↗

Towards Agentic AI on Particle Accelerators

As particle accelerators grow in complexity, traditional control methods face increasing challenges in achieving optimal performance. This paper envisions a paradigm shift: a decentralized multi-agent framework for accelerator control, powered by Large Language Models (LLMs) and distributed among autonomous agents. We present a proposition of a self-improving decentralized system where intelligent agents handle high-level tasks and communication and each agent is specialized control individual accelerator components. This approach raises some questions: What are the future applications of AI in particle accelerators? How can we implement an autonomous complex system such as a particle accelerator where agents gradually improve through experience and human feedback? What are the implications of integrating a human-in-the-loop component for labeling operational data and providing expert guidance? We show two examples, where we demonstrate viability of such architecture.

43 PARTICLE ACCELERATORS↗

LC-Opt: Benchmarking Reinforcement Learning and Agentic AI for End-to-End Liquid Cooling Optimization in Data Centers

Liquid cooling is critical for thermal management in high-density data centers with the rising AI workloads. However, machine learning-based controllers are essential to unlock greater energy efficiency and reliability, promoting sustainability. We present LC-Opt, a Sustainable Liquid Cooling (LC) benchmark environment, for reinforcement learning (RL) control strategies in energy-efficient liquid cooling of high-performance computing (HPC) systems. Built on the baseline of a high-fidelity digital twin of Oak Ridge National Lab's Frontier Supercomputer cooling system, LC-Opt provides detailed Modelica-based end-to-end models spanning site-level cooling towers to data center cabinets and server blade groups. RL agents optimize critical thermal controls like liquid supply temperature, flow rate, and granular valve actuation at the IT cabinet level, as well as cooling tower (CT) setpoints through a Gymnasium interface, with dynamic changes in workloads. This environment creates a multi-objective real-time optimization challenge balancing local thermal regulation and global energy efficiency, and also supports additional components like a heat recovery unit (HRU). We benchmark centralized and decentralized multi-agent RL approaches, demonstrate policy distillation into decision and regression trees for interpretable control, and explore LLM-based methods that explain control actions in natural language through an agentic mesh architecture designed to foster user trust and simplify system management. LC-Opt democratizes access to detailed, customizable liquid cooling models, enabling the ML community, operators, and vendors to develop sustainable data center liquid cooling control solutions.

Naug, Avisek [Hewlett Packard Enterprise]↗

A Novel LDPP-MADDPG Approach for Distributed Power Allocation in mmWave Cellular Networks

This paper considers the problem of distributed beam scheduling and power allocation problem in millimeter- Wave (mmWave) cellular networks, in which multiple Base Stations (BSs) operate as individual operators over a shared spectrum. We propose a novel learning-aided approach that integrates the Lyapunov Drift-Plus-Penalty (LDPP) framework and Multi-agent Deep Deterministic Policy Gradient (MADDPG) reinforcement learning algorithms. This offers a powerful approach to learning stable and constraint-aware policies, reaping the joint benefit of both LDPP and MADDPG, in complex multiagent environments. The major challenge for this approach is to integrate these two approaches in a meaningful and effective manner. The key idea to solve this problem is to introduce a novel feature of local observation that incorporates potential negative value of the reward function due to the stochastic constraints introduced by the LDPP framework. Empirical results demonstrate that our proposed scheme outperforms the baseline methods under various conditions.

99 - GENERAL AND MISCELLANEOUS↗

A non-cooperative meta-modeling game for automated third-party calibrating, validating and falsifying constitutive laws with parallelized adversarial attacks

The evaluation of constitutive models, especially for high-risk and high-regret engineering applications, requires efficient and rigorous third-party calibration, validation and falsification. While there are numerous efforts to develop paradigms and standard procedures to validate models, difficulties may arise due to the sequential, manual, and often biased nature of the commonly adopted calibration and validation processes, thus slowing down data collections, hampering the progress towards discovering new physics, increasing expenses and possibly leading to misinterpretations of the credibility and application ranges of proposed models. This work attempts to introduce concepts from game theory and machine learning techniques to overcome many of these existing difficulties. Here, we introduce an automated meta-modeling game where two competing AI agents systematically generate experimental data to calibrate a given constitutive model and to explore its weakness such that the experiment design and model robustness can be improved through competitions. The two agents automatically search for the Nash equilibrium of the meta-modeling game in an adversarial reinforcement learning framework without human intervention. In particular, a protagonist agent seeks to find the more effective ways to generate data for model calibrations, while an adversary agent tries to find the most devastating test scenarios that expose the weaknesses of the constitutive model calibrated by the protagonist. By capturing all possible design options of the laboratory experiments into a single decision tree, we recast the design of experiments as a game of combinatorial moves that can be resolved through deep reinforcement learning by the two competing players. Our adversarial framework emulates idealized scientific collaborations and competitions among researchers to achieve a better understanding of the application range of the learned material laws and prevent misinterpretations caused by conventional AI-based third-party validation. Numerical examples are given to demonstrate the wide applicability of the proposed meta-modeling game with adversarial attacks on both human-crafted constitutive models and machine learning models.

97 MATHEMATICS AND COMPUTING↗

Advancing Building Energy Modeling with Large Language Models: Exploration and Case Studies

The rapid progression in artificial intelligence has facilitated the emergence of large language models like ChatGPT, offering potential applications extending into specialized engineering modeling, especially physics-based building energy modeling. This paper investigates the innovative integration of large language models with building energy modeling software, focusing specifically on the fusion of ChatGPT with EnergyPlus. A literature review is first conducted to reveal a growing trend of incorporating large language models in engineering modeling, albeit limited research on their application in building energy modeling. We underscore the potential of large language models in addressing building energy modeling challenges and outline potential applications including simulation input generation, simulation output analysis and visualization, conducting error analysis, co-simulation, simulation knowledge extraction and training, and simulation optimization. Three case studies reveal the transformative potential of large language models in automating and optimizing building energy modeling tasks, underscoring the pivotal role of artificial intelligence in advancing sustainable building practices and energy efficiency. The case studies demonstrate that selecting the right large language model techniques is essential to enhance performance and reduce engineering efforts. The findings advocate a multidisciplinary approach in future artificial intelligence research, with implications extending beyond building energy modeling to other specialized engineering modeling.

building energy modeling↗

“Multiagent” Screening Improves Directed Enzyme Evolution by Identifying Epistatic Mutations

Enzyme evolution has enabled numerous advances in biotechnology and synthetic biology, yet still requires many iterative rounds of screening to identify optimal mutant sequences. This is due to the sparsity of the fitness landscape, which is caused by epistatic mutations that only offer improvements when combined with other mutations. We report an approach that incorporates diverse substrate analogues in the screening process, where multiple substrates act like multiple agents navigating the fitness landscape, identifying epistatic mutant residues without a need for testing the entire combinatorial search space. We initially validate this approach by engineering a malonyl-CoA synthetase and identify numerous epistatic mutations improving activity for several diverse substrates. The majority of these mutations would have been missed upon screening for a single substrate alone. We expect that this approach can accelerate a wide array of enzyme engineering programs.

60 APPLIED LIFE SCIENCES↗

The Cost of Scaling Up in Large-Format Additive Manufacturing

Additive manufacturing (AM) of large objects has, over the last decade, required the scaling of existing material extrusion processes. The current generation of large-scale printers are primarily gantry robots with high-throughput extrusion systems. With workspaces approaching 50 m 3 , these printers have pushed the boundaries of achievable print volume while allowing the utilization of low-cost feedstocks, such as cementitious materials and polymer pellets, like those used in injection molding. Continued workspace expansion requires an examination of the inherent trade-offs, which impact capital and operational costs. Here, in this work, the authors examine these trade-offs to determine fundamental scaling laws for existing system architectures, survey the state of the art for alternative system configurations, and pose recommendations for future system designers to continue the evolution of large-scale AM systems.

3D printing↗

Decentralized Voltage Control of Large-Scale Distribution System with PVs Based on MADRL

This paper proposes a model-free decentralized control framework for the voltage regulation of large-scale distribution systems through the coordinated control of PV inverters. This is achieved by developing a novel interaction mechanism between the surrogate model and the centralized training and decentralized execution multiagent deep reinforcement learning framework. Specifically, the sparse Gaussian processes regression method is first utilized to develop the surrogate model of the original distribution system for reward calculation during the training stage, where each agent represents a sub-region in the centralized fashion for coordination strategy learning. After that, the learned control rules are used to inform controllers within each sub-region for real-time decisions with only local measurements. Comparative tests among various methods on the EPRI Ckt5 test system demonstrate the effectiveness of the proposed method.

distribution system↗

Convex Decreasing Algorithms: Distributed Synthesis and Finite-Time Termination in Higher Dimension

Here we establish finite time termination algorithms for consensus algorithms based on geometric properties that yield finite-time guarantees, suited for use in high dimension and in the absence of a central authority. These pursuits motivate a new peer to peer convex hull algorithm which is utilized for one stopping algorithm. Further an alternative lightweight norm based stopping criteria is also developed. The practical utility of the algorithm is illustrated through MATLAB simulations.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Privacy-Preserving Average Consensus With Beaver Triple and Communication Obfuscation

A privacy-preserving average consensus algorithm is proposed that synergizes the Beaver triple in secret sharing theory and noise obfuscation. The algorithm safeguards the initial values of agents against passive adversaries in a multiagent system. It is proved that the proposed algorithm can concurrently ensure average consensus and privacy, while also reducing the online computation and communication overhead compared to encryption-based ones. In addition, it imposes a less stringent condition for privacy preservation compared to certain noise-obfuscation techniques.

Beaver triple↗

Network-Level Traffic Signal Cooperation: A Higher-Order Conflict Graph Approach

Traffic signal control and cooperation are extremely important to alleviate traffic congestion in a large traffic network. This study develops a higher-order conflict graph approach for network-wide traffic signal control and cooperation. A conflict graph is applied to model the traffic signal configurations, which identifies the conflict and unconflicted movements for each intersection. In conflict graph, the node represents each movement. The weight of each node can be defined as traffic volume, queue length, fuel consumption, or any weighted combinations of these measurements. The calculation of the optimal green light duration and green light sequence (for different movements) is equivalent to sequentially finding the maximum weight independent set (MWIS) in the conflict graph. The conflict graph also provides a uniform and efficient way to connect traffic signal operations among nearby intersections spatially. Then, we introduced the concept of the k -th order neighborhood to model the degree of connectivity between each movement to the movements at upstream or downstream intersections. The weight of each node in the higher-order conflict graph not only represents its own congestion level, but also relates to the traffic conditions of nearby intersections. Through this approach, the cooperation of multiple intersections can be realized by incorporating their spatial connectivity into conflict graph and solving the MWIS problem. A simulation network is built in SUMO to test the effectiveness of the proposed method. Results suggested that the proposed model outperformed other state-of-the-art signal control methods. Also, the scheme maintains good performance under varying traffic demands.

42 ENGINEERING↗

Distributed Finite-Time Termination for Consensus Algorithm in Switching Topologies

Here, in this article, we present a finite-time stopping criterion for consensus algorithms in networks with dynamic communication topology. Prior state of the art has established convergence to the consensus value; however, the asymptotic convergence of these algorithms poses a challenge in practical settings where the response from agents is required in finite time. To this end, we propose a maximum-minimum protocol that propagates the global maximum and minimum values of agent states (while running the consensus algorithm) in the network. This article focuses on establishing that the global maximum and minimum values are strictly monotonic even for a dynamic topology, and they can be used to distributively ascertain the closeness to convergence in finite time. We rigorously show that each node can have access to the global maximum and minimum by running the proposed maximum-minimum protocol to realize a finite-time stopping criterion for the otherwise asymptotic consensus algorithm. The practical utility of the algorithm is illustrated through experiments where each agent is instantiated by a NodeJS socket.io server.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Distributed Conditions for Small-signal Stability of Power Grids and Local Control Design

Operating modern power grids with stability guarantees is markedly important. Typical methods for analyzing and certifying power grid stability are largely centralized relying on the ability of the system operator to gather network-wide information and accurately compute the system's eigenvalues. These methods are oftentimes not privacy-preserving and computationally burdensome. They are therefore, not well-suited to modern power grids where small-signal stability has to be evaluated timely, efficiently and in a privacy-preserving fashion. Herein, we introduce a distributed methodology for certifying small-signal stability of power grids and designing the local controllers. First, we analytically derive distributed conditions for network-wide stability that bus agents can inspect using local information. By leveraging these conditions, we then introduce a distributed control design algorithm (DCDA) that can guide the local control design so that stability of the interconnected system is guaranteed. The agents that adopt the proposed distributed algorithm are responsible for tuning their local controllers, producing their local control commands and ensuring that their local stability condition is met. The system operator is only responsible for verifying network-wide stability upon receiving affirmative responses from all agents and, announcing, that the overall system is stable. The proposed DCDA algorithm is numerically validated via simulations using the IEEE 39-bus system.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Minimal Energy Routing of a Leader and a Wingmate with Periodic Connectivity

We consider a route planning problem in which two unmanned vehicles are required to complete a set of tasks present at distinct locations, referred to as targets, with minimum energy consumption. The mission environment is hazardous, and to ensure a safe operation, the UVs are required to communicate with each other at every target they visit. The problem objective is to determine the allocation of the tasks to the UVs and plan tours for the UVs to visit the targets such that the weighted sum of the distances traveled by the UVs and the distances traveled by the communicating signals between them is minimized. We formulate this problem as an Integer program and show that naively solving the problem using commercially available off-the-shelf solvers is insufficient in determining scalable solutions efficiently. To address this computational challenge, we develop an approximation and a heuristic algorithm, and employ them to compute high-quality solutions to a special case of the problem where equal weights are assigned to the distances traveled by the vehicles and the communicating signals. For this special case, we show that the approximation algorithm has a fixed approximation ratio of 3.75. We also develop lower bounds to the optimal cost of the problem to evaluate the performance of these algorithms on large-scale instances. We demonstrate the performance of these algorithms on 500 randomly generated instances with the number of targets ranging from 6 to 100, and show that the algorithms provide high-quality solutions to the problem swiftly; the average computation time of the algorithmic solutions is within a fraction of a second for instances with at most 100 targets. Finally, we show that the approximation ratio has a variable ratio for the weighted case of the problem. Specifically, if ρ denotes the ratio of the weights assigned to the distances representing the communication and travel costs, the algorithm has an a posteriori ratio of $3 + \frac{3ρ}{4}$ when ρ ≥ 1, and $\frac{3}{ρ}$ + $\frac{3}{4}$ when ρ ≤ 1.

42 ENGINEERING↗

Optimal Equilibrium Selection of Price-Maker Agents in Performance-Based Regulation Market

This paper analyzes the oligopolistic equilibrium of multiple price-maker agents in performance-based regulation (PBR) markets. In these markets, there are price-maker agents representing some frequency regulation (FR) providers and a number of independent price-taker FR providers. An equilibrium problem with equilibrium constraints (EPEC) model is employed in this paper to study the equilibria of a PBR market in the presence of price-maker agents and pricetaker FR providers. Due to the incorporation of the FR providers' dynamics, the proposed model is reformulated as a mixed-integer linear programming (MILP) problem over innovative mathematical techniques. An optimal equilibrium point is also selected for the market, where none of the agents is unique deviator and the dynamic performance of power system is improved simultaneously. The effectiveness of proposed optimal equilibrium point is evaluated in the numerical results section by comparing the outputs with conventional optimal dispatches of the FR providers.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Interpreting Primal-Dual Algorithms for Constrained Multiagent Reinforcement Learning: Preprint

We study multiagent reinforcement learning (MARL) with constraints. This setting is gaining importance as MARL algorithms find new applications in real-world systems ranging from power grids to drone swarms. Most constrained MARL (C-MARL) algorithms use a primal-dual approach to enforce constraints through a penalty function added to the reward. In this paper, we study the structural effects of the primal-dual approach on the constraints and value function. First, we show that using the constraint evaluation as the penalty leads to a weak notion of safety, but by making simple modifications to the penalty function, we can enforce meaningful probabilistic safety constraints. Second, we show that the penalty term changes the value function in a way that is easy to model, and demonstrate the consequences of not doing so. We conclude with simulations in a simple constrained multiagent environment to back up the theoretical results.

data-driven control↗