Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Multi-agent systems”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Traffic Signal Optimization by Integrating Reinforcement Learning and Digital Twins

Machine learning (ML) methods, especially reinforcement learning (RL), have been widely considered for traffic signal optimization in intelligent transportation systems. Most of these ML methods are centralized, lacking in scalability and adaptability in large traffic networks. Further, it is challenging to train such ML models due to the lack of training platforms and/or the cost of deploying and training in a real traffic networks. This paper presents an approach for the integration of decentralized graph-based multi-agent reinforcement learning (DGMARL) with a Digital Twin (DT) to optimize traffic signals for the reduction of traffic congestion and network-wide fuel consumption related to stopping. Specifically, the DGMARL agents learn traffic state patterns and make decisions regarding traffic signal control with assistance from a Digital Twin module, which simulates and replicates the traffic behaviors of a real traffic network. The proposed approach was evaluated using PTV-Vissim [1], a microscopic traffic simulation platform. PTV-Vissim is also the simulation engine of the DT, enabling emulation and optimization of the traffic signals on the MLK Smart Corridor in Chattanooga, Tennessee. Compared to an actuated signal control baseline approach, experiment results show that Eco_PI, a developed performance measure capturing the impact of stops on fuel consumption, was reduced by 44.27% in a 24-hour and an average of 29.88% in a PM-peak-hour scenario.

Multi-Agent Reinforcement Learning, Digital Twin, ↗

Decentralized Voltage Control with Peer-to-peer Energy Trading in a Distribution Network

Utilizing distributed renewable and energy storage resources via peer-to-peer (P2P) energy trading has long been touted as a solution to improve energy system’s resilience and sustainability. Consumers and prosumers (those who have energy generation resources), however, do not have expertise to engage in repeated P2P trading, and the zero-marginal costs of renewables present challenges in determining fair market prices. To address these issues, we propose a multi-agent reinforcement learning (MARL) framework to help automate consumers’ bidding and management of their solar PV and energy storage resources, under a specific P2P clearing mechanism that utilizes the so-called supply-demand ratio. In addition, we show how the MARL framework can integrate physical network constraints to realize decentralized voltage control, hence ensuring physical feasibility of the P2P energy trading and paving ways for real-world implementations.

Feng, Chen↗

ToPolyAgent: AI agents for coarse-grained bead-spring topological polymer simulations

We introduce ToPolyAgent, a multi-agent AI framework for performing coarse-grained molecular dynamics (MD) simulations of topological polymers through natural language instructions. By integrating large language models (LLMs) with domain-specific computational tools, ToPolyAgent supports both interactive and autonomous simulation workflows across diverse polymer architectures, including linear, ring, brush, and star polymers, as well as dendrimers. The system consists of four LLM-powered agents: a Config Agent for generating initial polymer–solvent configurations, a Simulation Agent for executing LAMMPS-based MD simulations and conformational analyses, a Report Agent for compiling markdown reports, and a Workflow Agent for streamlined autonomous operations. Interactive mode incorporates user feedback loops for iterative refinements, while autonomous mode enables end-to-end task execution from detailed prompts. We demonstrate ToPolyAgent's versatility through case studies involving diverse polymer architectures under varying solvent conditions, thermostats, and simulation lengths. Furthermore, we highlight its potential as a research assistant by directing it to investigate the effect of interaction parameters on the linear polymer conformation, and the influence of grafting density on the persistence length of the brush polymer. By coupling natural language interfaces with rigorous simulation tools, ToPolyAgent lowers barriers to complex computational workflows and advances AI-driven materials discovery in polymer science. It lays the foundation for autonomous and extensible multi-agent scientific research ecosystems.

Ding, Lijie [Oak Ridge National Laboratory (ORNL),↗

Scalable Approaches to Selecting Key Entities in Large Networked Infrastructure Systems

This work aims at bringing advances in discrete optimization algorithms to solving practical engineering problems at scale. Often times, in many engineering design problems, there is a need to select a small set of influential or representative elements from a large ground set of entities in an optimal fashion. Submodular optimization provides for a formal way to solve such problems. Common examples with infrastructure systems involve sensor placement and identification of key entities with certain objectives. However, scaling these approaches to large infrastructure systems can be challenging because of the high computational complexity of the overall framework that include the optimization algorithms as well as high-complexity compute-oracles that provide the necessary objective function values. In this work, we explore a well-studied and widely-applicable paradigm, namely leader-selection in a multi-agent networked setting in the context of scalable methodologies. We demonstrate novel frameworks that utilize variations of accelerated submodular optimization algorithms along with linear-algebraic methods that can help accelerate the oracle computations. We further explore this combination in conjunction with graph partitioning paradigms to take advantage of the accelerated algorithms in a distributed setting. Finally we demonstrate the key findings on a practical problem in an operational setting. For this, we leverage an example road network with approximately 18k nodes and 27k edges in a traffic control application, where we seek a limited number of k=200 key intersections. This problem can be solved in a serial setting in just under 5 hours providing more than 2 orders of magnitude speed-up over methods that do not consider acceleration techniques.

Visweswara Sathanur, Arun↗

CityLearn v2: energy-flexible, resilient, occupant-centric, and carbon-aware management of grid-interactive communities

As more distributed energy resources become part of the demand-side infrastructure, quantifying their energy flexibility on a community scale is crucial. CityLearn v1 provided an environment for benchmarking control algorithms. However, there is no standardized environment utilizing realistic building-stock datasets for distributed energy resource control benchmarking without co-simulation or third-party frameworks. CityLearn v2 extends CityLearn v1 by providing a stand-alone simulation environment that leverages the End-Use Load Profiles for the U.S. Building Stock dataset to create grid-interactive communities for resilient, multi-agent, and objective control of distributed energy resources with dynamic occupant feedback. While the v1 environment used pre-simulated building thermal loads, the v2 environment uses data-driven thermal dynamics and eliminates the need for co-simulation with building energy performance software. This work details the v2 environment and provides application examples that use reinforcement learning control to manage battery energy storage system, vehicle-to-grid control, and thermal comfort during heat pump power modulation.

Nweye, Kingsley↗

Non-Stationary Policy Learning for Multi-Timescale Multi-Agent Reinforcement Learning

In multi-timescale multi-agent reinforcement learning (MARL), agents interact across different timescales. In general, policies for time-dependent behaviors, such as those induced by multiple timescales, are non-stationary. Learning non-stationary policies is challenging and typically requires sophisticated or inefficient algorithms. Motivated by the prevalence of this control problem in real-world complex systems, we introduce a simple framework for learning non-stationary policies for multi-timescale MARL. Our approach uses available information about agent timescales to define and learn periodic multi-agent policies. In detail, we theoretically demonstrate that the effects of non-stationarity introduced by multiple timescales can be learned by a periodic multi-agent policy. To learn such policies, we propose a policy gradient algorithm that parameterizes the actor and critic with phase-functioned neural networks, which provide an inductive bias for periodicity. The framework's ability to effectively learn multi-timescale policies is validated on a gridworld and building energy management environment.

control↗

SMART-COM – Scalable Multi-Agent Adaptive Resolution Tools for Collaborative Outage Management

The purpose of this grant was to conduct scientific research and prototype applications to support NPP outage staff in their adaptive decision-making in efficient scheduling and resource allocation while preventing violation of safety technical specifications. The project contributed to scientific knowledge and engineering methods in (1) user interface design, (2) scheduling optimization and risk estimation, and (2) natural language processing that would benefit the nuclear power plants in minimizing schedule overruns and even unexpected shutdowns. The research team conducted site visits at a test reactor facility and an operating nuclear power plant to gather necessary information and inputs for research and development of a software application to support NPP staff in executing their outages. The final software application consisted of three modules. First, the natural language processing module supports interactive processing of technical documentation to build a database for outage staff to query non-permissible actions on system components. This module can alleviate outage staff from reviewing extensive documentation and minimize violation of technical specifications, especially in time-sensitive situations. Second, the schedule optimization module schedules outage activities and compute risk indices that outperform existing software and current practice. This module can reduce completion time of an outage that typically have too many activities for human to optimize based on current practice that does not apply the latest operations research. Finally, the visualization module presents progress and risk information of the overall outage and individual activities, as well as enabling access to the natural language processing and schedule optimization modules. This module can provide outage staff with situation awareness that are necessary to make risk-informed decisions in response to unexpected events during the execution of an outage.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Toward Intelligent Multimodal Holography for Real-Time Chemical Imaging of Dynamic Ion Separation

Molecular-level visualization of ion transport and separation dynamics in complex environments is crucial for advancing energy systems, water purification, and critical materials recovery. Achieving this requires imaging platforms that combine structural sensitivity, chemical specificity, and real-time operation. Digital off-axis holography (DOAH) provides high-throughput, label-free quantitative phase imaging but inherently lacks chemical selectivity. Integrating DOAH with complementary spectroscopic channels such as fluorescence or hyperspectral imaging introduces the needed molecular specificity, while also creating challenges in multimodal data fusion, synchronization, and computational throughput. Artificial intelligence offers a powerful route to address these limitations by uniting physics-based reconstruction with data-driven interpretation. In this Perspective, we outline a framework for intelligent multimodal holography and demonstrate its potential using a preliminary AI-driven test case. Raw DOAH holograms of lanthanide solutions subjected to magnetic field gradients were analyzed using multi-agent AI workflows that autonomously selected reconstruction tools, extracted NMF components, and generated scientific claims consistent with true paramagnetic and diamagnetic behavior. This demonstration shows how AI-enabled reasoning can deliver real-time chemical–structural interpretation directly from raw holograms. Together, these advances define a path toward adaptive, intelligent holography platforms capable of supporting in situ chemical separations, dynamic ion transport analysis, and next-generation interfacial science.

Ricchiuti, Giovanna↗

PowerNet: Multi-agent Deep Reinforcement Learning for Scalable Powergrid Control

This paper develops an efficient multi-agent deep reinforcement learning algorithm for cooperative controls in powergrids. Specifically, we consider the decentralized inverter-based secondary voltage control problem in distributed generators (DGs), which is first formulated as a cooperative multi-agent reinforcement learning (MARL) problem. We then propose a novel on-policy MARL algorithm, PowerNet, in which each agent (DG) learns a control policy based on (sub-)global reward but local states and encoded communication messages from its neighbors. Motivated by the fact that a local control from one agent has limited impact on agents distant from it, we exploit a novel spatial discount factor to reduce the effect from remote agents, to expedite the training process and improve scalability. Furthermore, a differentiable, learning-based communication protocol is employed to foster the collaborations among neighboring agents. In addition, to mitigate the effects of system uncertainty and random noise introduced during on-policy learning, we utilize an action smoothing factor to stabilize the policy execution. To facilitate training and evaluation, we develop PGSim, an efficient, high-fidelity powergrid simulation platform. Here, experimental results in two microgrid setups show that the developed PowerNet outperforms the conventional model-based control method, as well as several state-of-the-art MARL algorithms. The decentralized learning scheme and high sample efficiency also make it viable to large-scale power grids.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Interpreting Primal-Dual Algorithms for Constrained Multiagent Reinforcement Learning: Preprint

We study multiagent reinforcement learning (MARL) with constraints. This setting is gaining importance as MARL algorithms find new applications in real-world systems ranging from power grids to drone swarms. Most constrained MARL (C-MARL) algorithms use a primal-dual approach to enforce constraints through a penalty function added to the reward. In this paper, we study the structural effects of the primal-dual approach on the constraints and value function. First, we show that using the constraint evaluation as the penalty leads to a weak notion of safety, but by making simple modifications to the penalty function, we can enforce meaningful probabilistic safety constraints. Second, we show that the penalty term changes the value function in a way that is easy to model, and demonstrate the consequences of not doing so. We conclude with simulations in a simple constrained multiagent environment to back up the theoretical results.

data-driven control↗

The Cost of Scaling Up in Large-Format Additive Manufacturing

Additive manufacturing (AM) of large objects has, over the last decade, required the scaling of existing material extrusion processes. The current generation of large-scale printers are primarily gantry robots with high-throughput extrusion systems. With workspaces approaching 50 m 3 , these printers have pushed the boundaries of achievable print volume while allowing the utilization of low-cost feedstocks, such as cementitious materials and polymer pellets, like those used in injection molding. Continued workspace expansion requires an examination of the inherent trade-offs, which impact capital and operational costs. Here, in this work, the authors examine these trade-offs to determine fundamental scaling laws for existing system architectures, survey the state of the art for alternative system configurations, and pose recommendations for future system designers to continue the evolution of large-scale AM systems.

3D printing↗

Multi-agent AI collaboration for digital twin development and assessment

Developing a digital twin (DT) model involves different steps that encompass formulating requirements, model development, implementation, and assessment with respect to real applications. Human expertise is required to coordinate and implement different steps in the DT development and assessment process. However, certain parts of this process can be automated using artificial intelligence (AI) agents for efficient workflow development. In this work, we test and analyze a multiagent AI collaboration with humans in the loop to automate different elements of the DT development and assessment process. To implement the workflow for multiagent AI DT development and assessment, we use Autogen, a multiagent framework developed by Microsoft. Autogen offers a modular and flexible framework for configuring and designing task-specific multiagent workflows. In this framework, large language models (LLMs) form the core intelligence of the AI agents where the quality and performance of the automated element is governed by the inherent capabilities and knowledge base of the LLM. We use retrieval augmented generation to supplement the LLM with relevant domain-specific information for DT requirement formulation. We illustrate this multiagent workflow using a case study on a thermal energy storage system, focusing on how AI agents can collaborate with humans to expedite and optimize different elements of DT development and assessment process.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Nuclear microreactor transient and load-following control with deep reinforcement learning

The economic feasibility of nuclear microreactors will depend on minimizing operating costs through advancements in autonomous control, especially when these microreactors are operating alongside other types of energy systems (e.g., renewable energy). This study explores the application of deep reinforcement learning (RL) for real-time drum control in microreactors, exploring performance in regard to load-following scenarios. By leveraging a point kinetics model with thermal and xenon feedback, we first establish a baseline using a single-output RL agent, then compare it against a traditional proportional–integral–derivative (PID) controller. This study demonstrates that RL controllers, including both single- and multi-agent RL (MARL) frameworks, can achieve similar or even superior load-following performance as traditional PID control across a range of load-following scenarios. In short transients, the RL agent was able to reduce the tracking error rate in comparison to PID by one half to one third. Over extended 300-minute load-following scenarios in which xenon feedback becomes a dominant factor, PID maintained better accuracy, but RL still remained within a 1% error margin despite being trained only on short-duration scenarios. This highlights RL’s strong ability to generalize and extrapolate to longer, more complex transients, affording substantial reductions in training costs and reduced overfitting. Furthermore, when control was extended to multiple drums, MARL enabled independent drum control as well as maintained reactor symmetry constraints without sacrificing performance---an objective that standard single-agent RL could not learn. We also found that, as increasing levels of Gaussian noise were added to the power measurements, the RL controllers were able to maintain lower error rates than PID, and to do so with at least 10% and upwards of 150% less control effort. These findings illustrate RL's potential for autonomous nuclear reactor control, laying the groundwork for future integration into high-fidelity simulations and experimental validation efforts.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Safe Exploration Reinforcement Learning for Load Restoration using Invalid Action Masking

This paper addresses the load restoration problem after a power outage event. Our primary proposed methodology uses a multi-agent reinforcement learning method to make the optimal sequential decisions on picking up critical loads. Typically, a negative reward is provided to discourage the agents from selecting decisions that violate physical constraints during the restoration process. However, the main disadvantage of this approach is its difficulty in applying it to large-scale systems due to the curse of dimensionality. This paper introduces the invalid action masking technique to overcome this limitation. The features of this technique include zero physical constraint violations, reduced training time, and stabilization of the explo- ration process. Simulation results are performed in IEEE 13-node and IEEE 123-node systems showing the better performance of the proposed algorithm in comparison to the conventional approaches both in terms of restored power and learning curve.

reinforcement learning, blackstart, artificial int↗

mada-tools: MCP servers, configurations, skills, and examples for MADA

MADA-tools (Multi-Agent Design Assistant tools) is a library for defining MCP (Model Context Protocol) servers that can be used by AI agents in the MADA project. Each MCP server provides a focused set of tools that enhances an LLM's knowledge and capabilities for a specific domain, for example, how to launch jobs with Flux versus Slurm. The library makes it easy to configure and start multiple MCP servers using configuration files or command line options. Once running, these servers are intended to be consumed by one or more agents in the MADA ecosystem. The system is designed to be extensible so that future projects can contribute their own MCP servers, skills, and toolsets.

Gunnarson, BrianS [Lawrence Livermore National Lab↗

Building MCP-native hierarchical AI scientist ecosystems: a perspective on scaling multi-agent scientific discovery

Large language models (LLMs) are evolving from chatbots with limited tool-using capabilities to agentic AI systems that can perform deep research, assist in proposing hypotheses, help design experiments, automate data analysis, and draft scientific reports. However, there are currently two bottlenecks limiting LLMs' real-world impact on the broader scientific research community beyond academic demonstrations: lack of interoperability (repetitive manual tool-integration is required across scenarios) and the need for scalable coordination (unstructured communication and memory become brittle as the number of agents grows). In this Perspective, we argue that the next phase of agentic scientific discovery requires the development of an ecosystem of protocol-native agents and tools organized through hierarchies inspired by human society, beyond the current paradigm of a single monolithic “AI scientist”. We use Model Context Protocol (MCP) as a concrete example of an emerging interoperability layer for scientific tool and context exchange, and we propose three complementary pathways to increase the scaling capabilities of an MCP-native scientific ecosystem by addressing the composability issues: (1) MCP servers for high-value scientific tools maintained by domain experts, (2) automated transformation of existing code repositories into MCP services, and (3) autonomous invention and evolution of new agents and workflows. Finally, we provide a practical roadmap for scaling AI-driven scientific discovery by expanding tool supply and coordination in MCP-native scientific ecosystems.

97 MATHEMATICS AND COMPUTING↗

Towards Agentic AI on Particle Accelerators

As particle accelerators grow in complexity, traditional control methods face increasing challenges in achieving optimal performance. This paper envisions a paradigm shift: a decentralized multi-agent framework for accelerator control, powered by Large Language Models (LLMs) and distributed among autonomous agents. We present a proposition of a self-improving decentralized system where intelligent agents handle high-level tasks and communication and each agent is specialized control individual accelerator components. This approach raises some questions: What are the future applications of AI in particle accelerators? How can we implement an autonomous complex system such as a particle accelerator where agents gradually improve through experience and human feedback? What are the implications of integrating a human-in-the-loop component for labeling operational data and providing expert guidance? We show two examples, where we demonstrate viability of such architecture.

43 PARTICLE ACCELERATORS↗