Engineering PapersSearch

SEARCH · Engineering Papers

Results for “GAME THEORY”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Adaptive, Distributed Control of Constrained Multi-Agent Systems

Product Distribution (PO) theory was recently developed as a broad framework for analyzing and optimizing distributed systems. Here we demonstrate its use for adaptive distributed control of Multi-Agent Systems (MASS), i.e., for distributed stochastic optimization using MAS s. First we review one motivation of PD theory, as the information-theoretic extension of conventional full-rationality game theory to the case of bounded rational agents. In this extension the equilibrium of the game is the optimizer of a Lagrangian of the (Probability dist&&on on the joint state of the agents. When the game in question is a team game with constraints, that equilibrium optimizes the expected value of the team game utility, subject to those constraints. One common way to find that equilibrium is to have each agent run a Reinforcement Learning (E) algorithm. PD theory reveals this to be a particular type of search algorithm for minimizing the Lagrangian. Typically that algorithm i s quite inefficient. A more principled alternative is to use a variant of Newton's method to minimize the Lagrangian. Here we compare this alternative to RL-based search in three sets of computer experiments. These are the N Queen s problem and bin-packing problem from the optimization literature, and the Bar problem from the distributed RL literature. Our results confirm that the PD-theory-based approach outperforms the RL-based scheme in all three domains.

Bieniawski, Stefan

Product Distribution Theory for Control of Multi-Agent Systems

Product Distribution (PD) theory is a new framework for controlling Multi-Agent Systems (MAS's). First we review one motivation of PD theory, as the information-theoretic extension of conventional full-rationality game theory to the case of bounded rational agents. In this extension the equilibrium of the game is the optimizer of a Lagrangian of the (probability distribution of) the joint stare of the agents. Accordingly we can consider a team game in which the shared utility is a performance measure of the behavior of the MAS. For such a scenario the game is at equilibrium - the Lagrangian is optimized - when the joint distribution of the agents optimizes the system's expected performance. One common way to find that equilibrium is to have each agent run a reinforcement learning algorithm. Here we investigate the alternative of exploiting PD theory to run gradient descent on the Lagrangian. We present computer experiments validating some of the predictions of PD theory for how best to do that gradient descent. We also demonstrate how PD theory can improve performance even when we are not allowed to rerun the MAS from different initial conditions, a requirement implicit in some previous work.

Lee, Chia Fan

Adaptive Multi-Agent Systems for Constrained Optimization

Product Distribution (PD) theory is a new framework for analyzing and controlling distributed systems. Here we demonstrate its use for distributed stochastic optimization. First we review one motivation of PD theory, as the information-theoretic extension of conventional full-rationality game theory to the case of bounded rational agents. In this extension the equilibrium of the game is the optimizer of a Lagrangian of the (probability distribution of) the joint state of the agents. When the game in question is a team game with constraints, that equilibrium optimizes the expected value of the team game utility, subject to those constraints. The updating of the Lagrange parameters in the Lagrangian can be viewed as a form of automated annealing, that focuses the MAS more and more on the optimal pure strategy. This provides a simple way to map the solution of any constrained optimization problem onto the equilibrium of a Multi-Agent System (MAS). We present computer experiments involving both the Queen s problem and K-SAT validating the predictions of PD theory and its use for off-the-shelf distributed adaptive optimization.

Macready, William

Distributed Optimization

We demonstrate a new framework for analyzing and controlling distributed systems, by solving constrained optimization problems with an algorithm based on that framework. The framework is ar. information-theoretic extension of conventional full-rationality game theory to allow bounded rational agents. The associated optimization algorithm is a game in which agents control the variables of the optimization problem. They do this by jointly minimizing a Lagrangian of (the probability distribution of) their joint state. The updating of the Lagrange parameters in that Lagrangian is a form of automated annealing, one that focuses the multi-agent system on the optimal pure strategy. We present computer experiments for the k-sat constraint satisfaction problem and for unconstrained minimization of NK functions.

Macready, William

Improving Simulated Annealing by Replacing Its Variables with Game-Theoretic Utility Maximizers

The game-theory field of Collective INtelligence (COIN) concerns the design of computer-based players engaged in a non-cooperative game so that as those players pursue their self-interests, a pre-specified global goal for the collective computational system is achieved as a side-effect. Previous implementations of COIN algorithms have outperformed conventional techniques by up to several orders of magnitude, on domains ranging from telecommunications control to optimization in congestion problems. Recent mathematical developments have revealed that these previously developed algorithms were based on only two of the three factors determining performance. Consideration of only the third factor would instead lead to conventional optimization techniques like simulated annealing that have little to do with non-cooperative games. In this paper we present an algorithm based on all three terms at once. This algorithm can be viewed as a way to modify simulated annealing by recasting it as a non-cooperative game, with each variable replaced by a player. This recasting allows us to leverage the intelligent behavior of the individual players to substantially improve the exploration step of the simulated annealing. Experiments are presented demonstrating that this recasting significantly improves simulated annealing for a model of an economic process run over an underlying small-worlds topology. Furthermore, these experiments reveal novel small-worlds phenomena, and highlight the shortcomings of conventional mechanism design in bounded rationality domains.

Wolpert, David H.

Improving Simulated Annealing by Recasting it as a Non-Cooperative Game

The game-theoretic field of COllective INtelligence (COIN) concerns the design of computer-based players engaged in a non-cooperative game so that as those players pursue their self-interests, a pre-specified global goal for the collective computational system is achieved "as a side-effect". Previous implementations of COIN algorithms have outperformed conventional techniques by up to several orders of magnitude, on domains ranging from telecommunications control to optimization in congestion problems. Recent mathematical developments have revealed that these previously developed game-theory-motivated algorithms were based on only two of the three factors determining performance. Consideration of only the third factor would instead lead to conventional optimization techniques like simulated annealing that have little to do with non-cooperative games. In this paper we present an algorithm based on all three terms at once. This algorithm can be viewed as a way to modify simulated annealing by recasting it as a non-cooperative game, with each variable replaced by a player. This recasting allows us to leverage the intelligent behavior of the individual players to substantially improve the exploration step of the simulated annealing. Experiments are presented demonstrating that this recasting improves simulated annealing by several orders of magnitude for spin glass relaxation and bin-packing.

Wolpert, David

Improving Search Algorithms by Using Intelligent Coordinates

We consider algorithms that maximize a global function G in a distributed manner, using a different adaptive computational agent to set each variable of the underlying space. Each agent eta is self-interested; it sets its variable to maximize its own function g (sub eta). Three factors govern such a distributed algorithm's performance, related to exploration/exploitation, game theory, and machine learning. We demonstrate how to exploit alI three factors by modifying a search algorithm's exploration stage: rather than random exploration, each coordinate of the search space is now controlled by a separate machine-learning-based player engaged in a noncooperative game. Experiments demonstrate that this modification improves simulated annealing (SA) by up to an order of magnitude for bin packing and for a model of an economic process run over an underlying network. These experiments also reveal interesting small-world phenomena.

Wolpert, David H.

Time-optimal Aircraft Pursuit-evasion with a Weapon Envelope Constraint

The optimal pursuit-evasion problem between two aircraft including a realistic weapon envelope is analyzed using differential game theory. Six order nonlinear point mass vehicle models are employed and the inclusion of an arbitrary weapon envelope geometry is allowed. The performance index is a linear combination of flight time and the square of the vehicle acceleration. Closed form solution to this high-order differential game is then obtained using feedback linearization. The solution is in the form of a feedback guidance law together with a quartic polynomial for time-to-go. Due to its modest computational requirements, this nonlinear guidance law is useful for on-board real-time implementation.

Menon, P. K. A.

AGATE: Adversarial Game Analysis for Tactical Evaluation

AGATE generates a set of ranked strategies that enables an autonomous vehicle to track/trail another vehicle that is trying to break the contact using evasive tactics. The software is efficient (can be run on a laptop), scales well with environmental complexity, and is suitable for use onboard an autonomous vehicle. The software will run in near-real-time (2 Hz) on most commercial laptops. Existing software is usually run offline in a planning mode, and is not used to control an unmanned vehicle actively. JPL has developed a system for AGATE that uses adversarial game theory (AGT) methods (in particular, leader-follower and pursuit-evasion) to enable an autonomous vehicle (AV) to maintain tracking/ trailing operations on a target that is employing evasive tactics. The AV trailing, tracking, and reacquisition operations are characterized by imperfect information, and are an example of a non-zero sum game (a positive payoff for the AV is not necessarily an equal loss for the target being tracked and, potentially, additional adversarial boats). Previously, JPL successfully applied the Nash equilibrium method for onboard control of an autonomous ground vehicle (AGV) travelling over hazardous terrain.

Huntsberger, Terrance L.

A Game Theoretic Fault Detection Filter

The fault detection process is modelled as a disturbance attenuation problem. The solution to this problem is found via differential game theory, leading to an H(sub infinity) filter which bounds the transmission of all exogenous signals save the fault to be detected. For a general class of linear systems which includes some time-varying systems, it is shown that this transmission bound can be taken to zero by simultaneously bringing the sensor noise weighting to zero. Thus, in the limit, a complete transmission block can he achieved, making the game filter into a fault detection filter. When we specialize this result to time-invariant system, it is found that the detection filter attained in the limit is identical to the well known Beard-Jones Fault Detection Filter. That is, all fault inputs other than the one to be detected (the "nuisance faults") are restricted to an invariant subspace which is unobservable to a projection on the output. For time-invariant systems, it is also shown that in the limit, the order of the state-space and the game filter can be reduced by factoring out the invariant subspace. The result is a lower dimensional filter which can observe only the fault to be detected. A reduced-order filter can also he generated for time-varying systems, though the computational overhead may be intensive. An example given at the end of the paper demonstrates the effectiveness of the filter as a tool for fault detection and identification.

Chung, Walter H.

Optimizing Transportation Networks for E-Waste Reverse Logistics: A Multi-Modal Cost Allocation and Pricing Strategy

The exponential growth of electronic waste (e-waste) poses critical challenges for sustainable reverse logistics and transportation network optimization. This study develops a dual-channel transportation framework for e-waste logistics that integrates dynamic freight pricing, cost allocation mechanisms, and game-theoretic coordination. The model captures interactions between centralized hubs and distributed processing networks, accounting for freight rate elasticity, volume allocation, and capacity constraints. Using Stackelberg game theory and cost-sharing strategies, the framework optimizes transportation efficiency and profit distribution across logistics channels. Numerical simulations show that the dual-channel structure increases centralized hub profit by 226.8% compared to baseline single-channel operations, while boosting total transported volume by 1.2% and nearly doubling freight collector profit under cost-sharing. Scenario analyses across regional infrastructures reveal that network density, policy incentives, and logistics costs shape routing efficiency and profit allocation. These findings suggest that coordinated strategies combining dynamic pricing, targeted infrastructure investment, and strategic cost allocation are needed to design efficient, resilient, and regionally adaptable e-waste transportation systems.

29 ENERGY PLANNING, POLICY, AND ECONOMY

Machine learning-guided design of direct methanol fuel cells with a platinum group metal-free cathode

Direct methanol fuel cells (DMFCs) offer a promising solution for clean electricity generation, particularly in small electronics and remote auxiliary power units. However, optimizing their efficiency and performance is challenging due to the complex interactions between various factors. Here, we present a novel approach that integrates experiments with machine learning to model and predict the performance of these fuel cells using atomically dispersed platinum group metal (PGM)-free catalysts at the cathode. Further, our machine learning models, trained on diverse input parameters, allow for the comprehensive optimization of DMFC performance prior to fabrication and testing. Through extensive experimental validation, we demonstrate that this data-driven approach accurately predicts key performance metrics, such as maximum power output and polarization curves. By combining our models with interpretable game-theory methods, we provide deep insights into the factors governing fuel cell performance, ultimately paving the way for the design of scalable and efficient DMFC technologies.

25 ENERGY STORAGE

Impacts and emerging research opportunities in Vehicle-Grid Integration for transportation: A review

This review provides a comprehensive examination of Vehicle-Grid Integration (VGI) technologies and their impacts on transportation systems, with a particular emphasis on the transportation-energy nexus. It systematically explores how VGI affects key transportation applications such as charging infrastructure planning, electric vehicle (EV) routing, smart charging coordination, shared mobility, and dynamic pricing. By synthesizing recent literature from both transportation and energy systems perspectives, this study highlights how advanced methodologies, such as reinforcement learning, game theory, and optimization techniques, are used to model the complex interactions between EVs, mobility patterns, and distributed energy systems. Furthermore, the review also identifies critical challenges, including behavioral factors, data limitations, and system scalability. Drawing on these insights, the paper outlines emerging research opportunities to support the design of integrated, resilient, and user-centric VGI solutions that advance sustainable mobility and energy system efficiency.

Charging coordination

The Delicate Balance Redux: The Role of Nuclear Forces, Damage Limitation and Uncertainty in Future U.S.-China Crises

What is the impact of damage limitation capabilities like counterforce and missile defenses on deterrence, when their efficacy in stopping an adversary nuclear attack is uncertain? This is a key unanswered question to understand “how much is enough” for the United States to deter China and Russia in future nuclear crises. In this paper we extend an established, single move game theory model to capture the dynamics of two players in a nuclear crisis having varying damage limitation capabilities with uncertain effectiveness. Our model formalizes the logic of the “delicate balance” school of deterrence, which states that leverage in a crisis is driven by the risk each player can take with their combined strategic forces, and that those risks carry uncertainty as nuclear forces are hard to deliver against technologically advanced adversaries. Our model shows that damage limitation capabilities—even those with significant uncertainty around them like cyber or electronic warfare—can drive bargaining outcomes in an array of nuclear crises. We then apply these bargaining outcomes to the expected U.S.-China strategic balance as China builds out its nuclear force through 2035. We apply published force exchange models to determine the expected damage each side will be able to deliver, and we use these values to determine the likelihood that the U.S. can prevail in crises of varying stakes. Last, we show that U.S. policymakers have an array of options to improve future bargaining outcomes, evaluating how additional nuclear forces trade against improvements in damage limitation.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF

Nonzero-sum differential games.

Differential games theory with nonzero sum for application to economic analysis, discussing Nash equilibrium, minimax and noninferior strategies set

Ho, Y. C.

Differential games

Differential game theory, discussing two player zero sum situation

Varaiya, P. P.