Reinforcement learning for online adaptation of model predictive controllers: Application to a selective catalytic reduction unit
Here we present a novel application of reinforcement learning (RL) for online dynamic tuning of model predictive controllers (MPC). Applying a state-action-reward-state-action (SARSA) algorithm for temporal difference learning with a control-specific reward function improves the error tracking performance of a standard MPC formulation. The proposed RL approach is also readily adaptable to other MPCs, or entirely different control approaches. Practical details for the implementation of the RL-MPC algorithm are also presented. The proposed algorithm is applied to a case study of controlling nitrogen oxide (NO x ) emissions in an industrial selective catalytic reduction (SCR) unit, a control problem characterized by significant nonlinearity and time delay. Along with an RL-MPC formulation for NOx control, another MPC is proposed to mitigate ammonia slip and decrease ammonia consumption in the SCR. Results showing the efficacy of the RL-MPC for NO x control through learning and implementation on the nonlinear SCR dynamic model are presented.