back to home

Tencent Kaiwu AI Arena: Honor of Kings 3v3

Reinforcement Learning PPO Reward Design

Tencent's Kaiwu AI Arena is a reinforcement-learning contest built on the 3v3 mode of Honor of Kings. Teams train agents that play the game against each other. We were a team of five, about 100 teams entered, and we finished 6th. The setup used PPO with GAE, a CNN plus LSTM policy, and separate learner and actor processes for distributed training, following Tencent's own paper on mastering MOBA games with deep reinforcement learning.

My part was reward design and keeping training stable. Each role, the marksman, the mid laner and the jungler, got its own block of reward terms, and only the jungler was paid for attacking neutral monsters. The default rewards paid too much for things that look good in the next few seconds, like gold. I cut the gold reward sharply, raised the weight on damage to enemy heroes and on towers, added a penalty for dying, and added new rewards for attacking and destroying the enemy crystal.

The idea behind all of it was to move the reward away from short-term gains and toward the actual win condition, the towers and the base, and the steps that lead there. Most of the work in a reward-shaped game turned out to be deciding what you are willing to pay for.