
요약
Ataraxos, an AI for the board wargame Stratego, establishes a design pattern for reinforcement learning and search that is effective under large amounts of hidden information, a longstanding desideratum of the field of strategic decision-making.
본문
Design of Ataraxos Ataraxos consists of two interdependent self-play reinforcement learning processes, realized by transformer networks for set-up selection and move selection; a belief network trai…