Oh no! Stratego had been on my mind as something we just hadn't tried hard enough to make a winning bot for, including the DeepMind effort from 2022. I was planning to make the first one.
I thought this was slightly less crank-coded than trying to prove the Riemann Hypothesis, but maybe these days you just ask Claude to do that and it tells you there's a counterexample at 1 + πi that no one ever noticed before.
This approach also works for Hanabi, which is a very interesting game. You can't see your own cards, but the other players can. I bought the game because someone on a reinforcement learning podcast [2] mentioned it, and actually played it multiple times.
This puts the earlier "Mastering the Game of Stratego with Model-Free Multiagent Reinforcement Learning", 2022 [1] in some perspective. Apparently the "mastering" in 2022 wasn't quite there yet. Four years later, the new approach seems to actually be better than humans.
I thought this was slightly less crank-coded than trying to prove the Riemann Hypothesis, but maybe these days you just ask Claude to do that and it tells you there's a counterexample at 1 + πi that no one ever noticed before.
[1] https://en.wikipedia.org/wiki/Hanabi_(card_game)
[2] https://www.talkrl.com/episodes/jakob-foerster
[1] https://arxiv.org/abs/2206.15378