Reinforcement learning

Discussion about development of draughts in the time of computer and Internet.
Post Reply
Joost Buijs
Posts: 566
Joined: Wed May 04, 2016 11:45
Real name: Joost Buijs

Reinforcement learning

Post by Joost Buijs »

An experiment I've wanted to try for a while now is training an international draughts network using reinforcement learning. Since it's a one-off project and coding everything from scratch takes quite a bit of time, I just never got around to it. However, with the help of Claude, it’s now actually possible to run these kinds of experiments without it eating up an excessive amount of time.

Luckily, I had already written most of the core software needed for this experiment in the past:

A game runner that can play up to 32 games in parallel using multithreading.
A game converter that translates the games into labeled positions to use as input for the network trainer.
A network trainer that uses those labeled positions to optimize a neural network.

I asked Claude to build a wrapper around these routines so the whole setup could function as a reinforcement learning system.

Here is how the loop works:

A: Play 10 x 495 ballot3 games (4,950 games total) using only basic material evaluation.
B: Translate these games into a batch of labeled positions (roughly 460,000).
C: Use these positions as input for the trainer, and train a random network for 10 epochs using this batch.

For all subsequent iterations of 4,950 games, the neural network takes over as the evaluator instead of material value, and steps B and C are repeated. The games start out with 20K nodes per move. If the network hits a plateau and stops improving, the number of nodes per move is doubled until a specific maximum is reached.

So far, this setup is working surprisingly well! On my 32-core machine, I’m already getting a decent-playing network after just 5 iterations (about 33 Elo below Kingsrow). After 46 iterations (which took about 4 hours) it's already more or less on par with Kingsrow. Of course, to train a truly strong network, it’s going to need a lot more iterations (probably around 1,000).

I've attached a PDF with a full, detailed description of the whole project.
Attachments
AresRL-manual.pdf
(146.38 KiB) Downloaded 23 times
BertTuyt
Posts: 1681
Joined: Wed Sep 01, 2004 19:42

Re: Reinforcement learning

Post by BertTuyt »

Joost, a very interesting result. Really curious where this will end. Without any doubt, Claude could also program the game runner, game converter and network training tools. So basically only the Draughts program remains.
Think we really grow towards a "game, set and match" situation is several years.
But anyway in the meantime it was fun....

Bert
Joost Buijs
Posts: 566
Joined: Wed May 04, 2016 11:45
Real name: Joost Buijs

Re: Reinforcement learning

Post by Joost Buijs »

Bert,

I didn't expect it to play a reasonable game after just 5 iterations, which amounted to 24,750 games at 20k nodes per move. Naturally, the process is slowing down now because it doubles the nodes per move whenever the network's performance plateaus.

It has currently completed 80 iterations and still shows signs of improvement, though that doesn't guarantee it will actually play any stronger. I plan to let it run for a few more days just to see how the network behaves with extended training. That said, a draw is still a draw—there is no changing that.

Joost
BertTuyt
Posts: 1681
Joined: Wed Sep 01, 2004 19:42

Re: Reinforcement learning

Post by BertTuyt »

Another way of looking at it, the computer was able to learn high-level draughts by just playing against itself and started with a mediocre level.
It has never seen any game by humans.
And after 4 hours it reached almost world champion level (or at least grandmaster++).
Many years the verdict was, the computer only copies what the programmer told him/her to do, and therefore can never be better then its creator.
Think we reach a new era. Even mathematics is no longer safe, and the only question is "who/what is next".

Bert
Last edited by BertTuyt on Mon Sep 28, 2026 20:02, edited 1 time in total.
Joost Buijs
Posts: 566
Joined: Wed May 04, 2016 11:45
Real name: Joost Buijs

Re: Reinforcement learning

Post by Joost Buijs »

To be honest, I believe that the current top 10 International draughts engines will never lose a match against the world champion. In fact, there is a very high chance these engines would win, even though the game is notoriously drawish. This shows that a computer can reach a level through just a few hours of self-play that likely surpasses the human world champion.

Of course, a search engine is also part of the equation and needs to be developed too; a neural network alone is incapable of playing a single game.

Joost
Joost Buijs
Posts: 566
Joined: Wed May 04, 2016 11:45
Real name: Joost Buijs

Re: Reinforcement learning

Post by Joost Buijs »

After each iteration, a ballot match of 990 games is played to determine whether the new network is better than the old one. After 50 iterations, progress stalled because the learning rate decay was not functioning. Now that this has been fixed, it is progressing again.
Joost Buijs
Posts: 566
Joined: Wed May 04, 2016 11:45
Real name: Joost Buijs

Re: Reinforcement learning

Post by Joost Buijs »

Reinforcement learning for International Draughts is a fun experiment. The biggest problem with this is that even games with a random factor included eventually all end in a draw, meaning the network no longer learns anything from them. An option to break the dead draw can be adding a few extra plies of randomness. To find out, I have to experiment with it somewhat longer.

After a few days of computing on my 32-core machine, the RL network is marginally better than my old network, which is based on 2 million pre-played games. However, I do not expect reinforcement learning to lead to major breakthroughs in International Draughts because this game has too many draw mechanisms, especially in the endgame.
BertTuyt
Posts: 1681
Joined: Wed Sep 01, 2004 19:42

Re: Reinforcement learning

Post by BertTuyt »

Joost, thanks for sharing.
It seems that we have reached a ceiling, until proven otherwise.
Also possible that there is still ELO te gain, but at huge resource-costs.
As we discussed often, larger networks come at a cost, but with the AI-hype, we can expect future processors to have better capabilities.
A larger network without huge depth-penalty will become feasible.
The MCTS alternative seems not to work due to combination traps.
At least all attempts failed so far.
So I have my hope on a combined policy and value-network, also to replace the history-heuristics.
But in the end it might just be an alternative route, to the same ceiling.

Bert
Post Reply