policy network

Discussion about development of draughts in the time of computer and Internet.
Post Reply
BertTuyt
Posts: 1681
Joined: Wed Sep 01, 2004 19:42

policy network

Post by BertTuyt »

With the policy network developed I asked Codex to optimize the LMR in relation to the policy output, in such a way that at least the number of nodes and/or time was reduced with 50%.

Codex did several experiments, the total time spent by Codex was around 1 hour.
I also asked Codex to write a report in the end.
All code changes were also implemented by the AI in an experimental version of the Damage engine (version 18.4).

I did not have time to study the extensive report in detail, see attached.

As a proof of the pudding a 158 games DXP match was played against KR, settings no book, 6p DB, 1 core, and 80 moves for 1 minute.
Match result 158 draws. See also attached file.

This result is not yet optimized, but as I believe we have reached a ceiling I don't expect there is much ELO to gain.

Bert
Attachments
Damage184_LMR_experiments_and_recommendation.pdf
(1.21 MiB) Downloaded 10 times
dxpmatch_20261004.pdn
(158.44 KiB) Downloaded 8 times
Joost Buijs
Posts: 567
Joined: Wed May 04, 2016 11:45
Real name: Joost Buijs

Re: policy network

Post by Joost Buijs »

For running tests nowadays, I use 4" + 0.1' time control. This works out to about 10 seconds per game, which means 3 games per minute for both sides combined. With 32 cores, I can blast through a full ballot2 match within 2 minutes. Even then, there are at most 4 decisive games in a 158-game match. So, I think we really reached the ceiling.

Joost
Post Reply