There has been a lot of discussion on this board about the problem that the strong engines always draw when they play another strong engine in a tournament. A few months ago I ran some experiments to see if this problem could be addressed using somewhat lopsided start positions. Below is a summary of various positions that I tried and the results of short DXP matches between two strong engines. All test matches were run with 6pc dbs, time control of 70 moves in 5 minutes, and 1 search thread. Each engine got to play both white and black for an equal number of games.
[FEN "W:W31,32,33,34,35,36,37,38,39,40,41,42,44,45,46,47,50:B1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,18,19"]
1 wins, 0 losses, 9 draws, 0 unknowns
Surprisingly, this position was not unbalanced enough to get many decisive games.
[FEN "W:W31,32,33,34,35,36,37,38,39,41,42,43,46,47,48:B1,2,3,4,5,6,7,8,9,10,11,12,13,14,15"]
0 wins, 0 losses, 10 draws, 0 unknowns
This just wasn't a problem for either engine.
[FEN "W:W31,32,33,34,35,36,37,38,39,40,41,42,43,44,45,46,47:B1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,20"]
4 wins, 2 losses, 4 draws, 0 unknowns
[FEN "W:W26,27,29,30,31,32,33,34,35,36,37,38,39,40,41,43,44,45,46,50:B1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20"]
3 wins, 0 losses, 7 draws, 0 unknowns
[FEN "W:W26,27,29,30,31,32,33,34,35,36,37,38,39,40,41,42,44,45,46,50:B1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20"]
4 wins, 5 losses, 0 draws, 1 unknowns
Upon analysis with the 8pc db, the 1 unknown was a draw that was taking more than 70 moves for the 6pc engines to see.
My goal was to try to find some start positions that were on the borderline between draw and loss, so ideally about half of the games would be draws.
I have one more result to report, but the Forum will not let me add another image to this post. I will try to submit it in a follow up post.
-- Ed

Start positions
-
Ed Gilbert
- Posts: 878
- Joined: Sat Apr 28, 2007 14:53
- Real name: Ed Gilbert
- Location: Morristown, NJ USA
- Contact:
-
Ed Gilbert
- Posts: 878
- Joined: Sat Apr 28, 2007 14:53
- Real name: Ed Gilbert
- Location: Morristown, NJ USA
- Contact:
Re: Start positions
Here is one more:
[FEN "W:W31,32,33,34,35,36,37,38,39,40,41,42,43,44,45,46,47,48,49,50:B1,2,3,K4,5,6,7,8,9,10,11,12,13,14,15,17,18,19"]
It may not be clear from the image, but there is a black king on 4.
16 wins, 12 losses, 10 draws, 2 unknowns
All but 1 of the wins were black wins.
[FEN "W:W31,32,33,34,35,36,37,38,39,40,41,42,43,44,45,46,47,48,49,50:B1,2,3,K4,5,6,7,8,9,10,11,12,13,14,15,17,18,19"]
It may not be clear from the image, but there is a black king on 4.
16 wins, 12 losses, 10 draws, 2 unknowns
All but 1 of the wins were black wins.
-
Joost Buijs
- Posts: 530
- Joined: Wed May 04, 2016 11:45
- Real name: Joost Buijs
Re: Start positions
Hi Ed,
Neural networks perform best when they are trained on positions that resemble positions that can occur in practice during normal games. Artificial positions like these can never occur on the board in a normal game. Especially the last position, with the black king on square 4, a network will struggle greatly with this.
It is, of course, possible to train a network on artificial positions, but that likely degrades the quality of the network for positions encountered in practice.
I think the rules of the game need to be changed to solve the high draw rate of draughts.
Joost
Neural networks perform best when they are trained on positions that resemble positions that can occur in practice during normal games. Artificial positions like these can never occur on the board in a normal game. Especially the last position, with the black king on square 4, a network will struggle greatly with this.
It is, of course, possible to train a network on artificial positions, but that likely degrades the quality of the network for positions encountered in practice.
I think the rules of the game need to be changed to solve the high draw rate of draughts.
Joost
-
gwiesenekker
- Posts: 88
- Joined: Sun Feb 20, 2011 21:04
- Real name: Gijsbert Wiesenekker
Re: Start positions
Hi,
Although generally speaking you should not evaluate positions with a neural network that are 'far off' from the positions the network has been trained on it is quite instructive to evaluate the following 'sparse' positions with your NN including the empty board:
W:W:B {0.0}
W:W6:B {0.0}
W:W7:B {0.0}
W:W8:B {0.0}
W:W9:B {0.0}
W:W10:B {0.0}
W:W11:B {0.0}
W:W12:B {0.0}
W:W13:B {0.0}
W:W14:B {0.0}
W:W15:B {0.0}
W:W16:B {0.0}
W:W17:B {0.0}
W:W18:B {0.0}
W:W19:B {0.0}
W:W20:B {0.0}
W:W21:B {0.0}
W:W22:B {0.0}
W:W23:B {0.0}
W:W24:B {0.0}
W:W25:B {0.0}
W:W26:B {0.0}
W:W27:B {0.0}
W:W28:B {0.0}
W:W29:B {0.0}
W:W30:B {0.0}
W:W31:B {0.0}
W:W32:B {0.0}
W:W33:B {0.0}
W:W34:B {0.0}
W:W35:B {0.0}
W:W36:B {0.0}
W:W37:B {0.0}
W:W38:B {0.0}
W:W39:B {0.0}
W:W40:B {0.0}
W:W41:B {0.0}
W:W42:B {0.0}
W:W43:B {0.0}
W:W44:B {0.0}
W:W45:B {0.0}
W:W46:B {0.0}
W:W47:B {0.0}
W:W48:B {0.0}
W:W49:B {0.0}
W:W50:B {0.0}
B:W:B1 {0.0}
B:W:B2 {0.0}
B:W:B3 {0.0}
B:W:B4 {0.0}
B:W:B5 {0.0}
B:W:B6 {0.0}
B:W:B7 {0.0}
B:W:B8 {0.0}
B:W:B9 {0.0}
B:W:B10 {0.0}
B:W:B11 {0.0}
B:W:B12 {0.0}
B:W:B13 {0.0}
B:W:B14 {0.0}
B:W:B15 {0.0}
B:W:B16 {0.0}
B:W:B17 {0.0}
B:W:B18 {0.0}
B:W:B19 {0.0}
B:W:B20 {0.0}
B:W:B21 {0.0}
B:W:B22 {0.0}
B:W:B23 {0.0}
B:W:B24 {0.0}
B:W:B25 {0.0}
B:W:B26 {0.0}
B:W:B27 {0.0}
B:W:B28 {0.0}
B:W:B29 {0.0}
B:W:B30 {0.0}
B:W:B31 {0.0}
B:W:B32 {0.0}
B:W:B33 {0.0}
B:W:B34 {0.0}
B:W:B35 {0.0}
B:W:B36 {0.0}
B:W:B37 {0.0}
B:W:B38 {0.0}
B:W:B39 {0.0}
B:W:B40 {0.0}
B:W:B41 {0.0}
B:W:B42 {0.0}
B:W:B43 {0.0}
B:W:B44 {0.0}
B:W:B45 {0.0}
GWD's best neural network evaluates the empty position as -5, positions with a man that can promote around +200 and positions with man on the back-row at around +60. In contrast, for a weaker network (weaker meaning the result of a match against Kingsrow is worse) the numbers are -60(!), +180 and +20, so next to the 90-99% percentile errors I also use these numbers to gauge the quality of the network.
GW
Although generally speaking you should not evaluate positions with a neural network that are 'far off' from the positions the network has been trained on it is quite instructive to evaluate the following 'sparse' positions with your NN including the empty board:
W:W:B {0.0}
W:W6:B {0.0}
W:W7:B {0.0}
W:W8:B {0.0}
W:W9:B {0.0}
W:W10:B {0.0}
W:W11:B {0.0}
W:W12:B {0.0}
W:W13:B {0.0}
W:W14:B {0.0}
W:W15:B {0.0}
W:W16:B {0.0}
W:W17:B {0.0}
W:W18:B {0.0}
W:W19:B {0.0}
W:W20:B {0.0}
W:W21:B {0.0}
W:W22:B {0.0}
W:W23:B {0.0}
W:W24:B {0.0}
W:W25:B {0.0}
W:W26:B {0.0}
W:W27:B {0.0}
W:W28:B {0.0}
W:W29:B {0.0}
W:W30:B {0.0}
W:W31:B {0.0}
W:W32:B {0.0}
W:W33:B {0.0}
W:W34:B {0.0}
W:W35:B {0.0}
W:W36:B {0.0}
W:W37:B {0.0}
W:W38:B {0.0}
W:W39:B {0.0}
W:W40:B {0.0}
W:W41:B {0.0}
W:W42:B {0.0}
W:W43:B {0.0}
W:W44:B {0.0}
W:W45:B {0.0}
W:W46:B {0.0}
W:W47:B {0.0}
W:W48:B {0.0}
W:W49:B {0.0}
W:W50:B {0.0}
B:W:B1 {0.0}
B:W:B2 {0.0}
B:W:B3 {0.0}
B:W:B4 {0.0}
B:W:B5 {0.0}
B:W:B6 {0.0}
B:W:B7 {0.0}
B:W:B8 {0.0}
B:W:B9 {0.0}
B:W:B10 {0.0}
B:W:B11 {0.0}
B:W:B12 {0.0}
B:W:B13 {0.0}
B:W:B14 {0.0}
B:W:B15 {0.0}
B:W:B16 {0.0}
B:W:B17 {0.0}
B:W:B18 {0.0}
B:W:B19 {0.0}
B:W:B20 {0.0}
B:W:B21 {0.0}
B:W:B22 {0.0}
B:W:B23 {0.0}
B:W:B24 {0.0}
B:W:B25 {0.0}
B:W:B26 {0.0}
B:W:B27 {0.0}
B:W:B28 {0.0}
B:W:B29 {0.0}
B:W:B30 {0.0}
B:W:B31 {0.0}
B:W:B32 {0.0}
B:W:B33 {0.0}
B:W:B34 {0.0}
B:W:B35 {0.0}
B:W:B36 {0.0}
B:W:B37 {0.0}
B:W:B38 {0.0}
B:W:B39 {0.0}
B:W:B40 {0.0}
B:W:B41 {0.0}
B:W:B42 {0.0}
B:W:B43 {0.0}
B:W:B44 {0.0}
B:W:B45 {0.0}
GWD's best neural network evaluates the empty position as -5, positions with a man that can promote around +200 and positions with man on the back-row at around +60. In contrast, for a weaker network (weaker meaning the result of a match against Kingsrow is worse) the numbers are -60(!), +180 and +20, so next to the 90-99% percentile errors I also use these numbers to gauge the quality of the network.
GW
-
Joost Buijs
- Posts: 530
- Joined: Wed May 04, 2016 11:45
- Real name: Joost Buijs
Re: Start positions
Hi Gijsbert,gwiesenekker wrote: Fri Jul 24, 2026 10:56
GWD's best neural network evaluates the empty position as -5, positions with a man that can promote around +200 and positions with man on the back-row at around +60. In contrast, for a weaker network (weaker meaning the result of a match against Kingsrow is worse) the numbers are -60(!), +180 and +20, so next to the 90-99% percentile errors I also use these numbers to gauge the quality of the network.
It's nice to know how GWD thinks about these positions, but these values don't mean anything for Ares. Ares has a totally different evaluation curve and evaluation thresholds. When the network is trained on a well balanced dataset the evaluation of an empty board should be zero, in practice this is difficult to obtain, this is not very important anyway because it's only a small offset, and you will never analyse an empty board.
After being busy with it for quite some time, I get more and more the impression that the topology of a network for draughths is not very critical, and that the results are largely determined by the quality of the dataset the network is trained on. I tried many different network architectures, and they all performed within approx. 7 Elo points from each other.
Joost
Re: Start positions
Joost, I tend to agree.
The only remark is that the search depth-penalty when using a neural network should be limited.
Nothing can beat the patterns-evaluation in this respect.
I found that based upon my pruning settings, Damage should cross the 20 ply border, otherwise it can be confronted with some tactical traps.
Especially in short (bullet)-games (so 1 min/80 moves, or even less) a large float network therefore could be critical.
Also in the endgame, when there are sometimes narrow deep escapes or deep promotions, search depth helps a lot.
Most likely this will change in the future as the AI-Hype will force the processor industry to accommodate smaller networks with ultra-fast latency on the processor.
Bert
The only remark is that the search depth-penalty when using a neural network should be limited.
Nothing can beat the patterns-evaluation in this respect.
I found that based upon my pruning settings, Damage should cross the 20 ply border, otherwise it can be confronted with some tactical traps.
Especially in short (bullet)-games (so 1 min/80 moves, or even less) a large float network therefore could be critical.
Also in the endgame, when there are sometimes narrow deep escapes or deep promotions, search depth helps a lot.
Most likely this will change in the future as the AI-Hype will force the processor industry to accommodate smaller networks with ultra-fast latency on the processor.
Bert
-
Joost Buijs
- Posts: 530
- Joined: Wed May 04, 2016 11:45
- Real name: Joost Buijs
Re: Start positions
Bert, I fully agree with that.BertTuyt wrote: Fri Jul 24, 2026 15:01 Joost, I tend to agree.
The only remark is that the search depth-penalty when using a neural network should be limited.
Nothing can beat the patterns-evaluation in this respect.
I found that based upon my pruning settings, Damage should cross the 20 ply border, otherwise it can be confronted with some tactical traps.
Especially in short (bullet)-games (so 1 min/80 moves, or even less) a large float network therefore could be critical.
Also in the endgame, when there are sometimes narrow deep escapes or deep promotions, search depth helps a lot.
Most likely this will change in the future as the AI-Hype will force the processor industry to accommodate smaller networks with ultra-fast latency on the processor.
Bert
Speed is of utmost importance, with draughts you don't lose as much Elo as with chess, but the difference of a factor of two in N/s will usually make an engine lose a few games more in a 158 game match.
With pattern evaluation you only have to calculate the indices, which can easily be done with pext() and a precalculated lookup table. On machines where you can't use pext() (as the ones Krzysztof uses) you have to use shifts and masks, and the lookup tables have to be larger. Calculating the indices with AVX2 is probably faster because you can calculate the (I assume eight) indices in parallel and keep all the calculations in the AVX registers (which doesn't take much memory transfers).
Edit:
On the AMD 9960X Threadripper Ares currently does single thread 11.25 MN/s at the starting position.
On the AMD 3970X Threadripper (with turbo boost enabled) it does 5.9 MN/s.
For Kingsrow this is respectively 27.315 and 15.125 MN/s.
This is a speed difference of approx. 2.5, it still seems to be good enough not to lose any games with 30 sec. per game (90 moves) and 0.5 sec. increment. Ed's book is probably better, so I never test with books enabled, and I never tried ballot3 though, maybe this will be worse, I simply don't know.
Joost
Re: Start positions
Interesting idea. We could build a database of thousands of borderline positions, and use that as start positions for matches.Ed Gilbert wrote: Thu Jul 23, 2026 14:15 There has been a lot of discussion on this board about the problem that the strong engines always draw when they play another strong engine in a tournament. A few months ago I ran some experiments to see if this problem could be addressed using somewhat lopsided start positions. Below is a summary of various positions that I tried and the results of short DXP matches between two strong engines. All test matches were run with 6pc dbs, time control of 70 moves in 5 minutes, and 1 search thread. Each engine got to play both white and black for an equal number of games.
What would be a good criterium for finding these positions?
50% draw, 50% win would be good
50% win, 50% loss would be excellent but probably non-existent.
We would need a good decisiveness-score.
