Another experiment was to develop a Policy Network. This is a network which provides the probability for all (legal) moves in a position. This can be used to sort moves and also use the value to guide LMR.
The tool was developed by Codex.
In this example the network was 90 (binary) inputs, 45 for white and 45 for black. One layer with 64 neurons (relu activation), and 81 output neurons (softmax activation). In this implementation only man-moves were included (with white to move). For black to move the rotated input was used.
See first the training-tab, training with 10K games is in less then 1 minute.
Next we can for every position check the move probability distribution according to the network (policy inspector).
In this example the move 11-17 was played in the game, and also selected by the policy network (probability 44.4%). Also interesting that the 3 moves which lose 1 man 16-21, 23-29 and 22-28 have low probability (0.69% , 1.86% and 3.32%).
I will issue a separate post regarding match results.
Bert