Julia rather than Tensorflow for training the neural network

Discussion about development of draughts in the time of computer and Internet.
Post Reply
gwiesenekker
Posts: 132
Joined: Sun Feb 20, 2011 21:04
Real name: Gijsbert Wiesenekker

Julia rather than Tensorflow for training the neural network

Post by gwiesenekker »

Hi,

I was not too happy with 'our' Tensorflow training script (written by ChatGPT and me) for training GWD's neural network. It often broke after a Tensorflow update, I am running Tensorflow out of a Docker container so that I can revert easily. It was not very fast and for some reason training of the network always took 340 epochs, ChatGPT never identified why. I asked the AIs for any alternatives to Tensorflow and Julia looked quite promising also in terms of performance. Claude Opus converted my Tensorflow script to a Julia script but it was twice as slow as Tensorflow, but Claude Opus can (CUDA) profile and optimize Julia very well. The final script is three times faster than Tensorflow:

Code: Select all

Here's what changed in `tf/gpu-fit-embed.jl` since the first port, grouped by purpose rather than in order. The command line of the Python script still works unchanged. Everything new is an extra option with a default.

## Correctness fixes (from the Astra reviews)

- **Pattern counts and index maps** are built from the original training states before any remapping, so the exported index lists and the training rows match.
- **Validation records** that use a pattern unseen in training are removed in place, instead of making a second full-size copy. The log reports how many were removed per pattern.
- **Option checks** (learning rate, factors, patience, sizes, allowed values) now run before the data sets load. A typo no longer costs you the loading time first.
- **JSON strings are escaped** (`jstr`), so a `--data` or `--tag` containing quotes or backslashes can't produce invalid JSON.
- **The diagnostics copy far less data.** The percentile code used to copy the whole validation set around (about 12.8 GB of copies); it now looks values up through a sort order. The percentile rank is `round(p/100·(n−1))`, the same as your updated `fen2network`.
- **Timing:** a `CUDA.synchronize()` was added so the training time includes the last GPU work of the epoch.

## Learning-rate schedule (the 340-epoch fix)

- **`--schedule plateau` (the new default).** It halves the learning rate after `--patience` epochs in which val_loss improved by less than a relative 0.1%. It stops after 8 epochs without improvement, or when the learning rate drops below 1e-6.
- **`--schedule legacy`** reproduces the Keras behaviour exactly: factor 0.96, an absolute min_delta of 1e-4 and early-stop patience 1000. With losses around 1e-5, an absolute min_delta of 1e-4 means no epoch ever counts as an improvement. That is where your 340 epochs came from.
- **Overrides:** `--lr`, `--lr_factor`, `--min_delta`, `--min_delta_mode rel|abs`, `--lr_min` and `--es_patience` override either preset. The chosen settings are printed at startup.

## Output

- **`network-best.json`** is written whenever val_loss reaches a new minimum (`--save_best`, on by default).
- **The last epoch** always exports the network and runs the diagnostics, whether training ended normally or stopped early.
- **Per-epoch timing line** shows train, validation, export and diagnostics times.
- **`--shape`** (default `kingsrow2`, which you normally need to set to `6x4`) and **`--root`** were added.

## Profiling

- **`--profile_steps N`** runs 20 warm-up steps and then times N steps. It reports three things: wall time, per-phase times (upload, forward, backward, optimizer), and the CUDA.jl profiler tables. Then it exits without training.

## Speed, step 1: 42.3 → 37.4 s/epoch

- **Pinned memory:** the training and validation arrays are locked in RAM once at startup (about 19 GB), so batch uploads are true asynchronous copies.
- **Loss on the GPU:** losses are 1-element GPU arrays and are summed on the GPU. Validation predictions go into a GPU buffer, and both are read once per epoch.
- **Fused AdamW:** one kernel per parameter array, with the same arithmetic as before.
- **Partial sort in the diagnostics:** only the 90–99% ranks are sorted.

## Speed, step 2: 37.4 → 9.0 s/epoch (`--engine fast`, the default on a GPU)

- **Custom GPU kernels** for the whole training step:
  - the embedding lookup-and-sum (add and concat) and its backward pass with atomic adds;
  - the dense layers with their activations fused in;
  - the backward pass through the activations and the input gradients;
  - the weight and bias gradients, summed in a fixed order so they are deterministic;
  - the huber, mse and sccentropy losses.
- **Why it was needed:** cuBLAS copied its alpha/beta scalars from host memory on every matrix multiply, and NNlib's gather and scatter checked their indices and waited for the result. Together that meant about 30 host waits per step. There are now none, and all buffers are allocated once.
- **`--engine zygote`** is the previous code path, kept as the reference and for the CPU.
- **`--check_gradients N`** compares both engines on N batches and exits. Your run matched to within Float32 rounding.

## Current state

- A full plateau run takes 138 epochs at about 9.5 s each (about 23 minutes) and reaches val_loss 1.80e-6 with a 99th-percentile error of 0.0085.
GW
Post Reply