FLFrédéric Legrand
HOMEGAMESABOUT

© 2026 Frédéric Legrand. All rights reserved.

Get in touch
← Back to Blog

Teaching a neural net to play chess

September 22, 2026

How strong can a chess engine get when the real goal is just to learn PyTorch?

That was the plan, honestly. I wanted a project big enough to drag me through the parts of deep learning that tutorials skip: data pipelines, evaluation, scaling, deployment.

Chess is perfect for that. The rules never change, the data is free, and you can’t fool yourself about results. The engine wins games or it doesn’t.

Eight versions later, the engine you can play on this site sits at roughly 2040 Elo. The first version was 1096.

Most of what moved that number wasn’t what I expected.

The recipe: judgment plus calculation

The idea comes straight from AlphaZero, at about 1/1000th of the scale.

A neural network gives you judgment: how good is this position? A search gives you calculation: what happens a few moves from now?

The network reads a board as 17 planes of 8×8. Twelve planes hold the pieces, four hold castling rights, and one marks the en passant square. The board is always shown from the side to move, flipped when it’s Black’s turn, so the network learns chess once instead of once per color.

It has two heads. The value head outputs win, draw and loss probabilities. The policy head outputs 4096 logits, one per from-to square pair, and predicts the move a strong player would choose.

The current network is a ResNet with 8 residual blocks and 128 channels, about 5M parameters. I kept it small on purpose. It runs on a CPU inside a time budget, so every millisecond spent evaluating is a millisecond not spent searching.

Version 1: 20k games and a lot of noise

v1 learned from 20k Lichess games. Each sampled position was labelled with the game’s final result.

It worked, sort of. Validation accuracy sat around 67% and barely moved, whatever I changed.

That plateau was my first real lesson. Predicting the outcome from a single position has a noise floor. People throw away winning positions all the time, so the label is often wrong, and you can’t learn your way past bad labels.

The second lesson came right after. v2 trained on 168k games from players rated 1600+, with a bigger ResNet. Validation accuracy stayed flat. Playing strength jumped: v2 beat v1 +5 =5 −0.

Always measure the thing you actually care about. Loss curves are a proxy. Games are what count.

The policy head makes search affordable

v3 added the policy head. On its own, it picks the move strong players chose about 41% of the time, out of roughly 30 legal moves.

Where it really pays off is pruning. A full-width 3-ply search looks at about 27k positions per move. If you keep only the policy’s favourite moves at each level, that drops to around 1.9k, and the engine barely gets tactically blinder.

v3 beat v2 +5 =5 −0.

The odd-horizon bug

Then something weird showed up. Across every match in the v3 generation, every decisive game was won by Black. Ten out of ten.

That isn’t luck, it’s a bug in the search.

A 3-ply search ends on the engine’s own move. It never sees the opponent’s answer to its last idea, so it’s always a bit too optimistic. Both sides overextend, and whoever has to commit first (White) loses.

The fix is boring: search an even number of plies, so every line ends after the opponent replies.

Same network, 4-ply instead of 3-ply: +8 =1 −1. About +300 Elo for zero training.

Cheapest Elo in the project

Search depth beat every training trick I tried. If your engine is fast enough to look one ply deeper, do that first.

Reinforcement learning: four attempts, zero gains

Self-play is the famous part of AlphaZero, so of course I tried it.

The first attempt, plain REINFORCE on self-play games, lost to its own starting point (+0 =7 −3). The engine figured out that the safest way not to lose is to steer every game toward a draw. Checkmates per 300 games fell from 124 to 48 over the iterations. Textbook policy-gradient collapse.

The second attempt added guards: a draw penalty, a supervised anchor and an entropy bonus. The collapse went away, and so did any improvement. Dead tie.

The last two used expert iteration, where the search picks the moves during self-play and the network learns to imitate the search. That’s the actual AlphaZero loop. Both tied.

It took me a while to see why, and it wasn’t a bug. By then my training labels came from Stockfish, which is a far better teacher than my engine’s own search. Imitating my own search could only make things worse or keep them equal. AlphaZero’s loop worked because it had no better teacher to learn from.

Better labels beat clever tricks

Once game outcomes were clearly the bottleneck, the next step was obvious: stop learning from results.

v6 relabelled 1.6M positions with a local Stockfish at depth 6. Value accuracy on those cleaner targets jumped from about 65% to 83.8%.

v7 relabelled 4.3M positions at depth 10 and trained the value head on a soft target built from the centipawn score, so +99cp and +101cp stop being opposite labels.

v8 moved to the Lichess evaluations database: positions Lichess users asked the server to analyse, with Stockfish results usually at depth 20 to 60. I kept a hash-selected 25% of it, 102M positions, 24× more than v7.

I kept the network the same size through all of this. On a fixed time budget, strength is evaluation quality times evaluations per second, and a bigger network would have meant a shallower search.

v8 didn’t fit on my laptop anymore. It trained on a single L4 GPU with bf16, torch.compile and an exponential moving average of the weights, at 30.8k positions per second, 4 epochs in 3h45. The averaged copy did better on validation, so that’s the one that shipped.

The soft value target (added in v7, kept for v8) is only a few lines:

# A smooth (loss, balanced, win) distribution from the mover's centipawn score.
SOFT_K_CP = 100.0
SOFT_S_CP = 120.0
 
def cp_to_soft_targets(cp: torch.Tensor) -> torch.Tensor:
    p_win = torch.sigmoid((cp - SOFT_K_CP) / SOFT_S_CP)
    p_loss = torch.sigmoid((-cp - SOFT_K_CP) / SOFT_S_CP)
    p_balanced = (1.0 - p_win - p_loss).clamp_min(0.0)
    return torch.stack([p_loss, p_balanced, p_win], dim=1)

v8 beat v7 +6 =9 −0.

How strong is it, really?

Elo between engines is slippery, so I’ll be precise about how I got the number.

elo_tournament.py adds two Stockfish players capped at calibrated strengths (UCI_Elo 1320 and 1600) and plays them against the key versions. Then it merges those games with every ladder match and fits one Bradley-Terry model over all of it. As a sanity check, the 1600 player, left free in the fit, lands at 1605.

The ratings:

  • •v1: 1096
  • •v3: 1355
  • •v6 at 4-ply: 1742
  • •v7 at 4-ply: 1919. It swept Stockfish-1600 6-0 and beat Stockfish-1900 +3 =2 −1.
  • •v8 at 4-ply: about 2040, measured below

For a while v8’s number was the shakiest of the lot. It came only from its match against v7, with no Stockfish game of its own. So I sat it down against three handicapped Stockfish players, alternating colors:

  • •Stockfish 1600: +8 =2 −0 over 10 games
  • •Stockfish 1900: +9 =3 −0 over 12 games
  • •Stockfish 2200: +3 =5 −2 over 10 games

It never lost to 1600 or 1900, and it edged out 2200 by a single game.

Feeding those games back into the same Bradley-Terry fit puts v8 at about 2040, give or take 150. So the earlier estimate was about right, and now it actually rests on games against Stockfish.

Take the number with a grain of salt anyway. Thirty-two games still leave wide error bars. The fit rates the 2200 player at around 2000, because Stockfish’s Elo limiter is calibrated for slower games than the 50 milliseconds a move I gave it, while v8 took about two seconds a move. Games that hit 150 moves are scored on material. And engine-vs-engine Elo only loosely carries over to humans.

Serving it

The trained network is exported to ONNX (about 20MB). The first version ran it with onnxruntime on a small cloud CPU.

The depth you pick on the site is really a wall-clock budget. The engine runs iterative deepening: a 1-ply pass first, then 2, 3 and 4-ply beam searches, going deeper only while the next rung is predicted to fit the budget. The same request comes back fast on a shared vCPU and searches deeper on a faster machine, and I never had to retune it.

Before any of that, it checks every legal move for mate in one. Nothing looks more broken than an engine that declines checkmate.

Here’s the kind of position it catches instantly. White to move, and the rook on a1 has a one-move answer:

Back-rank mate: Ra8#.

Production taught me a few things the hard way.

Batch everything, but in chunks. Scoring a whole search level in one batch spiked memory into the hundreds of MB and got the 512MB instance killed. Chunks of 64 positions keep memory flat at about the same speed.

Never do slow work inside the request. The game-over email was sent before the final move’s response went out, so when the host blocked outbound email, every checkmate froze the board.

Keep the real game state. The API rebuilt the board from a plain 8×8 array on every request, which quietly dropped the en passant square. The UI would offer you an en passant capture, and then the server would reject it as illegal.

Moving it into your browser

The server was the weakest part of the whole project. It slept when nobody played, so the first visitor of the day stared at a loading screen while it woke up.

So now there’s no server. The search is ported to TypeScript, the same ONNX network runs with onnxruntime-web (WebAssembly) in a Web Worker, and chess.js handles the rules. Your browser downloads about 20MB of network and 14MB of WebAssembly once, then keeps them cached.

A port like this is only worth it if it plays the same moves, so I checked it against the Python engine. On 30 random games, the TypeScript encoder produces exactly the same input planes as python-chess, and the network’s evaluations match to four decimals.

One thing didn’t carry over: the cost model for deepening. In Python, a lot of the time went into python-chess generating moves, which is why the first rung was predicted to cost 8x. In the browser the network is about 95% of the work, so a rung costs roughly its number of positions, and I retuned the factors to match.

It runs on a single thread for now. More threads need the page to be cross-origin isolated, and when I turned that on, the browser blocked the worker script itself. In testing, Easy answers in about 150ms at 1 ply, and Medium and Hard in about 0.9s at 2 ply.

As a bonus, the en passant bug went away. The game is replayed from its full move history now, so there’s no bare board to lose the en passant square from.

What I’d tell you if you’re starting

  • •Get a baseline that plays real games on day one. Matches are your only honest metric.
  • •Fix your labels before your architecture.
  • •Look for free Elo in the search before you train anything.
  • •Don’t expect self-play to beat a teacher that’s stronger than you.

Play the engine and tell me if it’s really 2000.

Comments (0)

0/1000
Loading comments...
Back to Blog
Published on September 22, 2026