Timman vs Short: When Six Chess Models Miss the Move
A real GPU-backed failure study: six custom neural evaluators agree on the wrong capture, while the Stockfish safety gate preserves a famous attacking finish.
Reproducible analysis run complete
6 saved models on NVIDIA GeForce RTX 5080 · Stockfish 18 · 250,000 nodes per checked line.
Why this game?
The pieces around the king become the trap
This game is a useful stress test because material captures look attractive while a forcing passed-pawn and mating attack demand deeper calculation. The result is honest and instructive: all six lightweight evaluators repeat the same mistake, and the independent engine gate catches it every time.
- White
- Jan Timman
- Black
- Nigel Short
- Event
- Tilburg
- Result
- 1–0
Interactive comparison
Choice one: pass the pawn
White to move on move 21. All six neural evaluators choose the immediate capture 21.Qxc4; the game move 21.e6 is Stockfish’s first choice.
3n1rk1/1pp3pp/8/p1qNPpN1/2b2Qn1/6P1/P3PPBP/3R2K1 w - - 2 21Model candidates
Choose a viewpoint
V6 Active: Qxc4
Measured GPU output. This evaluator ranked the visible capture first, but the independent engine gate finds that choice losing.
Model top three: Qxc4 · Qxg4 · Nxc7
Stockfish 18 safety gate
CompleteGate stopped Qxc4: 776 cp behind the best line. Stockfish keeps 21.e6 at about +4.09 for White in this fixed-node snapshot.
- Engine move
- e6
- Depth
- 15
- Node budget
- 250,000
- Evaluation
- +4.09
Verification line
21.e6 Re8 22.Qxf5 Qxf2+ 23.Qxf2 Nxf2 24.Kxf2 Nxe6 25.Nxe6 Rxe6
Position lesson
The advanced e-pawn matters more than the bishop on c4. After 21.e6, promotion threats and tactical pressure keep Black tied down. The material-first capture lets the initiative disappear.
Interactive comparison
Choice two: one square from promotion
White to move on move 25. Every model takes the knight with 25.Qxg4; the played move 25.e7 preserves the forcing promotion threat.
5rk1/2pR2pp/2p1P3/p4pN1/5Qn1/q5P1/P3PP1P/6K1 w - - 0 25Model candidates
Choose a viewpoint
V6 Active: Qxg4
Measured GPU output. The capture is legal and looks profitable, but the engine gate shows that it gives away White’s winning attack.
Model top three: Qxg4 · Qxc7 · Rxc7
Stockfish 18 safety gate
CompleteGate stopped Qxg4: 1,226 cp behind the best line. The quiet-looking 25.e7 keeps a decisive promotion-and-mate attack.
- Engine move
- e7
- Depth
- 16
- Node budget
- 250,000
- Evaluation
- +7.76
Verification line
25.e7 h6 26.Qxf5 Nf6 27.exf8=R+ Qxf8 28.Qe6+ Kh8 29.Rd8
Position lesson
A passed pawn on the seventh rank changes the calculation. The tempting knight capture wins material now, but 25.e7 forces Black to answer promotion and keeps the king exposed.
Interactive comparison
Choice three: find the mating route
White to move on move 27. All six models choose 27.Qxg4. Stockfish and the published game find 27.Nf7+, beginning a forced mate in four.
4r2k/2pRP1pp/2p5/p4pN1/2Q3n1/q5P1/P3PP1P/6K1 w - - 3 27Model candidates
Choose a viewpoint
V6 Active: Qxg4
Measured GPU output. The model’s capture misses the forcing check sequence; the engine gate rejects it and preserves the published mating line.
Model top three: Qxg4 · Rxc7 · Qxc6
Stockfish 18 safety gate
CompleteGate stopped Qxg4. Stockfish finds mate in 4 with the same forcing sequence played in the game. The large numeric loss uses the documented mate comparison value.
- Engine move
- Nf7+
- Depth
- 44
- Node budget
- 250,000
- Mate
- M4
Verification line
27.Nf7+ Kg8 28.Nh6+ Kh8 29.Qg8+ Rxg8 30.Nf7#
Position lesson
Forcing moves outrank material. Nf7+ sends the king toward g8, Nh6+ uncovers the queen, and Qg8+ forces the rook to occupy its king’s last flight square.
Published finish
The forcing line
This is the published game finish. The queen sacrifice forces Black’s rook onto the final flight square before the knight returns with mate.
Review method
Source, models, then verification
Start with a published game
The position remains connected to the complete ChessBox game record and PGN.
Run six saved models on GPU
Every legal move is ranked by each scalar evaluator with the same 30% hybrid scoring formula.
Apply the Stockfish gate
Every distinct model choice and the played move receive an independent fixed-node verification.
What this review can and cannot show
Clear boundaries keep a useful experiment from becoming a bigger claim than the evidence supports.
- These checkpoints are lightweight scalar position evaluators, not move-policy, search, or explanation models.
- Their inputs encode piece placement only; they do not receive move history, clocks, repetition state, castling rights, or en-passant state.
- The profile labels describe intended training behavior, not proven human-like personalities or internal reasoning.
- Stockfish used a fixed 250,000-node budget per line with two threads; this is a reproducible snapshot, not a claim of perfect play.
- Two-thread engine search can produce small score, depth, node-count, and principal-variation differences between reruns.
- This review measures three positions from one game. It does not establish Elo, general playing strength, or unseen-data accuracy.
- The prose explains observable chess consequences, not hidden neural-model reasoning.
Continue learning
Put the position on your own board
Start from the exact FEN and test whether you can find and convert the idea against a human-style bot.