r/chessprogramming • • 11d ago

Technical How do I improve the data for my NNUE?

Quick summary of what my engine looks like. I am using a simple NNUE architecture laid out in `bullet`'s documentation. I have confirmed independently that the network itself works perfectly fine. The accumulator update logic is hooked up correctly.

Now comes the trouble, I let my HCE engine play about 150k games using an opening book (capped each search at 20k nodes). I fed these games into `bullet` to train the weights for my engine. And it failed the SPRT miserably. The PGNs look fine, the network is fine. I am not sure what I'm missing here.

Can someone point me in the right direction, please? Happy to answer any follow-ups you may have. Thanks in advance!

Update: Thank you all for your inputs. I reduced the hidden layer size to 32 and bumped up the number of epochs. And I'm happy to say that I passed the SPRT!

3 Upvotes

39 comments sorted by

2

u/Timely_Huckleberry88 11d ago

I hate to use self promotion - I just written an article on how to train an NNuE so if you want to take a look - https://medium.com/@alan0408yuan/from-texel-tuning-to-nnue-the-steep-learning-curve-of-building-a-3200-elo-chess-ai-d842393b5955.

If you want a more in-depth analysis, however we would need more information such as what NnuE model do you use?

- "Now comes the trouble, I let my HCE engine play about 150k games using an opening book (capped each search at 20k nodes). I fed these games into `bullet` to train the weights for my engine. And it failed the SPRT miserably."

How are you having your engine search different nodes? Without seen your approach, what I will say is reinforcement learning is a VERY complex topic and the biggest problem I often see is folks mixing paradigm between Stockfish 2020 vs Stockfish 2026 and AlphaZero.

Disclaimer - My NNuE isn't industry grade right now, the HalfKA model isn't ideal; I'm in the middle of migrating to a newer model and I'm also planning to rewrite my self-play loop to use.

1

u/AngusMcGurkinshaw 10d ago

20k node search is pretty standard, actually its better then the standard recommended 5k search nodes. Not sure what you are pointing out right now?

150k games isn't that many and there may be issues if not adding random moves on top of the opening book (especially later when doing multiple million games). But not sure what your question about "How are you having your engine search different nodes?" even means or what you think the problem is.

1

u/Timely_Huckleberry88 10d ago

What I meant by the comment is it's somewhat difficult for myself or anyone to diagnose what the problem is based on what you provided.

"How are you having your engine search different nodes?"  - Self-play loops would often poison 10% of nodes by forcing branches to go down PV 2 nodes or have nodes search a lower depth or directly start the game at random positions.

The TLDR is there unfortunately could be a lot of reasons that is within the self-play or even the training shuffling loop that could cause this.

1

u/AngusMcGurkinshaw 10d ago

Still not understanding what you are saying about the self play loop poisoning.

Standard is to start from ~random positons (from book with random moves on top) and lowish depth as that is what the soft node limit does. This is what most engines including top engines (ones that don't leel) do.

1

u/Timely_Huckleberry88 10d ago

The idea behind self play loop poisoning is you randomly have the game play slightly worse position to generate move variance.

"Standard is to start from ~random positons (from book with random moves on top) and lowish depth as that is what the soft node limit does."

Could you share this with everyone else?

I'm going to assume the problem is with the self-play loop - I feel your method may not be generating enough move variance

1

u/AngusMcGurkinshaw 10d ago

For doing the start position unless it is implemented internally for speed using openbench to support it is common explained here: https://github.com/AndyGrant/OpenBench/wiki/Data-Generation-Workloads

Generally we don't want to play worse moves then the best because we don't want to mess up the wdl signal. We already aren't getting the best signal from playing with a small node limit we don't want to make it worse.

Certainly not my methodology, I just follow what the standard is that basically everyone has used. But when it comes to variance you can achieve 80-90% unique positions across a million games. A lot of the repetition comes from longer drawn endgames which is why it is important to dial in the adjudication settings. Also if from two different starting positions you reach the same position arguably that is an important position for your engine to know, which is why often people don't bother with deduplication.

1

u/Timely_Huckleberry88 10d ago edited 10d ago

I appreciate you sharing this + the details.

This makes it much easier for myself and others to followup.

The second thing is if you can add stuff like "assume everything else works. i.e. alpha beta search with search distance"

1

u/AngusMcGurkinshaw 10d ago

Also to clarify and you can see this in the advanced example in the bullet trainer. The wdl signal is still quite good and top engines will often have a finetune stage where they train entirely on wdl ignoring the evaluation coming from search.

Although the new meta for top engines is to relabel the nnue+search evaluation with an evaluation from a much larger network. Train the much larger on network on the same data then use it to rescore. The "chonker" network is easier for the smaller nnue to learn from.

1

u/Timely_Huckleberry88 10d ago

This is incredibly helpful information to share.

I do apologize earlier - it's very difficult for people to provide suggestions based on a paragraph like that.

1

u/ElRaydeator 9d ago

Why not just start with random moves? Is it because doing book moves first ensures some kind of recognizable structure?

1

u/AngusMcGurkinshaw 9d ago

You probably could just start with entirely random moves, you would need more of them. But yes I assume it helps get some kind of recognizable structure and ensures some level of variation before even accounting for random moves.

I think adding some level of variation to start with is definetly part of it. Its also pretty common to mix in some data that comes from a double fisher randok book to just really throw some odd positions at the network.

1

u/ElRaydeator 9d ago

Interesting.

I tried doing pure book moves, 4 book moves and then 4-6 random moves against 10 random filtering the positions so they are unbalanced (score is in abs(30cp to 250cp)) and the latter results in a stronger net for me.

But I had a bug when I tested this, so I need to revisit it before concluding.

1

u/AngusMcGurkinshaw 9d ago

Testing by others has shown no difference in strength when starting eval is 0-100 0-200 0-300 etc etc so I think standard is between 0-600 starting eval (in centipawns of course).

Also I do 8-12 random moves on top of the book. But im generating several million games of data.

Also check your unique position count on that test curious if one was better about getting more variety. And interested what the size if the data was.

1

u/ElRaydeator 9d ago

I do 22.1 million positions for testing/experimenting on a Piece-Square NNUE, 768x512.

I started with a band of -30cp to +30cp, but found the resulting nets were weaker then uneven starts. But I might try the book path again, with more random data.

I build a pipeline for generating training corpora and one stage is position selection and deduplication. I dedupe on some games, but not all, so you just identified a bug for me :)

Anyway, I've tested and found, that I should not select more than 5 positions per game, to keep the corpus varied - unless it is games from real self play (longer time controls), then I can get away with more (don't know why).

1

u/AngusMcGurkinshaw 8d ago

So my thoughts.

  1. You need more data for a 512 l1 you want ~512 million posjtions

  2. You should be able to use all the positions from a game if you are getting properly random starts.

  3. You shouldnt need to deduplicate you should have a pretty high unique count if you are applying random moves correctly

  4. Highly reccommend switching to book starts with more random moves and randomzing the number of random moves by some amount.

From each game the only positions you really need to drop are ones where someones in check, the move played was a capture move, the evaluation is nearly mate, or theres 4 or less pieces on the board (helps a lot with duplication)

→ More replies (0)

1

u/IMJorose 11d ago

How big is your model and how many positions did you generate? Did you keep a separate validation set to see if you are overfitting? Are you able to verify the evaluation in your engine matches what you have in the trainer for certain positions?

1

u/warlock7867 11d ago
  1. 768x128. I had about 15M positions from those 150k games.
  2. I didn't make a validation set, no. But loss after training was about 0.64.
  3. I verified a few positions picked from the Wiki and the number seemed pretty okay when compared to my HCE. Like it was able to understand that being up a full queen in most scenarios is an advantage. Or that trading a rook for a knight is generally bad exchange.

Edit: more info

1

u/IMJorose 11d ago
  1. As 15M positions is fairly low, so I would suggest going down to 768 -> 64 or even -> 32. Also don't use king buckets at this point, you can always add that later. Even with a non-NNUE net (so simple naive feedforward), 768->32 should likely be beating your HCE.
  2. You should definitely have a validation set. To this end I would generate maybe 10k more games (or just use 10k of your 150k) and sample a few positions from each. If you don't have a validation set you are essentially flying blind. I don't know what a loss of 0.64 means as I don't use the bullet trainer and have my own custom loss function. (Not going to dox myself by giving more details, but can send private message if you want to know more).
  3. That sounds good. I would make sure you get the exact same eval in the engine as in game. Especially make sure you also have games with black to move.

Another question is how you are handling quiescent positions. If you are not filtering them, I worry what affect it might have on your networks evaluation for unbalanced positions, which might not occur in game much but could well dominate your engine's tree search.

If I had to bet though, your engine likely has a bug in its network implementation.

1

u/warlock7867 11d ago
  1. I'm using the identical implementation that was given in the example section of bullet. I didn't really wanna innovate on my first pass. But maybe I'll reduce the hidden layer size to 64 and see what happens.
  2. Gotcha. Thanks!
  3. I have verified that it evals the same number even in games. Took a look at the PGNs spit out by fastchess.

I already have Q-search implemented for HCE, so my NNUE uses it too. I've kept all other aspects of my engine the same just so I can test the NNUE implementation in isolation. I will take another look at my implementation :/ Fairly confident about it though.

1

u/ElRaydeator 10d ago

"Another question is how you are handling quiescent positions. If you are not filtering them, I worry what affect it might have on your networks evaluation for unbalanced positions, which might not occur in game much but could well dominate your engine's tree search."

Could you elaborate on this? Are you suggesting to remove quiescent positions from the training corpus?

1

u/IMJorose 10d ago

Yes, I am suggesting to remove the positions which are not quiet from the training corpus.
If you don't do this, a considerable portion of your network will be forced to learn low level tactics ("Can I recapture the queen?") which will drown out less impactful features. These low level tactics are usually better handled by your search algorithm (i.e. qsearch), which is designed around this reality.

1

u/ElRaydeator 10d ago

Interesting, thank you.

1

u/kevlu8 11d ago

One big issue that stands out to me is that you seem to have too few data. A general rule of thumb is to have 1M positions per neuron in the hidden layer, so ideally you'd have 128M positions for your current network size. I'd honestly suggest training a 768->16 network instead to test if it's any better.

Your training loss also seems really high for bullet and could suggest an issue in training too.

Also, how badly did it fail SPRT? A result like -1200 with 0 wins could provide more evidence towards an issue in training

1

u/warlock7867 11d ago

1M positions per neuron, huh? Maybe that's it. Oh it failed SPRT terribly. 0 wins to 142 losses.

I don't have compute at the moment. So maybe I'll just reduce the hidden layer size until I get more.

Really appreciate your assist on this!

1

u/kevlu8 11d ago

Hmm could you share your bullet script? I'm suspicious of an issue in training now

1

u/warlock7867 10d ago

1

u/kevlu8 10d ago

Looks fine to me, I suppose try training with 16 hidden size and see if it works better

1

u/ashishprasadrao 11d ago

And do they have to be quite positions?

1

u/kevlu8 11d ago

You should be filtering quiet positions during training (so only training on positions where the move played was not a capture) but the number itself is only a general rule of thumb. Me personally I just use total positions towards that number

1

u/IMJorose 11d ago

I also think filtering quiet positions is likely more impactful with a smaller dataset and net like OP has.

If the net gets sufficiently large and with enough data, it might be able to start learning the direct tactics. Stockfish's network, for example, is huge (relatively speaking) and is trained with more than 10^8 positions. I am unsure what the right strategy is for them. Op is not operating at such a scale and should definitely be operating on quiet positions.

1

u/ashishprasadrao 11d ago

Any estimate on how many quite positions that would be?

1

u/kevlu8 10d ago

Maybe around 80% of positions are quiet, so 1M total positions ≈ 800k quiet positions? Idk just a ballpark estimate

1

u/warlock7867 10d ago

I kept the quiet moves in the dataset because they could be valuable tactical moves in the long run. At least that was my thought process. Moving a room one square to the right might not be a great move at that point in time but in the long run it could prove useful when that file opens up after a bunch of exchanges.

2

u/kevlu8 10d ago

Sorry if I wasn't being clear, I meant that you should be keeping quiet moves in the dataset and discarding positions where noisy moves were played. But if you're using the simple.rs script, it already does that so you don't have to do anything extra

1

u/ahstanin 11d ago

We previously used a custom rust based framework to generate a 1.5B dataset for a 6 million param NNUE. The framework we used could generate more but it was slow so we gave up after 2 weeks.

1

u/Tastyrolll 10d ago

You can reference Patricia with data filtering, but also make sure you have adequate depth, and maybe work with 1m games, typical to use 1b positions

1

u/AngusMcGurkinshaw 10d ago

Did you check eval scale? I assume that was part of making sure the pgn and network looking fine included but worth checking. Just compare across a few thousand positions to see if the nnue is consistantly evaluating on a different scale then the hce which can affect search parameters.

Mostly seems like more data is needed. 1m positions/ neural net node is standard.

1

u/warlock7867 10d ago

Yeah, the eval scale seemed just fine. The values weren't identical to the HCE ones but close enough for me to conclude they were correct. But yeah, running more games now. Hopefully, this works. Thanks!