Handle examples originated from the same board state #115
evg-tyurin
started this conversation in
Conceptual
Replies: 3 comments
|
I believe that having different labels for the same state has an effect such that the output gets averaged out. For example if one state has labels 1 and and -1, it will be pushed towards 0 instead. I don’t think the current implementation suffers from any issue as such. But do let me know if you notice any improvements by preprocessing. |
0 replies
|
Shouldn't we aim to get from NNet predicted reward values (v) as integers from fixed set of possible game results? Can it be arbitrary averaged float values from -1 to 1? |
0 replies
|
The NN predicts a winning probability in [0,1], or the expected reward in [-1,1]. |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
During self-play phase we usually collect different examples for the same board states. Should we preprocess such examples before optimizing the NNet? In the current implementation, we don't preprocess them so we train NNet and expect different output from the same input values. I think this may be wrong.
All reactions