OR
Off-policy reinforcement learning
Topic
What experts have said about Off-policy reinforcement learning
1 statement · 1 negative
Training on unreachable off-policy states wastes model capacity.
“if the current model is looking at states that it would never reach, then it's kind of wasting capacity”
Open the episode · Eric Jang – Building AlphaGo from scratchListen at 2:10:51
Statements are attributed to the speaker as said on the episode and reflect their view at the time, not PodLume's. They are not advice.
