OS
On-policy self-distillation
Topic
What experts have said about On-policy self-distillation
2 statements · 2 positive
On-policy self-distillation can improve continual learning without requiring an outer-loop verifiable reward.
“This is better than RLVR for two reasons. One OPSD doesn't require us to have some outer loop verifiable reward.”
Open the episode · The next big breakthrough will be AIs learning on the jobListen at 13:28
On-policy self-distillation provides denser supervision than naive reinforcement learning.
“OPSD provides a much denser supervision signal than naive RO.”
Open the episode · The next big breakthrough will be AIs learning on the jobListen at 13:49
Statements are attributed to the speaker as said on the episode and reflect their view at the time, not PodLume's. They are not advice.
