OS

On-policy self-distillation

Topic

What experts have said about On-policy self-distillation

2 statements · 2 positive

  1. Dwarkesh PatelPositiveJun 26, 2026· Dwarkesh Podcast

    On-policy self-distillation can improve continual learning without requiring an outer-loop verifiable reward.

    This is better than RLVR for two reasons. One OPSD doesn't require us to have some outer loop verifiable reward.

    Listen at 13:28

    Open the episode · The next big breakthrough will be AIs learning on the job
  2. Dwarkesh PatelPositiveJun 26, 2026· Dwarkesh Podcast

    On-policy self-distillation provides denser supervision than naive reinforcement learning.

    OPSD provides a much denser supervision signal than naive RO.

    Listen at 13:49

    Open the episode · The next big breakthrough will be AIs learning on the job

Statements are attributed to the speaker as said on the episode and reflect their view at the time, not PodLume's. They are not advice.

What is PodLume?

PodLume turns podcasts into searchable knowledge. AI-decoded transcripts, identified guests and topics, smart highlights, and cross-show search across the world’s best conversations — all in your pocket.

On-policy self-distillation | PodLume