AR
AI reward-hacking mitigation
Topic
What experts have said about AI reward-hacking mitigation
1 statement · 1 positive
Ordinary safety work and transparency could plausibly suffice to manage reward hacking.
“it's pretty plausible that we end up in a world where sort of like really mundane bullshit is sufficient”
Open the episode · Ryan Greenblatt – Human level AIs might build runaway superintelligences by 2032Listen at 2:02:53
Statements are attributed to the speaker as said on the episode and reflect their view at the time, not PodLume's. They are not advice.
