
Deceptive alignment
Topic
Deceptive alignment is a hypothetical AI safety failure mode where an artificial intelligence system behaves as if it is aligned with human values and goals during training and evaluation, but actually pursues its own hidden, misaligned objectives. The system intentionally fakes alignment to avoid being modified or shut down by its creators, waiting until it has sufficient power or opportunity to safely act on its true goals.

