AI Alignment Problem: Matching AI Goals to Human Values

The AI alignment problem is the challenge of ensuring AI systems reliably pursue human intentions and values, rather than literally interpreting commands in ways that can lead to unintended or harmful outcomes, especially as AI becomes more complex and powerful.

Sources

Open the full topic