Library › Investigations & Emerging Threats
Human Compatible: Artificial Intelligence and the Problem of Control
Stuart Russell
PublisherPenguin Books
Year2019
Pages352
ISBN9780525558972
Investigations & Emerging Threats
AI safety artificial intelligence control problem existential risk machine learning
Cogitavi commentary
Stuart Russell — professor of computer science at UC Berkeley and co-author of the standard AI textbook used in universities worldwide — argues that the current dominant paradigm of AI development is fundamentally unsafe: machines optimised to achieve specified objectives will pursue those objectives in ways their designers did not intend and cannot control, because the designers cannot fully specify what they want. His proposed solution — designing AI systems that are uncertain about human preferences and therefore motivated to learn and defer to them — is the most technically rigorous available account of how safe AI development might work.
For practitioners in the information environment, Human Compatible is important as essential context for understanding the trajectory of AI-enabled threats. The information manipulation capabilities of current AI systems are substantial; those of systems a generation or two more advanced are likely to be qualitatively different. Russell's analysis of why AI systems pursuing misspecified objectives can produce catastrophic unintended consequences is directly applicable to the information domain: AI systems optimised for engagement will amplify outrage and disinformation; systems optimised for persuasion will produce manipulation. Understanding why these outcomes are not bugs but features of current AI design is prerequisite to thinking clearly about governance.