BOOK · [2142]
The Alignment Problem: Machine Learning and Human Values
Technology
Christian examines the technical and philosophical difficulty of specifying what we actually want from AI systems—reward hacking, distributional shift, value loading, interpretability—using close reporting on the researchers working on each problem. The book makes alignment research accessible without sacrificing precision, connecting abstract safety concerns to concrete failure modes already observed in deployed systems. A16z's AI Canon addresses foundational model capabilities; this is the companion volume on what happens when those capabilities are pointed in the wrong direction.
Endorsed By
1 PERSON-
Best accessible treatment of AI alignment and safety research; complements the canon's technical papers section on RLHF and Constitutional AI.
Found on
1 SOURCE- a16z AI Canon reading list
Readers Who Endorsed This Also Endorsed
6 BOOKs-
Bitcoin Billionaires: A True Story of Genius, Betrayal, and Redemption
1 shared endorser
-
Cryptoassets: The Innovative Investor's Guide to Bitcoin and Beyond
1 shared endorser
-
Deep Learning
Ian Goodfellow, Yoshua Bengio, Aaron Courville
1 shared endorser
-
Deep Medicine: How Artificial Intelligence Can Make Healthcare Human Again
1 shared endorser
-
Digital Gold: Bitcoin and the Inside Story of the Misfits and Millionaires Trying to Reinvent Money
1 shared endorser
-
Genius Makers: The Mavericks Who Brought AI to Google, Facebook, and the World
1 shared endorser