Bibliography
Foundational AI Safety
Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., & Mane, D. (2016). Concrete Problems in AI Safety. arXiv preprint arXiv:1606.06565.
Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., DasSarma, N., … & Kaplan, J. (2022). Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback. arXiv preprint arXiv:2204.05862.
Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., … & Kaplan, J. (2022). Constitutional AI: Harmlessness from AI Feedback. arXiv preprint arXiv:2212.08073.
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., & Amodei, D. (2017). Deep Reinforcement Learning from Human Preferences. Advances in Neural Information Processing Systems, 30.
Irving, G., Christiano, P., & Amodei, D. (2018). AI Safety via Debate. arXiv preprint arXiv:1805.00899.
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., … & Lowe, R. (2022). Training Language Models to Follow Instructions with Human Feedback. Advances in Neural Information Processing Systems, 35.
Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., & Finn, C. (2023). Direct Preference Optimization: Your Language Model is Secretly a Reward Model. arXiv preprint arXiv:2305.18290.
Alignment Measurement and Evaluation
Askell, A., Bai, Y., Chen, A., Drain, D., Ganguli, D., Henighan, T., … & Kaplan, J. (2021). A General Language Assistant as a Laboratory for Alignment. arXiv preprint arXiv:2112.00861.
Bowman, S. R. (2023). Eight Things to Know about Large Language Models. arXiv preprint arXiv:2304.00612.
Lin, S., Hilton, J., & Evans, O. (2022). TruthfulQA: Measuring How Models Mimic Human Falsehoods. Proceedings of ACL, 3214–3252.
Perez, E., Ringer, S., Lukosuite, K., Nguyen, K., Chen, E., Heiner, S., … & Kaplan, J. (2022). Discovering Language Model Behaviors with Model-Written Evaluations. arXiv preprint arXiv:2212.09251.
Sharma, M., Tong, M., Korbak, T., Duvenaud, D., Askell, A., Bowman, S. R., … & Perez, E. (2023). Towards Understanding Sycophancy in Language Models. arXiv preprint arXiv:2310.13548.
Scalable Oversight and Superalignment
Burns, C., Ye, H., Klein, D., & Steinhardt, J. (2022). Discovering Latent Knowledge in Language Models Without Supervision. arXiv preprint arXiv:2212.03827.
Greenblatt, R., Shlegeris, B., Sachan, K., & Roger, F. (2024). AI Control: Improving Safety Despite Intentional Subversion. arXiv preprint arXiv:2312.06942.
Leike, J., Krueger, D., Everitt, T., Martic, M., Maini, V., & Legg, S. (2018). Scalable Agent Alignment via Reward Modeling: A Research Direction. arXiv preprint arXiv:1811.07871.
Representation Engineering and Mechanistic Interpretability
Li, K., Patel, O., Vieillard, N., Liu, H., Bolukbasi, T., Ranzato, M., & Weston, J. (2024). Inference-Time Intervention: Eliciting Truthful Answers from a Language Model. Advances in Neural Information Processing Systems, 36.
Nanda, N., Chan, L., Lieberum, T., Smith, J., & Steinhardt, J. (2023). Progress Measures for Grokking via Mechanistic Interpretability. ICLR 2023.
Zou, A., Phan, L., Chen, S., Campbell, J., Guo, P., Ren, R., … & Fredrikson, M. (2023). Representation Engineering: A Top-Down Approach to AI Transparency. arXiv preprint arXiv:2310.01405.
Mathematical Foundations
do Carmo, M. P. (1992). Riemannian Geometry. Birkhauser.
Hohfeld, W. N. (1919). Fundamental Legal Conceptions as Applied in Judicial Reasoning. Yale University Press.
Nakahara, M. (2003). Geometry, Topology and Physics. CRC Press.
Nickel, M., & Kiela, D. (2017). Poincare Embeddings for Learning Hierarchical Representations. Advances in Neural Information Processing Systems, 30, 6338–6347.
Von Neumann, J., & Morgenstern, O. (1944). Theory of Games and Economic Behavior. Princeton University Press.
The Geometric Series
Bond, A. H. (2026a). Geometric Methods for Complex Systems. [Series Book 1.]
Bond, A. H. (2026b). Geometric Ethics: The Moral Manifold and the Geometry of Value. [Series Book 3.]
Bond, A. H. (2026c). Geometric Reasoning: Heuristic Fields, Geodesic Deviation, and the Structure of Informed Search. [Series Book 2.]
Bond, A. H. (2026d). Geometric Cognition: Attention, Memory, and the Architecture of Thought. [Series Book 4.]
Bond, A. H. (2026e). Geometric Communication: From Whale Song to Cuneiform, the Manifold of Meaning. [Series Book 5.]
Bond, A. H. (2026f). Geometric Medicine: Clinical Decision-Making on the Nine-Dimensional Manifold. [Series Book 6.]
Bond, A. H. (2026g). Geometric Education: Learning, Assessment, and the Topology of Growth. [Series Book 7.]
Bond, A. H. (2026h). Geometric Economics: Markets, Equilibrium, and the Curvature of Exchange. [Series Book 8.]
Bond, A. H. (2026i). Geometric Law: Rights, Symmetry, and the Gauge Structure of Justice. [Series Book 9.]
Bond, A. H. (2026j). Geometric Politics: Democracy, Power, and the Fiber Bundle of Collective Choice. [Series Book 10.]
Behavioral Economics and Decision Theory
Allais, M. (1953). Le Comportement de l’Homme Rationnel devant le Risque. Econometrica, 21(4), 503–546.
Ellsberg, D. (1961). Risk, Ambiguity, and the Savage Axioms. Quarterly Journal of Economics, 75(4), 643–669.
Kahneman, D., & Tversky, A. (1979). Prospect Theory: An Analysis of Decision under Risk. Econometrica, 47(2), 263–291.
Rawls, J. (1971). A Theory of Justice. Harvard University Press.
Ross, W. D. (1930). The Right and the Good. Clarendon Press.
Historical AI Safety
Asimov, I. (1950). I, Robot. Gnome Press.
Yudkowsky, E. (2001). Creating Friendly AI 1.0. Machine Intelligence Research Institute.
Empirical Foundations
Bond, A. H. (2026k). Measuring AGI: Five Cognitive Dimensions, Five Models, 25 Subtasks. [agi-hpc benchmark suite documentation.]
This bibliography is representative, not exhaustive. For complete references for specific results, see the footnotes in the relevant chapters.