[1] GABRIEL I,GHAZAVI V. The Challenge of Value Alignment: From Fairer Algorithms to AI Safety[EB/OL].(2021-01-15)[2026-05-30].arXiv:2101.06060. https://arxiv.org/abs/2101.06060. [2] RUSSELL S.Human-Compatible Artificial Intelligence[M]//MUGGLETON S,CHATER N(eds.). Human-like Machine Intelligence. Oxford: Oxford University Press,2021. [3] GABRIEL I. Artificial Intelligence,Values,and Alignment[J].Minds and Machines,2020,30. [4] 柏拉图.游叙弗伦;苏格拉底的申辩;克力同[M].严群,译.北京:商务印书馆,1983. [5] 袁旭亮.人工智能价值对齐的伦理挑战及其消解路径[J].伦理学研究,2024(6). [6] YUDKOWSKY E.Coherent Extrapolated Volition[R/OL].San Francisco: Singularity Institute for Artificial Intelligence,2004. [7] CASPER S,DAVIES X,SHI C,et al. Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback[EB/OL].2023[2026-06-30].Transactions on Machine Learning Research,2023. arXiv:2307. 15217v2. https://arxiv.org/abs/2307.15217. [8] DAI J,FLEISIG E. Mapping Social Choice Theory to RLHF[EB/OL]//ICLR Workshop on Reliable and Responsible Foundation Models2024.(2024-04-19)[2026-05-30].arXiv: 2404.13038. https://arxiv.org/abs/2404.13038. [9] FRANKFURT H G.Freedom of the Will and the Concept of a Person[J].The Journal of Philosophy,1971,68(1). [10] 国家互联网信息办公室,中华人民共和国国家发展和改革委员会,中华人民共和国教育部,等.生成式人工智能服务管理暂行办法[EB/OL].(2023-07-13)[2026-05-30].https://www.cac.gov.cn/2023-07/13/c_1690898327029107.htm. [11] ARROW K J.Social Choice and Individual Values[M].New York: John Wiley & Sons,1951. [12] ECKERSLEY P. Impossibility and Uncertainty Theorems in AI Value Alignment: Or Why Your AGI Should Not Have a Utility Function[EB/OL].2018[2026-06-30].arXiv: 1901.00064. https://arxiv.org/abs/1901.00064. [13] SHAFER-LANDAU R.Moral Realism: A Defence[M].New York:Oxford University Press,2003. [14] GABRIEL I,MANZINI A,KEELING G,et al. The Ethics of Advanced AI Assistants[R/OL].London: Google DeepMind,2024[2026-05-30].arXiv: 2404.16244. https://arxiv.org/abs/2404.16244. [15] 闫坤如.人工智能对齐的哲学困境及其出路[J].学术研究,2026(4). [16] 张凌寒. 生成式人工智能的法律定位与分层治理[J].现代法学,2023,45(4). [17] SIMON H A.Rational Decision Making in Business Organizations[J].American Economic Review,1979,69(4). [18] BROPHY M.Wide Reflective Equilibrium in LLM Alignment: Bridging Moral Epistemology and AI Safety[J].Ethics and Information Technology,2026,28(2). [19] 陈夕朦,屈彦璋,周鹏鹏,等. “人在环路”AI价值对齐的可行性与合理性[J].自然辩证法研究,2026,42(1). |