Weaknesses
The title is not proper. It is really vague to me by framing something as “from a deep learning perspective”. Frankly speaking it could use more work to become concise and precise. There is perhaps no such thing known as “deep learning perspective” as many things in deep learning are either not unified or still in debate and there is a mixture of knowledge used in deep learning. For example, is this paper based on any statistical learning theories or optimization works? I would suggest using a better term like “policy learning” or “in using human feedback”. But, I will let the authors decide.
Before the discussion of power-seeking models, the discussion more or less relates to the underspecification issue [1]. There might be a lot of ways to define the word underspecification, but one explanation from me is simply as a situation our training objectives fail to specify what we really desire for. This happens not just recently in the most advanced models tuned with RLHF but ubiquitously on many other tasks. For example, we train a model to classify ImageNet images hopefully with features we use but the model could “hack” the task by using spurious features and ends up being adversarially vulnerable. Many short-cuts described in the reward hack look really similar to the underspecification issue. It would be nice to relate some discussions here to the broader underspecification discussion.
A minor point: there are many works, e.g. [2], showing lies / deceptions are detectable in LLMs. How does these works influence people to analyze AGI’s intentions or situation-awareness.
[1] D'Amour, A., Heller, K.A., Moldovan, D.I., Adlam, B., Alipanahi, B., Beutel, A., Chen, C., Deaton, J., Eisenstein, J., Hoffman, M.D., Hormozdiari, F., Houlsby, N., Hou, S., Jerfel, G., Karthikesalingam, A., Lucic, M., Ma, Y., McLean, C.Y., Mincu, D., Mitani, A., Montanari, A., Nado, Z., Natarajan, V., Nielson, C., Osborne, T.F., Raman, R., Ramasamy, K., Sayres, R., Schrouff, J., Seneviratne, M.G., Sequeira, S., Suresh, H., Veitch, V., Vladymyrov, M., Wang, X., Webster, K., Yadlowsky, S., Yun, T., Zhai, X., & Sculley, D. (2020). Underspecification Presents Challenges for Credibility in Modern Machine Learning. J. Mach. Learn. Res., 23, 226:1-226:61.
[2] Zou, A., Phan, L., Chen, S., Campbell, J., Guo, P., Ren, R., Pan, A., Yin, X., Mazeika, M., Dombrowski, A., Goel, S., Li, N., Byun, M.J., Wang, Z., Mallen, A., Basart, S., Koyejo, S., Song, D., Fredrikson, M., Kolter, Z., & Hendrycks, D. (2023). Representation Engineering: A Top-Down Approach to AI Transparency. ArXiv, abs/2310.01405.