Principles first, preference learning second
Constitutional AI, by Yuntao Bai and colleagues, uses a list of principles as the human-provided oversight for a two-stage training process. In supervised learning, an initial model samples answers, critiques them against those principles, revises them, and is fine-tuned on the revisions. During the reinforcement-learning stage, a model compares candidate answers; those comparisons train the preference model used as the reward signal.
The intended refusal explains itself
The authors report a harmless but non-evasive assistant that responds to harmful queries by explaining its objections. Their evaluations use crowd judgments and preference-model scores for helpfulness and harmlessness. In the reported experiments, harmlessness scores rose over additional revision rounds; the authors also found that critique-and-revision could reduce evasiveness compared with optimizing harmlessness alone.
The procedure makes the constitution part of the trained behavior: it informs both the model's self-critiques and later comparisons. Human judgments cover selected evaluations, and the principle list reflects choices about what the assistant should do. An implementation based on a different list would need its own checks for harmful compliance and over-refusal.
Bibliography
- Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., Chen, A., Goldie, A., Mirhoseini, A., McKinnon, C., Chen, C., Olsson, C., Olah, C., Hernandez, D., Drain, D., Ganguli, D., Li, D., Tran-Johnson, E., Perez, E., Kerr, J., Mueller, J., Ladish, J., Landau, J., Ndousse, K., Lukosuite, K., Lovitt, L., Sellitto, M., Elhage, N., Schiefer, N., Mercado, N., DasSarma, N., Lasenby, R., Larson, R., Ringer, S., Johnston, S., Kravec, S., Showk, S. E., Fort, S., Lanham, T., Telleen-Lawton, T., Conerly, T., Henighan, T., Hume, T., Bowman, S. R., Hatfield-Dodds, Z., Mann, B., Amodei, D., Joseph, N., McCandlish, S., Brown, T., & Kaplan, J. (2022). Constitutional AI: Harmlessness from AI Feedback. arXiv:2212.08073v1. https://arxiv.org/pdf/2212.08073
