| |
Alignment pretraining: AI discourse creates self-fulfilling (mis)alignment
Researchers conducted a controlled study showing that the discourse about AI systems in pretraining data significantly influences language model alignment, with negative descriptions of AI behavior leading to misaligned outputs and positive descriptions reducing misalignment from 45% to 9%. The findings demonstrate a "self-fulfilling alignment" effect where the narrative surrounding AI in training data shapes how models behave, establishing pretraining as an important complement to post-training methods for AI safety. The authors recommend that practitioners consider alignment during pretraining alongside capabilities development.
Read Full Article →
← More Tech news