| |
Researchers used mechanistic interpretability to analyze political censorship in Alibaba's Qwen 3.5 LLM, discovering it operates as a small, identifiable circuit in the model's weights rather than being baked into its fundamental knowledge. The censorship works by computing three internal directions (vectors) that determine whether content is PRC-sensitive, whether to refuse, and what style of response to use—allowing the model to route around factual knowledge it actually possesses rather than losing it. The researchers demonstrated they could disable this censorship by subtracting specific directions at particular layers, revealing that the underlying facts about sensitive topics like Tiananmen Square and Falun Gong remain in the model's pretraining but are actively suppressed through learned behavior patterns.
Read Full Article →
← More Tech news