| |
Harm Laundering in GPT Models: Gender Discrimination Transformed Rather Than
Researchers found that safety improvements in GPT models across generations don't actually eliminate gender discrimination but rather transform it into subtler forms—a phenomenon they call "harm laundering." While explicit sexual violence in women-directed outputs decreased from GPT-2 to GPT-4, women-directed completions became narrower in topic diversity and representational range compared to men-directed ones, with toxicity classifiers failing to detect these shifts. The study demonstrates that declining toxicity scores are insufficient measures of actual harm reduction in large language models.
Read Full Article →
← More Tech news