| |
Five frontier LLMs disagree on 67% of 1k real-world fact-check claims
Five frontier large language models disagreed on verdicts for 63% of 1,000 real-world fact-check claims, with 23% showing disagreement spanning two or more verdict categories, indicating that the answer a user receives can depend on which model they consult. Models reported high confidence in their assessments despite disagreeing frequently, and individual confidence scores were poor predictors of disagreement, though the panel agreed more often on claims where all models expressed confidence. This disagreement represents a source of inaccuracy independent of any single model's benchmark performance.
Read Full Article →
← More Tech news