| |
Researchers benchmarked various large language models on their ability to play Magic: The Gathering, finding that GPT-5.5 (medium) achieved the highest score of 95.4, while performance varied significantly across models. The study revealed that LLMs struggle with rule enforcement and accurate game state tracking, despite being better at evaluating legality than executing legal turns, highlighting challenges in complex decision-making tasks requiring strict adherence to game rules.
Read Full Article →
← More Tech news