| |
A viral "GeoGuessr" prompt engineered for OpenAI's o3 model to identify photo locations was tested against a basic prompt using 200 benchmark images, and the results showed the elaborate prompt actually performed slightly worse than a simple "think carefully" instruction. The finding highlights how easily people can overestimate prompt engineering's effectiveness when models are already strong at a task, especially since models tend to confirm that tweaks are helpful when asked directly.
Read Full Article →
← More Tech news