Are AI Labs Pelicanmaxxing?
# Summary
A researcher tested whether AI labs have optimized their models specifically for the famous "pelican on a bicycle" benchmark by generating 1,008 SVG images across seven frontier models using a grid of 48 different animal-vehicle combinations. The experiment used an LLM judge to score image quality and analyzed whether the pelican-bicycle prompt showed suspiciously better performance than similar prompts with different animals or vehicles. The results of this informal investigation reveal whether AI labs may be "pelicanmaxxing"—optimizing specifically for this viral benchmark.
Read Full Article →