Show HN: I put a $2.43 necklace on 3 outfits. VLMs priced it at $19 to $104
A researcher tested six major vision language models (Claude, GPT-4o/5.6, Grok, Kimi, DeepSeek) by asking them to price a $2.43 necklace photographed in three different outfits—formal, party, and casual—plus a standalone flat-lay image. The models' valuations varied dramatically based on context, with estimates ranging from $19 to $104 for the identical item, demonstrating a significant "halo bias" where clothing and setting influenced price assessments by up to 3.6 times. This controlled study of ~1,500 API sessions revealed that VLMs fabricate material justifications that shift with context, suggesting their valuations are driven by visual framing rather than objective product properties.
Read Full Article →