🔍Pelican-Bicycle Prompt Fails to Impress in LLM Benchmark
The famous prompt isn't as special as everyone thought
TL;DR
A new experiment challenges the relevance of the 'pelican-on-a-bicycle' benchmark, showing mixed results across seven models. The test reveals that no model significantly outperforms others on this specific prompt.
The long-standing 'pelican-on-a-bicycle' prompt used to evaluate large language models (LLMs) has been put under scrutiny in a recent experiment involving 1,008 SVGs across seven frontier models. The results show that while the pelican and bicycle combination isn't as unique or challenging as previously thought, it still garners attention from AI labs looking for easy benchmarks to beat. However, the data reveals no significant advantage for any model on this specific prompt, suggesting a need for more diverse testing methods. Each SVG was generated with identical phrasing but varying animals and vehicles, resulting in 1,008 unique prompts.

Key Points
7 models tested: GPT-5.6 Terra, Claude Sonnet 5, Gemini 3.5 Flash, Grok 4.5, Qwen3.7-Max, GLM-5.2, and DeepSeek V4 Pro
1,008 SVGs generated across 48 prompts (8 animals × 6 vehicles) with identical phrasing
Each model scored images using a 1-5 rating system for animal, vehicle, and coherence of action
Pelicans ranked 6th out of 8 animals; bicycles second from last in vehicle ratings
Only Gemini 3.5 Flash showed significant improvement on the pelican-bicycle prompt
Why It Matters
If you're relying on the 'pelican-on-a-bicycle' benchmark to evaluate LLMs, think again. The results suggest that this specific test may not accurately reflect model performance across diverse use cases. Teams should consider a broader range of prompts and scenarios for more reliable evaluations.
Frequently Asked Questions
Why does this matter?
If you're relying on the 'pelican-on-a-bicycle' benchmark to evaluate LLMs, think again. The results suggest that this specific test may not accurately reflect model performance across diverse use cases. Teams should consider a broader range of prompts and scenarios for more reliable evaluations.
What happened?
A new experiment challenges the relevance of the 'pelican-on-a-bicycle' benchmark, showing mixed results across seven models. The test reveals that no model significantly outperforms others on this specific prompt.
Comments
Be the first to comment
Enjoyed this article?
Get it daily. 7am. Free. Reads in 5 minutes.
Join 2,209 builders reading daily.