Skip to content
Dylan Castillo·

🔍Pelican-Bicycle Prompt Fails to Impress in LLM Benchmark

The famous prompt isn't as special as everyone thought

TL;DR

A new experiment challenges the relevance of the 'pelican-on-a-bicycle' benchmark, showing mixed results across seven models. The test reveals that no model significantly outperforms others on this specific prompt.

The long-standing 'pelican-on-a-bicycle' prompt used to evaluate large language models (LLMs) has been put under scrutiny in a recent experiment involving 1,008 SVGs across seven frontier models. The results show that while the pelican and bicycle combination isn't as unique or challenging as previously thought, it still garners attention from AI labs looking for easy benchmarks to beat. However, the data reveals no significant advantage for any model on this specific prompt, suggesting a need for more diverse testing methods. Each SVG was generated with identical phrasing but varying animals and vehicles, resulting in 1,008 unique prompts.

Pelican-Bicycle Prompt Fails to Impress in LLM Benchmark — Dylan Castillo

Key Points

1

7 models tested: GPT-5.6 Terra, Claude Sonnet 5, Gemini 3.5 Flash, Grok 4.5, Qwen3.7-Max, GLM-5.2, and DeepSeek V4 Pro

2

1,008 SVGs generated across 48 prompts (8 animals × 6 vehicles) with identical phrasing

3

Each model scored images using a 1-5 rating system for animal, vehicle, and coherence of action

4

Pelicans ranked 6th out of 8 animals; bicycles second from last in vehicle ratings

5

Only Gemini 3.5 Flash showed significant improvement on the pelican-bicycle prompt

Why It Matters

If you're relying on the 'pelican-on-a-bicycle' benchmark to evaluate LLMs, think again. The results suggest that this specific test may not accurately reflect model performance across diverse use cases. Teams should consider a broader range of prompts and scenarios for more reliable evaluations.

LLMBenchmarkPelican-Bicycle PromptAI Labs

Frequently Asked Questions

Why does this matter?

If you're relying on the 'pelican-on-a-bicycle' benchmark to evaluate LLMs, think again. The results suggest that this specific test may not accurately reflect model performance across diverse use cases. Teams should consider a broader range of prompts and scenarios for more reliable evaluations.

What happened?

A new experiment challenges the relevance of the 'pelican-on-a-bicycle' benchmark, showing mixed results across seven models. The test reveals that no model significantly outperforms others on this specific prompt.

Comments

Subscribe to join the conversation...

Be the first to comment

Enjoyed this article?

Get it daily. 7am. Free. Reads in 5 minutes.

Join 2,209 builders reading daily.

Also get