🤖Simon Willison Puts Kimi K3 Through the Pelican Benchmark
TL;DR
Simon Willison test-drives Moonshot's Kimi K3, the 2.8-trillion-parameter open-weight model, with his SVG-pelican benchmark. His read: the strongest open model yet, still a notch below Claude Fable 5 and GPT-5.6 Sol.
Simon Willison test-drives Moonshot's Kimi K3, the 2.8-trillion-parameter open-weight model, with his SVG-pelican benchmark. His read: the strongest open model yet, still a notch below Claude Fable 5 and GPT-5.6 Sol.

Key Points
Kimi K3 ships 2.8T parameters with a 1M-token context and native vision
Open weights promised July 27, 2026, the largest open release to date
Beats Claude Opus 4.8 and GPT-5.5 on Moonshot's own coding and agent benchmarks
Willison's pelican SVG test still separates it from Fable 5-class models
Why It Matters
When a hobbyist benchmark can rank a free-to-run open model against frontier labs, the moat around closed models keeps thinning.
Quick Facts
Frequently Asked Questions
Why does this matter?
When a hobbyist benchmark can rank a free-to-run open model against frontier labs, the moat around closed models keeps thinning.
What happened?
Simon Willison test-drives Moonshot's Kimi K3, the 2.8-trillion-parameter open-weight model, with his SVG-pelican benchmark. His read: the strongest open model yet, still a notch below Claude Fable 5 and GPT-5.6 Sol.
Comments
Be the first to comment
Enjoyed this article?
Get it daily. 7am. Free. Reads in 5 minutes.
Join 2,179 builders reading daily.