
Clef Decision Models: Route Support Tickets in Python
Summary
Cloudflare's open Clef models return typed probabilities, not text. Build a cheap cascade router.
Most agent code has a quiet weak spot: the moment where an LLM has to decide something small. Is this ticket urgent? Which team owns it? Should the agent continue or stop? We usually answer by prompting a large model, asking for JSON, and parsing whatever comes back. It works, mostly, and it costs a full generation every time.
On October 1, Cloudflare released Clef and Clef-flash, two open-weight decision models hosted on Workers AI. You give them a state and a schema of typed questions. They return a probability for every allowed option, with no free-form text and nothing to parse. Cloudflare reports a median latency of 38.8 ms for Clef-flash and 209.3 ms for Clef, against 524.1 ms for Jev, the model that started this category. Both are Apache 2.0 on Hugging Face.
Keep reading — it's free
Enter your email to keep reading — plus the best of AI & tech, daily. Free, forever.
Already a member? Sign in
Comments
Be the first to comment