Cloudflare has released Clef and Clef-flash, two models built to make bounded decisions inside agent workflows rather than generate open-ended text. A developer can define fields such as urgency, destination team or severity, and receive typed choices with probabilities that ordinary code can use.

Both models are available through Workers AI and as Apache 2.0-licensed downloads on Hugging Face. They use vision as well as text, support a 64,000-token context window and are API-compatible with TypeSafe AI’s Jev decision model. Clef targets accuracy, while Clef-flash is intended for latency-sensitive decisions.

Cloudflare reports median benchmark latency of 209.3 milliseconds for Clef and 38.8 milliseconds for Clef-flash, compared with 524.1 milliseconds for Jev. Results varied by task: the models led several classification tests, while Jev remained ahead on some retrieval, escalation and trace-observability evaluations. These are Cloudflare-run benchmarks rather than independent measurements.

The company is also offering hands-on fine-tuning for customers using labeled data from their own workflows. A self-service reinforcement-learning platform is planned later, combining traffic captured through AI Gateway, container sandboxes for scoring and Workers AI deployment. Teams can use the general models now, but customized training initially requires working with Cloudflare’s engineering group.