Cloudflare has released two open-weight Clef models built to choose among predefined options rather than generate free-form text. The design is intended for decisions such as routing a support request, classifying a bot or escalating a case to a person.
The 27-billion-parameter Clef model accepts text, JSON, images or video plus a typed question schema. It returns probabilities for each allowed answer in one forward pass, avoiding the need to parse prose. Clef-Flash is a smaller 9-billion-parameter version aimed at lower latency. Cloudflare reports median latency of 209.3 milliseconds for Clef and 38.8 milliseconds for Flash on its own infrastructure.
Both models have a 64,000-token context window and are available through Workers AI or as downloadable weights on Hugging Face. Cloudflare also offers customer-specific fine-tuning with help from its engineers; a self-service system is planned but has no announced date.
Vendor benchmarks do not show how safely the models handle unfamiliar or high-cost edge cases. For production decisions, teams will need to measure calibration, false positives and whether confidence falls when inputs shift. Typed outputs simplify integration, but they do not make the underlying judgment automatically reliable.